Peer-to-peer ad hoc network distributed conference method, device and system based on multi-modal perception
By using Bluetooth self-organizing network and multimodal sensing technology, the problem of insufficient collaborative processing capabilities in multi-terminal video conferencing is solved, achieving stable audio and video output and speaker identity management in weak network environments, thereby improving the automation level of the meeting and the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VISION INTELLIGENCE CO LTD
- Filing Date
- 2025-12-12
- Publication Date
- 2026-05-01
AI Technical Summary
Existing multi-party video conferencing technologies suffer from several issues in scenarios involving multiple terminals, including insufficient collaborative processing capabilities, imperfect online voice identity attribution and audio-video binding, and poor stability of automatic speaker framing and voice enhancement. In particular, they lack effective collaboration mechanisms and task allocation details under weak network or serverless conditions.
A distributed conferencing method based on multimodal perception and peer-to-peer self-organizing network is adopted. A serverless network is established through Bluetooth self-organizing network. Each terminal processes visual and audio data locally, performs quality assessment and negotiation, selects the optimal node for audio and video fitting and synchronization, combines visual and auditory information for speaker tracking and targeted speech enhancement, manages speaker identity through voiceprint features, and employs causal constraint noise reduction and Kalman filtering techniques to ensure stability and consistency.
It achieves stable and clear audio and video output under weak network conditions with multiple terminals, ensures the accuracy of automatic framing of speakers and the stability of voice enhancement, provides low-latency distributed processing capabilities and structured identity labeling, and improves the automation level and user experience of multi-party meetings.
Smart Images

Figure CN121967630A_ABST
Abstract
Description
Multimodal sensing-based distributed conferencing method, device, and system for peer-to-peer ad hoc networks Technical Field
[0001] This invention relates to the field of multimodal perception and machine learning technology in the field of artificial intelligence and information networks, specifically to distributed intelligent conferencing communication and edge inference technology. Background Technology
[0002] Existing multi-party video conferencing technologies have evolved around three main lines: "who is speaking, how the image is presented, and how to make the sound clearer." On the audio side, methods such as speech activity detection, frame energy and signal-to-noise ratio estimation, time difference of arrival, and beamforming are typically used. On the visual side, face and lip movement detection are commonly used to help determine the speaker and automatically compose the frame and switch the view accordingly. As the form of terminals expands from traditional fixed conferencing equipment to multiple portable terminals collaborating, conferencing scenarios exhibit characteristics such as multi-viewpoints, multiple microphones, dynamic network access, and weak network / offline conditions. The limitations of single-endpoint or centralized processing paradigms in terms of scalability, latency, and robustness are becoming increasingly prominent, necessitating a systematic approach that balances multi-terminal collaboration with low-latency processing at the endpoint.
[0003] Chinese patent document CN113920560B proposes a method to simultaneously acquire video and audio data in a conference session scenario. Face detection and lip detection are used to obtain face bounding boxes and lip shape sequences, which are then fused with audio features to determine the current speaker and their identity. This approach demonstrates good stability in complex acoustic environments and can provide a basis for subsequent automatic framing and recording annotation. However, such solutions typically complete audio and video feature extraction and fusion in a single-end or centralized processing flow, focusing more on determining "who the current speaker is." They lack online speaker identity attribution and dynamic management mechanisms for continuous sessions, such as segmenting continuous speech streams, assigning temporary identities to segments that do not match existing identities, and writing back corrections after subsequent confirmation. Furthermore, the binding relationship between voice identity results and face trajectories and avatar displays is often loosely correlated on a single device, lacking mechanisms for task collaboration, load migration, and consistency constraints between multiple terminals under weak network or serverless conditions. Cross-device time alignment, cross-viewpoint threshold linkage, and smooth switching strategies have not been engineered into a closed loop.
[0004] Among existing voice-driven speaker tracking and framing technologies, US Patent No. US9723260B2 proposes methods such as speech detection, spectral features, and sound source localization to automatically track speakers and automatically switch between panoramic views of the room and close-ups of the speaker. This approach is widely deployed in commercial conference terminals, focusing on the active directing capability of a single terminal over the room and emphasizing rapid pointing and framing control with "audio priority." Some solutions also incorporate voiceprint recognition to distinguish and label different speakers, achieving a certain degree of voice identity attribution. However, the core assumption of these solutions is still a single-device or small-scale controlled device environment, making it difficult to directly adapt to peer-to-peer network scenarios with multiple independent conference terminals participating collaboratively. They lack systematic regulations on unified threshold systems between multi-source audio and multi-view video, cross-device timestamp alignment, and switching jitter suppression, and do not adequately address the need for "unified maintenance of speaker identity labels across devices and binding them to multiple video feeds."
[0005] Key technologies for multi-party video conferencing mainly revolve around automatic speaker identification and image presentation, multi-device networking and collaborative audio pickup and playback, array-based voice processing and noise reduction / echo cancellation, edge-cloud hybrid intelligent deployment, voiceprint-driven role differentiation and transcription, and contactless gesture control. Regarding voice identity attribution, existing technologies have explored methods such as voice activity detection and segmentation of continuous audio streams, extracting speaker voiceprint features from each segment, and comparing similarity with a candidate database to achieve online speaker identification and attribution. Some have even proposed clustering in the voiceprint space to distinguish different roles. However, these technologies are mostly based on single-channel audio or central server-side processing, primarily addressing the problem of "distinguishing different speakers on a single stream." They rarely consider how to collaboratively maintain the candidate database, uniformly assign temporary and formal identities, and rewrite and correct historical segments in scenarios with multiple terminals simultaneously picking up audio and multiple audio streams coexisting.
[0006] Taking automatic speaker framing as an example, some publicly available methods use a single camera to capture wide-view images and perform region-of-interest slicing and composition optimization on the device side to center the speaker's display. For instance, CN104349112B proposes determining the framing area based on visual features such as motion and facial features and outputting optimized video. Meanwhile, some publicly available methods attempt to complete participant annotation and display from multi-source media streams. For example, CN101952852A proposes annotating the input media stream at the identity and coordinate levels for better interface presentation. Some solutions further introduce voice cues or voiceprint information, binding the speaker's identity with name tags or avatars on the interface, which can be considered a form of audio-visual joint presentation. However, overall, these technologies are still mainly based on single-terminal or centralized processing. Speaker identification on the audio side and facial trajectory binding on the visual side are usually completed at a single point, lacking asynchronous correction and history writing mechanisms for unified temporal alignment, cross-terminal evidence fusion, and error binding across multiple terminals and perspectives.
[0007] Regarding multi-device collaboration, existing solutions mostly focus on wireless interconnection and link synchronization, emphasizing expanding sound pickup / playback and coverage over a larger space. However, their implementation is often limited by hardware form factors and network architecture. For example, CN111586523B uses a specific form of "wireless headset + battery box" to achieve local / near-end dual-source acquisition and threshold fusion, solving the problem of local sound pickup and transmission. Another example is CN110310637A, which proposes using Bluetooth Mesh networking to aggregate multi-point voice and control, suggesting a feasible decentralized path. Overall, existing publications have shown a trend of expanding the meeting space with short-range wireless or Mesh methods. However, under the premise of "multiple independent meeting terminals as a unified system," there is still a lack of an engineering framework with operational details regarding voice identity attribution, task orchestration for joint audio and video attribution, unified time reference and alignment across terminals, and robust collaboration under weak network / offline conditions.
[0008] In microphone array speech processing and far-field pickup, published literature generally employs a combination of sound source localization, beamforming, noise suppression, and echo cancellation to improve intelligibility and stability. Some solutions introduce additional sensors to assist localization; for example, CN102800325A improves azimuth estimation accuracy through ultrasonic ranging and time difference compensation. However, these solutions mostly use "single device / single array" as the basic unit, lacking systematic processing of "cross-device energy trajectory fusion, consistency constraints, and switching jitter suppression" in multi-terminal parallel pickup scenarios. They also fail to integrate array processing results with voiceprint-based online voice identity attribution and audio-video joint attribution strategies, resulting in limited adaptability to the constantly changing speaker positions and sound field conditions in complex conference rooms.
[0009] In terms of interactive control, existing gesture recognition controls are mostly based on the idea of comparing front-end cameras with templates / features, which is suitable for close-range and stable lighting conditions. In meeting environments, the risks of crowd occlusion, posture changes, and false triggers are high, usually requiring joint verification with facial recognition or voiceprint cues to suppress erroneous operations. However, existing publicly available regulations on the design of security thresholds, time window alignment, and multi-factor linkage constraints for the specific application of "meeting control" are not detailed enough, resulting in the need for strong secondary engineering in actual deployment to ensure reliability and usability.
[0010] In summary, while existing technologies can achieve automatic speaker framing, localized wireless extended audio pickup and playback, single-end array voice enhancement, general end-to-cloud offloading, and speaker recognition based on voiceprint, as well as audio-visual multimodal judgment, they generally suffer from the following shortcomings: First, they lack a peer-to-peer self-organizing network collaboration mechanism and dynamic task allocation details for "multiple independent conference terminals" under serverless or weak network conditions. They cannot migrate computational tasks such as speaker location, voice identity attribution, image stitching, and voice enhancement to the optimal node for execution based on energy, signal-to-noise ratio, and node capability scores, and maintain a unified speaker identity label across the entire network. Second, they lack cross-terminal evidence fusion and... The stable switching rules are incomplete. A linkage threshold system for "frame energy difference, signal-to-noise ratio difference, arrival time difference, face detection confidence and lip movement activity correlation" has not yet been established. There is also a lack of mechanisms for performing temporal filtering on multi-source audio and multi-view video under a unified time reference, realizing audio-video joint attribution, asynchronous error correction and history writing. Thirdly, the audio front end lacks distributed collaboration and consistency constraints of "depth mask estimation - statistical post-processing" under causal constraints, making it difficult to maintain low latency and robustness under conditions of multiple people, multiple views and dynamic noise. It also fails to effectively integrate the noise reduction output quality with the upper-level voice identity attribution and audio-video joint attribution strategies.
[0011] Therefore, it is necessary to propose a distributed conferencing technology solution that integrates visual and auditory cues and enables task collaboration, timing alignment, and robust switching among multiple terminals in a peer-to-peer self-organizing network (including Bluetooth self-organizing network) environment. On the terminal side, it realizes online voice identity attribution based on voiceprint, temporary identity management, and history writing, and tightly binds and jointly corrects the facial trajectory and representative avatar on the video side. At the same time, it combines gesture or voiceprint multi-factor verification to improve the security and reliability of conference control. Summary of the Invention
[0012] To address the aforementioned technical problems, this invention provides a peer-to-peer self-organizing network distributed conferencing method, device, and system based on multimodal perception to solve the problems of insufficient collaborative processing capabilities of existing multi-party video conferencing under weak network conditions with multiple terminals, imperfect online voice identity attribution and audio-video joint binding, and poor stability of automatic speaker framing and voice enhancement.
[0013] A distributed conferencing method for peer-to-peer ad hoc networks based on multimodal awareness, the method comprising:
[0014] Multiple conference terminals establish a serverless peer-to-peer network via Bluetooth self-organizing network protocol, and each conference terminal has local data processing and decision-making capabilities.
[0015] Each conference terminal in the peer-to-peer network collects its own image data and array microphone audio data, performs visual speaker detection and array microphone-based sound source localization locally, and preprocesses and evaluates the image data and audio data to obtain image quality evaluation results and audio quality evaluation results.
[0016] Each conference terminal negotiates with peers based on local quality evaluation results, selects the terminal with the best image quality as the best image node and the terminal with the best audio quality as the best audio node in the peer network, and performs audio and video fitting and synchronous output between the best image node and the best audio node.
[0017] Speaker tracking and directional speech enhancement are performed based on visual and auditory information. The conference terminal determines the speaker's position in an image frame based on visual detection and estimates the speaker's spatial azimuth angle accordingly. The spatial azimuth angle is used as guidance information to control the array microphones to perform directional sound source localization and beamforming to directionally enhance the speaker's speech signal.
[0018] Furthermore, the method also includes an identity attribution step, which includes:
[0019] By combining lip movement information from the visual side with the speech energy timing from the auditory side, the current speaker is determined and a binding relationship is established between the current speaker's identity and their corresponding facial trajectory.
[0020] The voiceprint features extracted from the current speech segment are compared with a dynamically maintained speaker candidate library. The speaker candidate library includes voiceprint features of confirmed identities and voiceprint features of unconfirmed temporary identities. When the similarity with all identities in the candidate library is lower than a preset similarity threshold, a temporary identity is created for the current speech segment and the temporary identity is marked and output in real time. At the same time, the accumulated speech data corresponding to the temporary identity is continuously monitored.
[0021] When the temporary identity meets the preset confirmation conditions and is upgraded to a formal identity, an identity confirmation event is triggered within the peer-to-peer network, and the historical voice segment identifier corresponding to the identity is updated collaboratively.
[0022] Furthermore, the peer negotiation and task allocation include:
[0023] Each conference terminal independently calculates its own image quality score and audio quality score.
[0024] The image quality score is calculated based on a preset meeting mode and includes at least a weighted evaluation of face detection confidence, the proportion of the target in the image, and the stability of the image composition.
[0025] The audio quality score is calculated based on a preset noise reduction mode, including at least the signal-to-noise ratio after patterned noise reduction processing, the consistency with the azimuth locked according to the spatial azimuth angle, and the evaluation of residual interference.
[0026] Based on the above scoring results, the terminal with the highest image quality score is selected as the optimal image node in the peer-to-peer network, and the terminal with the highest audio quality score is selected as the optimal audio node. The data of the optimal image node and the data of the optimal audio node are then combined for fitting and output.
[0027] Furthermore, the multimodal fusion and switching rule is as follows: when the speaker's mouth movement is detected on the visual side, and / or the face detection confidence reaches the first threshold, and / or the target's proportion in the frame exceeds the second threshold and continues to exceed the first time threshold, a new candidate speaker is determined to have appeared.
[0028] Within the time window corresponding to the first time threshold, if the audio side's orientation estimation based on the array microphone, beamforming output energy, or signal-to-noise ratio after noise reduction is consistent with the spatial location of the candidate speaker and reaches the third threshold, then the switch to the candidate speaker's image is confirmed, and its corresponding terminal is designated as the speaker's associated node.
[0029] Furthermore, the speech signature detection on the visual side employs a determination method based on the correlation between mouth shape activity and speech pattern activity, including:
[0030] Calculate the temporal sequence of lip movement activity for the mouth region of the face trajectory on the visual side;
[0031] The correlation between the lip-shape activity time sequence and the corresponding speaker's speech activity time sequence is measured within the allowable time lag range;
[0032] Only when the correlation is greater than the first correlation threshold, and the deviation between the speaker's position estimated by vision and the spatial orientation obtained by the sound source localization on the auditory side is less than a preset angle threshold, a three-way binding relationship between the current speaker's identity, facial trajectory and spatial location is established.
[0033] Furthermore, the triggering conditions for identity verification include:
[0034] Condition 1: When the cumulative speaking time of a temporary identity reaches the first duration threshold, the temporary identity will be automatically upgraded to a formal identity and a global history rewrite will be triggered.
[0035] Condition 2: When the cumulative speaking time is between the first time threshold and the second time threshold, the identity verification event is triggered during the silence period after the end of the current round of speaking.
[0036] Condition 3: When the cumulative speaking time is lower than the second duration threshold, the voice segment corresponding to the temporary identity is determined to be an invalid segment, or when the identity confirmation event is triggered, the voice segment is merged back and assigned to the existing speaker with the highest similarity and a similarity greater than or equal to the preset merge similarity threshold.
[0037] Wherein, the first duration threshold is greater than the second duration threshold, the triggering logic is executed at the optimal audio node or the current computing node, and the result is synchronized to other terminals through the peer-to-peer network.
[0038] Furthermore, the voiceprint feature comparison employs a dual discrimination logic, including:
[0039] When the voiceprint similarity of the current speech segment is greater than or equal to the first similarity threshold, the speech segment is directly assigned to the corresponding existing identity;
[0040] If the voiceprint similarity of the current speech segment is between the second similarity threshold and the first similarity threshold, and the current speech segment is immediately following the previous speech segment in time and the previous speech segment has been attributed to a certain speaker, the current speech segment is determined to belong to the same speaker as the previous speech segment based on the contextual adjacency relationship.
[0041] A new speaker is identified and a corresponding temporary identity is created only when the voiceprint similarity of the current speech segment is lower than the second similarity threshold and there is no contextual basis that satisfies the adjacency rule;
[0042] Wherein, the first similarity threshold is greater than the second similarity threshold.
[0043] Furthermore, the conference terminal also uses a single camera with a wide field of view to intelligently crop and correct the perspective of the current speaker, so that the speaker is located in the central area of the output screen and forms a stable viewfinder; the audio and video data streams are aligned according to timestamps and time series fusion and smoothing are performed using Kalman filtering and hysteresis mechanisms to suppress frequent shaking and switching of the speaker's image.
[0044] Furthermore, the audio-side noise reduction employs a patterned noise reduction strategy, including at least one of the following operating modes:
[0045] Directional mode: Based on the target orientation obtained by visual detection, the beam is locked and directional enhancement is performed at the target orientation;
[0046] Voiceprint mode: Under pre-registration or password triggering conditions, noise reduction and enhancement are applied to the speech components based on the voiceprint characteristics of the target speaker;
[0047] General mode: When stable visual or voiceprint cues are not available, general noise reduction is applied to the omnidirectionally acquired speech signals.
[0048] Furthermore, the noise reduction task can be collaboratively executed by multiple conferencing terminals within the peer-to-peer network, including:
[0049] Based on a weighted score composed of frame energy, signal-to-noise ratio, terminal processing capability, and task priority of the speech frame, noise reduction tasks are dynamically allocated to each terminal. The task negotiation adopts a point-to-point communication protocol and carries task description, device status, priority, and threshold parameter fields.
[0050] The terminals use a unified time base or timestamp alignment to synchronize the speech frames and fuse the temporal energy trajectories from multiple terminals to constrain the consistency of the noise reduction output.
[0051] The noise reduction model uses a joint loss function that includes a time-domain signal-to-noise ratio term, a speech component term, and a noise component term during training, and assigns higher weights to low-frequency subbands and subbands with prominent non-steady-state noise.
[0052] The end-to-end processing latency of the noise reduction front-end is no higher than a preset latency threshold, and it is deployed between the audio input and speech recognition modules.
[0053] Furthermore, when the audio-side noise reduction is in general mode, the acquired speech signal undergoes real-time noise reduction processing, and the noise reduction processing method includes:
[0054] Perform a short-time Fourier transform on noisy speech to obtain complex spectral features;
[0055] The complex spectral features are input into a neural network that includes encoding, enhancement, and decoding and has skip connections to generate an initial spectral mask.
[0056] The initial spectral mask is optimized based on causal time-weighted filtering that uses only the time-frequency information of the current frame and its preceding frames, without using information from subsequent frames.
[0057] The optimized spectral mask is corrected according to the minimum mean square error criterion, the prior signal-to-noise ratio and the posterior signal-to-noise ratio are calculated, and the final complex spectral mask is obtained according to the preset frequency point gain formula, so that the complex spectral mask weights the real part and the imaginary part of the spectrum respectively within the amplitude limit range.
[0058] The final complex spectral mask is multiplied by the noisy spectrum and then subjected to an inverse short-time Fourier transform to output the denoised speech.
[0059] Furthermore, the noise reduction method also includes the following training data construction process:
[0060] Acquire room impulse response, clean speech and noise data, and adjust the reverberation duration of the room impulse response to keep the target reverberation time between 0.10 seconds and 0.60 seconds and maintain its decay pattern;
[0061] The clean speech was convolved with the adjusted room impulse response to obtain the reverberant speech, and one or more noise sources were selected with a probability of less than 0.7. The signal-to-noise ratio of the training samples was set in the range of -10 dB to +20 dB.
[0062] The reverberant speech and / or noise are enhanced and normalized in the spectral domain, and at least one signal level is normalized to a predetermined range. Then, the signal is mixed according to the signal-to-noise ratio strategy to form the training input, with the clean speech as the training target.
[0063] Furthermore, the noise reduction neural network used in the noise reduction processing method adopts a U-shaped structure that includes encoding, enhancement, and decoding and has skip connections, wherein: the encoding part and the decoding part are composed of multiple levels of convolutional layers and deconvolutional layers, and except for the first level of convolutional layer of the encoding part and the last level of deconvolutional layer of the decoding part, the size of the convolutional kernel and deconvolutional kernel of the other levels is 3×1, and the size of the first level of convolutional kernel of the encoding part and the last level of deconvolutional kernel of the decoding part is 5×2;
[0064] The enhancement part is grouped along the channel dimension, and each group is configured with a cyclic unit for independent calculation. The outputs of each group are then fused after instantaneous normalization.
[0065] The skip connection performs channel matching through 1×1 convolution and fuses with the decoding side features by feature addition. The network outputs a complex spectral mask and limits the amplitude of the mask to the range of 0 to 1.
[0066] The causal temporal weighted filtering applies a weighted average of the number of back-look-back frames L and the neighborhood width K of adjacent frequency points to each frequency point. The weights are non-negative and normalized on a frequency-by-frequency-point basis, and are generated by gated convolution or attention mapping.
[0067] The minimum mean square error post-processing adopts any one of Wiener gain, logarithmic amplitude minimum mean square error, or spectral amplitude minimum mean square error. The prior signal-to-noise ratio is obtained by exponential smoothing of the previous time-estimation and the current posterior signal-to-noise ratio using the decision-oriented method. The noise power spectrum is estimated using online noise tracking or speech presence probability methods.
[0068] Furthermore, the identity attribution step also includes face capture and avatar generation based on audio-video joint binding, used to generate representative avatars corresponding to the speaker's identity. Specifically, for each speaking event detected by the audio side, within a preset capture window after the start of the speech, i.e., within the time interval from the start time of the speech plus a first time offset to the start time of the speech plus a second time offset, face detection is performed on the video stream; target face trajectories are filtered based on the voiceprint identity to which the current speech segment belongs and the degree of matching with the lip movement activity of the speech segment, excluding faces of non-current speakers who have just finished speaking; among the retained faces, candidate avatars are determined based on image clarity, frontal angle, and detection confidence, and the face image with the highest comprehensive quality index and the number of consecutive frames reaching a preset frame threshold is selected as the representative avatar bound to the current speaker's identity, and is output synchronously along with the identity identifier.
[0069] Furthermore, the facial image corresponding to the representative avatar is bound to the current speaker's identity and stored as a mapping relationship between the speaker identifier and the facial trajectory identifier, while the binding confidence level is recorded. When a higher quality facial image of the speaker is captured in a subsequent meeting, or when an error is found in the initial binding based on long-term clustering analysis, the mapping relationship and its corresponding facial image and binding confidence level are updated asynchronously in the background. The updated binding information is used to refresh the interface display, but does not interrupt the current real-time transcription output and speaker attribution results.
[0070] Furthermore, when the video signal is missing or the current video quality is below a preset threshold, the speaker attribution result on the audio side continues to be output, and the face image corresponding to the most recent valid audio-video joint binding relationship is reused as the display prior; when the video signal is restored and the detected face meets the preset association conditions in terms of temporal correlation and similarity with the current speaker, the audio-video related matching is re-executed to correct the binding relationship between the speaker identity and the face capture.
[0071] A multimodal sensing-based peer-to-peer self-organizing network distributed conferencing terminal device, used to implement any of the above methods, includes:
[0072] An array microphone assembly is used to acquire local conference audio signals and support time difference of arrival measurement and sound source location estimation.
[0073] Camera components are used to acquire image data related to speaker detection;
[0074] The Bluetooth communication unit is used to establish a serverless peer-to-peer network with other conference terminals according to the Bluetooth self-organizing network protocol and to conduct point-to-point communication in order to exchange local image quality evaluation results and audio quality evaluation results and negotiate the optimal node.
[0075] The synchronization unit is used to provide a unified time reference and / or align with external timestamps to ensure the time comparability of audio and video data collected by each terminal.
[0076] A processor and a memory, wherein the memory stores program instructions that run on the processor, and the processor is configured to:
[0077] Local preprocessing and quality assessment are performed on the data collected by the camera component and array microphone component to generate image quality assessment results and audio quality assessment results;
[0078] The Bluetooth communication unit negotiates with other conference terminals, selects the terminal with the best image quality as the best image node and the terminal with the best audio quality as the best audio node in the peer network, and completes audio and video fitting and synchronous output between the best image node and the best audio node.
[0079] Based on visual detection, the current speaker's position in the image frame is determined and its spatial azimuth angle is estimated. The spatial azimuth angle is used as guidance information to control the array microphone to perform directional sound source localization and beamforming, thereby directionally enhancing the speaker's voice signal.
[0080] Furthermore, when executing the program instructions, the processor is configured to implement the following functional modules:
[0081] The audio acquisition and processing module is used to receive continuous audio streams and perform speech activity detection and speech segmentation;
[0082] The voiceprint feature extraction module is used to extract the speaker's voiceprint feature vector from the speech segment;
[0083] The candidate database comparison and identity determination module is used to compare the voiceprint features with the maintained speaker candidate database. The speaker candidate database includes confirmed speaker voiceprint features and unconfirmed temporary speaker voiceprint features. Based on the similarity and time context, it determines whether the current speech segment belongs to an existing speaker identity or creates a new temporary speaker identity.
[0084] The temporary identity management module is used to manage the identity of the temporary speaker, including recording its cumulative voice duration, monitoring its status, and triggering an identity confirmation event when preset confirmation conditions are met.
[0085] The decision and write-back module is used to upgrade the corresponding temporary identity to the formal speaker identity when the identity confirmation event is triggered, and to batch update the identity identifier of the historical voice segment corresponding to the temporary identity to the confirmed identity and store it in a fixed manner.
[0086] The central scheduling module is used to drive the online processing pipeline in a sliding window manner, serially traversing each processing stage and triggering corresponding processing and identity verification events based on conditions.
[0087] The output module is used to output the speech recognition transcription results with speaker identification information.
[0088] The global post-processing module is used to perform consistency checks and final cleanup on the real-time speaker attribution results at the end of audio stream processing or at the end of a predetermined time window, so as to improve the consistency and accuracy of the results.
[0089] Furthermore, when executing the program instructions, the processor is configured to further implement the following functional modules:
[0090] The video acquisition and face trajectory module is used to acquire a video stream synchronized with the audio stream, perform face detection on the video frames, and output the face trajectory and its quality score.
[0091] The lip-sync activity calculation and audio-visual correlation matching module is used to calculate the lip-sync activity time series of the mouth region of the face trajectory, and to measure the correlation and threshold determination of the lip-sync activity time series with the speech activity or speech energy time series corresponding to each speaker identity within an allowable time lag range, so as to establish a binding relationship between speaker identity and face trajectory; the face selection and binding management module is used to screen candidate faces within a preset time window, complete the binding of speaker identity and face capture, record and update the binding confidence, and when a higher quality face image is subsequently obtained or an error is found in the initial binding through long-term statistical analysis, the binding relationship and its face capture are asynchronously corrected and updated. When the video signal is missing or the quality is substandard, the speaker attribution result on the audio side is maintained and the most recent valid face binding is reused as the display prior. When the video signal is restored and the preset correlation conditions are met, the audio-visual correlation matching is re-executed to correct the binding relationship.
[0092] A peer-to-peer self-organizing network distributed conferencing system based on multimodal awareness, used to execute the method described in any of the above embodiments, includes multiple conferencing terminal devices and at least one conferencing display device, wherein:
[0093] The multiple conference terminal devices are deployed in the same offline physical conference room. Each conference terminal device establishes a serverless local peer-to-peer network through its own Bluetooth communication unit according to the Bluetooth self-organizing network protocol. This network is used to collect audio and / or video data from the local end and to perform control processing corresponding to the method described above locally.
[0094] The conference display device is located in the physical conference room and is connected to at least one conference terminal device in the local peer-to-peer network to receive conference processing results generated by the conference terminal device.
[0095] The beneficial effects of this invention are as follows:
[0096] This invention utilizes Bluetooth self-organizing networking to form a serverless peer-to-peer network. Multiple conference terminals automatically establish point-to-point connections and share status information under the mechanism of "connection upon joining, discovery upon joining, and collaboration upon joining." Each terminal possesses local data processing and decision-making capabilities, maintaining continuous and stable operation of the conference service even when the external network is unavailable or there is significant link jitter. By eliminating the central node and distributing processing power among the terminals, the single-point bottleneck and backhaul latency of the central node are avoided. Simultaneously, it supports plug-and-play and horizontal expansion when adding or removing terminals in the conference room, providing a low-latency, scalable distributed operating foundation for subsequent edge-side algorithms such as visual inspection, array positioning, noise reduction, and broadcast directing.
[0097] Each conference terminal locally evaluates the quality of its images and audio and participates in peer-to-peer negotiation. Based on image metrics such as face detection confidence, target proportion in the frame, and compositional stability in conference mode, and audio metrics such as processed signal-to-noise ratio, consistency with visual orientation, and residual interference in noise reduction mode, the system selects the optimal image node and optimal audio node within the peer-to-peer network. The system then performs audio-visual fitting and unified time-base synchronization on the optimal image and audio, ensuring that the speaker remains prominent and speech clarity is maintained even under conditions of personnel movement, changing viewpoints, and sound field disturbances. This reduces the probability of accidental switching and frequent re-switching, thereby improving the overall visual appeal of the conference screen and the intelligibility of the speech.
[0098] To improve the accuracy and stability of speaker tracking and switching, this invention uses visual evidence as the primary method for speaker tracking. When mouth movements are detected, face detection confidence reaches a threshold, or the target proportion consistently exceeds a threshold, a new candidate speaker is generated, and their spatial azimuth is estimated. This azimuth is used to guide the array microphones in performing directional sound source localization and beamforming, achieving focused sound pickup and interference suppression for the candidate speaker. Within a limited time window, if the audio-side azimuth estimation, beam output energy, or signal-to-noise ratio after noise reduction matches the visual position and meets a set threshold, a smooth switch to the speaker's visual and audio path is completed. This "visual-driven, audio-supported" strategy significantly reduces misjudgments and incorrect switching, maintaining stable focus on the current true speaker even in multi-person conversations or with strong background noise.
[0099] Furthermore, in terms of end-side noise reduction, the invention employs a causal-constrained complex spectrum masking noise reduction scheme at the terminal side. After the noisy speech undergoes a short-time Fourier transform to obtain a complex spectrum, an initial complex spectrum mask is first generated through a neural network containing encoding / enhancement / decoding and skip connections. Then, causal time-weighted filtering is performed using only the local and adjacent frequency information of the current frame and its preceding frames to obtain an optimized mask. Subsequently, the prior signal-to-noise ratio and posterior signal-to-noise ratio are calculated based on the minimum mean square error criterion and corrected according to a preset frequency gain, so that the final mask weights the real and imaginary parts of the complex spectrum within a limited amplitude range. The entire processing flow does not depend on future frames, has low end-to-end latency and high stability, and can perform consistent fusion of multi-source energy trajectories under a unified time reference, thus achieving robust end-side noise reduction and higher speech intelligibility even in noisy, reverberant, and multi-person conversation environments.
[0100] This invention implements global timing management for the "speaker positioning—command issuance—video / audio switching" link. Terminals use a unified timestamp or clock alignment. On the output side, strategies such as Kalman filtering, hysteresis, and minimum dwell time are used to perform time-series fusion and smoothing of the candidate speaker's confidence trajectory, suppressing frequent jitter and jumps caused by short-cycle fluctuations in scores. In terms of visual presentation, intelligent cropping and perspective correction are applied to wide-viewing-angle images to ensure the current speaker is stably centered. Under multi-terminal conditions, it supports image stitching and consistent presentation, achieving a more focused and continuous viewing experience while reducing the interference of image switching on participants' attention. During peer-to-peer negotiation, a lightweight consensus protocol or token-based arbitration is used to resolve simultaneous competition among multiple terminals. The arbitration result is valid within a limited time window. When the computing power or latency of the current optimal node does not meet the threshold, the system can automatically switch to a nearby suboptimal node to maintain continuous output.
[0101] Furthermore, to ensure accurate, continuous, and easily traceable speaker identification during meetings, this invention introduces an online speaker attribution and identity management mechanism based on voiceprint features. The continuous audio stream is segmented by speaking events, and voiceprint features are extracted from each speech segment and compared with a dynamically maintained candidate library. A dual threshold and contextual adjacency rule are used to determine whether an existing identity is assigned or a new temporary identity is created. When the accumulated speech of a temporary identity meets preset conditions, an identity confirmation event is triggered, upgrading the temporary identity to a formal identity, and performing a batch "write-back" of its corresponding historical speech segments. This mechanism fully utilizes posterior information to correct early judgments while maintaining real-time output, significantly improving the accuracy and consistency of speaker identification in long-duration conversations, reducing the workload of manual verification in the post-processing stage, and providing version-managed identity records for audit traceability.
[0102] Furthermore, to facilitate the matching of voices with specific speakers in a meeting setting, this invention achieves a joint binding of "speaker identity—facial trajectory—spatial location" by measuring the correlation between lip movement activity and the timing of speech activities. Within the capture window after the speaker begins, facial image capture and quality optimization are performed, generating a representative avatar with high clarity, appropriate angle, and confidence for each confirmed identity. The system supports asynchronous correction when higher-quality facial images are subsequently captured or when errors in the initial binding are discovered through long-term cluster analysis. The mapping relationship between identity and facial capture images is updated in the background, and the displayed avatar is maintained without interrupting real-time transcription output. In the event of missing or substandard video signals, the speaker attribution result on the audio side is maintained, and the most recent valid facial binding is reused as the display prior. Once the video signal is restored and the correlation meets a preset threshold, audio-video association matching is re-executed to correct the binding relationship. Through the above design, consistency and error correctability between the audio-side identity results and the front-end interface display are ensured, enabling each speaker to obtain a stable, clear, and updatable avatar presentation.
[0103] Overall, this invention uses Bluetooth peer-to-peer self-organizing networks as a distributed foundation, combined with a multimodal speaker tracking strategy of "visual-driven and audio-supported", a voiceprint-driven online voice identity attribution and history writing mechanism, an audio-video joint binding and asynchronous correction mechanism, and causal low-latency end-side noise reduction processing. Even in weak network / offline and complex meeting environments, it can still stably highlight the speaker, continuously output clear audio and video, and provide structured identity labeling and avatar display. It takes into account the system's real-time performance, robustness, maintainability, and security controllability, and improves the automation level and user experience of multi-party meetings. Attached Figure Description
[0104] Figure 1 is a schematic diagram of the overall system architecture of the present invention.
[0105] Figure 2 is a structural block diagram of the conference terminal of the present invention.
[0106] Figure 3 is a schematic diagram of the task negotiation and allocation process of the present invention.
[0107] Figure 4 is a schematic diagram of the multimodal speaker determination process of the present invention.
[0108] Figure 5 is a schematic diagram of the audio noise reduction process of the present invention.
[0109] Figure 6 is a schematic diagram of the U-shaped network structure of the present invention, which includes encoding-enhancement-decoding and skip connections.
[0110] Figure 7 is a schematic diagram of the skip connection fusion mechanism of the present invention.
[0111] Figure 8 is a schematic diagram of the image processing and timing synchronization process of the present invention.
[0112] Figure 9 is the main flowchart of the speaker attribution method based on voiceprint recognition of the present invention.
[0113] Figure 10 is a schematic diagram of the audio and video joint binding and face image generation process of the present invention. Detailed Implementation
[0114] The technical solution of the present invention will now be clearly and completely described in conjunction with the accompanying drawings. In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention. The terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying their relative importance.
[0115] In an embodiment of a multimodal awareness-based peer-to-peer self-organizing network distributed conferencing method of the present invention, multiple conferencing terminals complete network construction, task coordination, and multimodal processing without the participation of a central server. A typical process is as follows: after powering on, participating terminals 200 automatically discover and securely pair according to the Bluetooth self-organizing network protocol, sequentially completing discovery, handshake, link establishment, and network entry to form a decentralized peer-to-peer topology. As shown in the overall system architecture diagram in Figure 1, the peer-to-peer self-organizing network 110 establishes point-to-point data channels between terminals and maintains adjacency relationships and routing tables. It uses lightweight heartbeats to periodically broadcast node scores, available bandwidth, received signal strength, packet loss rate, and latency jitter, combined with rapid failure detection and shortest path recalculation to achieve instant node joining, leaving, and topology self-healing. The coordination process is driven by a message state machine: six stages—candidate, competition, grant, execution, monitoring, and transfer / takeover—are advanced across the entire network in a timestamp-ordered sequence. When multiple nodes simultaneously announce the same task, the overall score is compared first, followed by the timestamp. Conflicts are then resolved using the deterministic order of node identifiers. Unselected nodes are moved to hot standby for seamless subsequent takeover. To ensure the stability and timing consistency of audio and video processing, the system enables parallel LAN data synchronization 120: when LAN coverage is available, it undertakes auxiliary transmission of video and metadata and distributes time bases to each terminal. Time synchronization adopts a two-level approach of "coarse alignment + fine calibration." First, the deviation is converged to the millisecond level through multicast time synchronization. Then, fine calibration is performed by exchanging timestamps and round-trip delays on the point-to-point link to stabilize the relative deviation to the sub-millisecond level. Synchronization parameters are carried with the heartbeat and used to uniformly stamp speaker switching, subtitle anchors, video splicing, and audio segmentation. When the LAN is unavailable or the link quality degrades, the system automatically reverts to pure peer-to-peer synchronization, retaining only the data channels for audio and control priorities, and enabling bandwidth adaptation and buffer shaping to ensure continuous execution of multimodal processing and task collaboration under weak network conditions.
[0116] Based on the above, after completing network access, each conference terminal 200 broadcasts its own capabilities and environmental indicators in two ways: fixed periodicity and event triggering. These include processor usage and temperature, battery level and health, available memory and storage space, microphone array channel count and calibration status, current signal-to-noise ratio and frame energy statistics, echo cancellation convergence, uplink and downlink bandwidth, round-trip latency and jitter, packet loss rate, and local clock deviation. Each indicator is first smoothed locally using a sliding window to remove outliers, and then mapped to a node capability vector. This vector is then aggregated by other terminals within the network to form a comprehensive view of the network's node capabilities. The task orchestration module calculates a weighted score based on this and dynamically allocates tasks according to task characteristics and thresholds: tasks with high computational and memory usage, such as image stitching, multi-view cropping, and rendering, are prioritized for terminals with sufficient computing power and good heat dissipation. For latency-sensitive and low-bandwidth tasks like voice front-end noise reduction, echo suppression, dual-talk detection, and sound source localization, these are prioritized for nearby terminals with high signal-to-noise ratios and better pickup conditions. Recognition and translation are matched in a tiered manner based on the bundle width and window size, while meeting the latency budget, in order to balance real-time performance and stability.
[0117] To avoid frequent jitter, task allocation adopts a three-stage rule of "trigger threshold - hold threshold - minimum dwell time," and uses a message state machine to complete closed-loop control of candidate selection, competition, granting, execution, monitoring, and transfer or takeover. When multiple nodes simultaneously announce the same task, priority is given based on comprehensive score. If the scores are the same, priority is given to the one with the earlier timestamp; if they are still the same, the node identifier deterministic order is used to resolve the conflict. Unselected nodes enter hot standby state and wait for switchover at a preset takeover point. During execution, the main execution node continuously reports output confidence, processing latency, and link quality. If the threshold is exceeded consecutively, a rapid transfer is triggered, and the hot standby node seamlessly takes over under a unified timestamp. The subtitle anchor point, screen framing, and audio segments are aligned accordingly to ensure visual and auditory continuity.
[0118] In scenarios involving network degradation or short-term offline events, the peer-to-peer self-organizing network (110) prioritizes audio and control signaling channels, while video and auxiliary streams are downgraded according to a strategy: first reducing resolution and frame rate, then switching to a local single-view mode. When LAN data synchronization (120) is available, it continues to provide a global time base and video auxiliary transmission, achieving redundant support for clock and image stitching. When bandwidth is insufficient, the system enables adaptive bitrate and forward error correction, and retains only subtitles and translated audio streams when necessary. When computing power is limited or temperature exceeds thresholds, it automatically reduces the number of noise-reduced playback frames and neighborhood width, and shrinks the translation beam width and window to ensure end-to-end latency does not exceed budget. The entire process of "declaration—granting—execution—switching" is recorded with key timestamps and parameters for fault tracing and parameter tuning, thus maintaining continuous audio and video processing and stable image switching even under weak network and heterogeneous device conditions.
[0119] The hardware structure of each conference terminal 200 is shown in Figure 2. It consists of a camera 201, a microphone array 202, a processing unit 203, a communication unit 204, and a display speaker 205, forming a closed loop from acquisition, processing, to presentation. The camera 201 acquires wide-angle video signals and outputs visual elements such as faces, mouth shapes, and subject positions. The microphone array 202 acquires multi-channel audio, providing basic data for voice activity detection, frame energy and signal-to-noise ratio evaluation, time difference of arrival, and sound source direction determination. The two are installed adjacent to each other at the front end of the terminal, and their geometric positions and delay errors are calibrated at the factory. After network access, the processing unit 203 sends out timestamp alignment parameters to align the audio and video on the same timeline for multi-terminal speaker determination and screen switching. To ensure framing quality, the camera 201 supports wide dynamic range and distortion correction, and the framing window and cropping parameters are transmitted back with timestamps along with the frames. The microphone array 202 features channel amplitude and phase calibration and unified clock distribution. Board-level equal-length traces and isolation design reduce crosstalk, providing stable input for subsequent echo cancellation, dual-talk detection and causal noise reduction.
[0120] The processing unit 203 integrates computing and storage resources, pre-configured with speech denoising, recognition, translation, and speech synthesis models, as well as multimodal fusion and image processing algorithms, supporting local offline inference. Internally, it prioritizes audio front-end and timing alignment execution using thread priority and affinity binding, dynamically adjusting the decoding window, beamwidth, and number of denoising playback frames based on temperature, power consumption, and load to ensure end-to-end latency meets thresholds. The storage area includes secure boot and trusted storage, with model weights and meeting control credentials encrypted and stored. Anomaly reset is triggered by a watchdog timer and retains critical logs. The processing unit 203 interconnects with the perception module via a high-speed video and audio bus, issuing calibration coefficients, AEC parameters, and framing commands through the control bus, and uniformly distributing time bases and frame-level timestamps externally to drive consistent rendering of subtitles, translations, and images on the rendering end.
[0121] Communication unit 204 handles both peer-to-peer ad hoc network and local area network (LAN) access: On the short-range wireless side, it performs automatic discovery, pairing, and establishes point-to-point audio and control channels, prioritizing low-latency audio and negotiation signaling. On the LAN side, it provides auxiliary transmission of video and metadata and distributes time bases, converging clock deviations to sub-millisecond levels when coverage is available. Communication unit 204 reports real-time link metrics such as received signal strength, packet loss, round-trip time, and jitter to processing unit 203, which are used as input for task allocation, bitrate adaptation, and buffer control. When the link degrades, it automatically degrades to local framing and caption / translation-only mode to ensure conference continuity.
[0122] The display speaker 205 is responsible for text and voice output. The display side supports subtitle overlay, speaker identification, and multilingual switching, presented aligned with the timestamp and viewfinder window. The speaker side supports segmented playback and loudness management, and is linked with echo cancellation parameters to prevent backflow. The entire unit adopts zoned heat dissipation and overcurrent, overvoltage, and undervoltage protection in its structure and power supply. The heat source area is equipped with heat conduction channels and temperature monitoring. In long conversation scenarios, bypass power supply and low power mode can be enabled. In privacy mode, the terminal only processes and temporarily stores necessary features locally, without persisting facial and original voice data. Anonymous images and numbered identifiers are output when necessary. Through the above hardware and control collaboration, the terminal shown in Figure 2 can complete multimodal acquisition, peer-to-peer collaboration, and low-latency presentation under weak network and heterogeneous conditions, providing a stable foundation for subsequent dynamic task allocation, speaker tracking, real-time noise reduction, and cross-language processing.
[0123] The aforementioned components are connected via an internal bus to form a unified data and control path. The camera 201 and microphone array 202 continuously provide a synchronous data stream to the processing unit 203. The communication unit 204 maintains a self-organizing network and time synchronization in the background. The processing unit 203 completes noise reduction, recognition, translation, and multimodal fusion according to the method flow of this invention and generates control commands. The display speaker 205 performs subtitle overlay and voice playback according to the commands. This structure can operate in a closed loop within a single terminal. When multiple terminals collaborate, consistent screen control and voice output are achieved through peer-to-peer self-organizing networks and local area network synchronization, providing hardware and data path guarantees for dynamic task allocation, speaker tracking, and real-time noise reduction and cross-language processing under serverless conditions.
[0124] After the peer-to-peer network is formed, each conference terminal negotiates and completes task allocation based on real-time environmental indicators, as shown in Figure 3. Each terminal continuously extracts audio features such as frame energy, signal-to-noise ratio, time difference of arrival, and frequency band energy distribution using microphone array 202. Simultaneously, it aggregates the instantaneous load, temperature, power consumption, available bandwidth, packet loss rate, and jitter data from processing unit 203, combining them into a node capability vector and calculating a weighted score. The scoring employs sliding window smoothing and outlier suppression to avoid short-term peak interference. Each terminal periodically exchanges scores and key statistics using small broadcast messages, forming a distributed node capability view to provide a basis for subsequent decision-making. To reduce jitter, each terminal locally maintains a task and priority table, listing allocable items for speaker positioning and multimodal fusion, image cropping or splicing, and voice front-end enhancement (including noise reduction, echo suppression, and dual-talk detection). Trigger thresholds, hold thresholds, minimum dwell time, and maximum processing latency limits are set for each type of task. Only when environmental indicators continuously meet the trigger conditions and do not exceed the predetermined latency limit does the competition process begin. Terminals that meet the requirements first initiate a candidate declaration, with the message carrying the task identifier, current score, estimated processing latency, planned dwell time, and output interface description. If multiple terminals declare simultaneously, they are sorted by score from highest to lowest; if scores are the same, the earlier timestamp takes precedence; if still the same, a deterministic sequence is used for decision-making. Unselected terminals enter hot standby mode and remain on standby at the agreed takeover point.
[0125] After the selected terminal broadcasts confirmation to the network, tasks are initiated under a unified time base. Audio tasks are prioritized for processing on nodes with higher signal-to-noise ratios and better sound pickup conditions, and the enhanced results are streamed back to the terminals that need them. Video tasks are prioritized for nodes with more computing power and shorter links to the display, reducing cross-terminal transmission and rendering pressure. All results and control information carry a unified timestamp to ensure the consistency of speaker determination, screen switching, subtitle overlay, and translated playback timing. The network continuously monitors output confidence, processing latency, and link quality through heartbeats and quality feedback: when the confidence continuously falls below the hold threshold, the average processing latency exceeds the upper limit, or the link quality deteriorates significantly, a rapid reallocation is triggered, with a hot standby node seamlessly taking over at the reserved takeover time point, and the original node exiting after completing the current segment. To suppress frequent switching, the system only performs handover when the indicators continuously exceed the hold threshold. Short-term fluctuations are filtered through inertial hold and debouncing strategies. Negotiation and heartbeat messages use redundant transmission and small window confirmation, so a single packet loss does not affect the consistency of the entire network. When the consecutive loss reaches the limit, the adjacent node infers and temporarily takes over based on the most recent reliable state and local quality metric. Through the closed loop of "evaluation-competition-confirmation-execution-monitoring-takeover" described above, the task can stably fall to the optimal node, and a smooth migration can be achieved when the node fails or the environment changes abruptly. At the same time, the timing of each link is aligned with a unified time base, thereby maintaining low latency and high robustness in multi-terminal collaborative scenarios.
[0126] Speaker tracking and screen switching rely on multimodal fusion. As shown in Figure 4, the multimodal speaker determination process integrates visual and auditory information. Visual face confidence 410 is obtained by detecting face position and features using camera 201, while auditory confidence 430 estimates the sound source location based on the energy value, signal-to-noise ratio, and time-of-arrival difference of the microphone array 202. The linkage threshold and duration capture 430 checks whether the energy difference or signal-to-noise ratio difference between any two terminals is not less than a preset threshold and the duration meets the requirements. Simultaneously, the visual face detection confidence is required to be not lower than the threshold before determining the current speaker 440 and triggering a switch. When the conditions are met, the processing unit 203 generates a switch command, which is broadcast to all terminals via the communication unit 204, achieving automatic focusing. This fusion process aligns the audio and video streams using timestamps and employs Kalman filtering to smooth the target trajectory, suppressing switching jitter and ensuring a smooth transition from a multi-person environment to the current speaker's associated node. This multimodal linkage not only improves the accuracy of speaker identification, but also solves the problem of mis-slicing caused by inconsistent evidence across devices in the background technology, ensuring a stable visual presentation in dynamic meeting scenarios.
[0127] To improve audio quality, the system performs real-time noise reduction processing on the acquired speech at the edge, and the overall link is shown in Figure 5. The audio front-end 700 performs a short-time Fourier transform on the noisy speech to obtain a complex spectrum 710. The spectrum is fed into a U-shaped neural network containing encoding, enhancement, and decoding stages with multi-stage skip connections to generate an initial complex-valued spectral mask 720. The mask amplitude is limited to the range of zero to one to ensure numerical stability. Subsequently, it enters a causal time-series weighted filter 730, which uses only the time-frequency information of the current frame and its preceding frames at this frequency point and in the neighborhood for non-negative and frequency-point-normalized weighted fusion, without using any subsequent frames. Then, minimum mean square error post-processing 740 is implemented, calculating the gain at the frequency point based on the posterior and prior signal-to-noise ratios, and statistically correcting the mask in the real and imaginary parts respectively. Finally, the corrected complex-valued mask is multiplied by the noisy spectrum, and the denoised speech is output 750 after inverse short-time Fourier transform. To reduce end-to-end latency, the noise reduction front-end is placed between the audio input and speech recognition, and works in conjunction with echo cancellation and dual-talk detection in a serial parallel manner. The silence segment is gated by speech activity detection to suppress invalid calculations and bandwidth usage.
[0128] The network backbone unfolds as shown in Figure 6, with a U-shaped network structure of encoding-enhancement-decoding and skip connections. The input consists of complex time-frequency features. The encoding end uses multi-level small-size convolutions to progressively downsample and extract multi-scale representations. The first layer convolution kernel is 5x2, and the rest are 3x1. The top layer of the encoding enters a causal loop unit for temporal modeling, receiving only current and historical information. The decoding end restores resolution through progressive upsampling, ensuring the output matches the input in the frequency dimension. Multiple skip connections are established between the encoding and decoding ends. Same-scale features from the encoding side are injected into the decoding backbone via bypass, compensating for details and stabilizing training. In the figure, solid lines represent the backbone information flow, and dashed lines represent cross-layer injection paths. The network terminal directly outputs a complex mask weighted for both the real and imaginary parts, with amplitudes limited to zero to one, to ensure numerical stability and phase consistency.
[0129] The specific method of cross-layer fusion is implemented as shown in the schematic diagram of the skip connection fusion mechanism in Figure 7. Features from a certain coding layer are first aligned by a one-to-one channel matching convolution, and then channel-wise weights are generated by gated convolution, with weight values between zero and one. The coding-side features are first channel-weighted, and then added element-wise with the decoding backbone features of the same scale to form the decoding layer input at that scale. After convolution and upsampling in this layer, the fused features are obtained and continue to be passed to the next level of decoding layer. Compared with simple concatenation, this "channel matching, gated weighting, and element-wise addition" path reduces the computational load of channel stacking and upsampling at the decoding end, and can adaptively adjust the intensity of cross-layer information injection, maintaining detail and stability under the premise of lightweight computation.
[0130] The training phase corresponds to the input-output relationship shown in Figure 5 and the structural settings in Figures 6 and 7. The training data consists of room impact response, clean speech, and noise. The room impact response is target-processed for reverberation time, ranging from 0.1 to 0.6 seconds. The clean speech is convolved with the processed room impact response to obtain reverberated speech. One or more noise classes are selected according to a set probability, and a target signal-to-noise ratio is configured, commonly ranging from -5 to 15 dB, but can be extended to -10 to 20 dB. Before mixing, amplitude consistency normalization and lightweight spectral domain enhancement are performed on both speech and noise. The complex features after short-time Fourier transform are used as network input. The loss function consists of a time-domain signal-to-noise ratio term, a speech component term, and a noise component term, combined with frequency band weights to highlight low-frequency and non-steady-state noise bands. These are used for inference after convergence on the validation set.
[0131] Real-time speech is input into the network after short-time Fourier transform to obtain the initial mask. Then, it undergoes causal-time weighted summation and least mean square error correction to obtain the final mask. Finally, it is output as denoised speech after inverse transform. Frame segmentation parameters remain consistent with training, while thresholds and weights are loaded during device initialization. In multi-terminal collaboration, all outputs carry a unified timestamp. When necessary, the nearest node performs front-end enhancement and transmits the results back. When performance degrades, a hot standby node smoothly switches over at a preset takeover point. Through the processes, structures, and fusion mechanisms shown in Figures 5, 6, and 7, the edge device effectively suppresses non-stationary noise and reverberation tails without introducing future frame waiting, maintaining clear speech boundaries and phase consistency, providing stable, low-latency, high-quality input for subsequent recognition and translation.
[0132] Image processing 500 and time synchronization ensure the smoothness of visual output. As shown in the image processing and time synchronization flowchart in Figure 8, after the current speaker is determined 800, the current speaker area is extracted from the single-camera image captured by camera 201 using wide-angle intelligent cropping 801. Perspective correction 802 adjusts distortion to center the target, and a stable viewfinder window is output 804. Cross-terminal image stitching 803 synchronizes multi-source video streams through communication unit 204 in multi-terminal scenarios, aligns the fusion trajectory according to a unified time base timestamp 806, uses Kalman filtering to generate a smooth trajectory 807, and uses switching command smoothing 809 to suppress jitter and send it to each terminal. This process supports triggering cloud assistance when local computing is insufficient, uploading features and returning parameters for local execution, ensuring degraded availability when end-to-end latency or confidence is below the threshold. The speaker 805, as the core target, is maintained in a centered display through the above interactions, reducing mis-slicing and jitter, and improving focus.
[0133] Furthermore, to enhance control security, the system introduces a gesture command verification mechanism. Before executing meeting control, facial recognition is required to confirm the authorized operator. The face detection confidence level must be no less than a first threshold, and the similarity to the template must be no less than a second threshold. Simultaneously, the second feature recognition must be either a gesture or a voiceprint: the gesture detection preset posture duration must be no less than a holding time threshold, and the voiceprint similarity under prompts or commands must be no less than a third threshold. The command is executed only after the two results are time-aligned within a preset time window. This multi-factor verification, through the collaboration of the processing unit 203, camera 201, and microphone array 202, suppresses false triggers and improves reliability.
[0134] In terms of speaker attribution based on voiceprint recognition, in a preferred embodiment, as shown in the main flowchart of the speaker attribution method based on voiceprint recognition in Figure 9, the system may include an audio acquisition and processing module, a voiceprint feature extraction module, a candidate database comparison and identity determination module, a temporary identity management module, a decision and write-back module, a central scheduling module, an output module, and a global post-processing module connected in sequence. The above modules are preferably implemented by the processing unit 203 of the terminal executing program instructions stored in the memory.
[0135] Specifically, in step 900, the audio acquisition and processing module receives the continuous audio stream acquired by the microphone array 202, performs speech activity detection (VAD) and speech segmentation, dividing the continuous audio stream into independent speech segments. Preferably, a sliding window method is used to filter silent segments and pure noise segments to avoid invalid computation. For each obtained speech segment, the voiceprint feature extraction module extracts its speaker's voiceprint feature vector. Preferably, a deep neural network based on the CAM++ model can be used. This model is pre-installed on the terminal, supports multiple languages (e.g., at least 10), and can map variable-length speech segments into fixed-dimensional (e.g., 512-dimensional) vectors to comprehensively represent speaker personality characteristics such as pitch, speech rate, and formant distribution.
[0136] In step 901, the candidate database comparison and identity determination module compares the voiceprint features of the current speech segment with the maintained speaker candidate database. The speaker candidate database includes voiceprint features of confirmed speakers and voiceprint features of unconfirmed temporary speakers. Preferably, the comparison uses cosine similarity and a dual-threshold decision strategy: when the similarity is greater than or equal to the first threshold (e.g., 0.45), the speech segment is determined to belong to an existing official identity, and the process proceeds to step 902. When the similarity is between the second threshold and the first threshold, and the time interval between the speech and the speaker's most recent speech is less than the silence threshold, the speech segment can be attributed to that existing identity based on the temporal context, and the process also proceeds to step 902. When the similarity is less than or equal to the second threshold, the current speech segment is considered not to belong to any existing identity, and the process proceeds to step 903.
[0137] In step 902, when the current speech segment is determined to have an existing speaker identity, the system directly assigns the identity tag to the speech segment and sends the speech recognition transcription result with the identity tag to the output module, which is then used in step 909 to generate the final meeting record with the identity tag.
[0138] In step 903, when an existing identity cannot be assigned, the candidate database comparison and identity determination module creates a new temporary speaker identity (temporary ID) for the current speech segment and registers the temporary identity's voiceprint features, first appearance time, and other metadata in the speaker candidate database. Subsequently, in step 904, the temporary identity management module continuously monitors the temporary identity, records its cumulative effective speech duration T_total, and monitors its speech quality indicators (such as signal-to-noise ratio, voiceprint similarity stability, etc.). If the confirmation conditions are not yet met, the duration and quality information continue to be accumulated in subsequent speech segments, as shown in the figure as "No / Continue Accumulation and Monitoring".
[0139] In step 905, the temporary identity management module determines whether the confirmation conditions are met based on the cumulative voice duration and voice quality. Preferably, a T1 / T2 dual threshold strategy is adopted: when the cumulative duration is greater than the first duration threshold T1 (e.g., 5s) and the similarity with the voiceprint of a previously confirmed speaker is greater than a preset threshold (e.g., 0.45), the temporary identity can be directly determined to be confirmed. When the cumulative duration is in the second range T2 (e.g., 1.5–5s), confirmation can be delayed until the next silence period to comprehensively consider the similarity and continuity of several recent voice segments. When the cumulative duration is less than the third duration threshold (e.g., 1.5s), the temporary identity management module can attempt to merge with other temporary identities (e.g., merge if the similarity is greater than or equal to 0.6). If merging is not possible, it is marked as an invalid or low-priority identity. To optimize early continuous speaking scenarios, in a specific example, when the accumulated duration reaches a first duration threshold and the comprehensive score (weighted by indicators such as voiceprint similarity, signal-to-noise ratio, and speaking continuity) is greater than or equal to 0.85, the temporary identity can be directly confirmed as a formal identity in advance. For example, in the first 0–5 seconds of the meeting, if a temporary identity has a voiceprint similarity of 0.88, a signal-to-noise ratio of 30dB, and a comprehensive score of 0.90, it can be upgraded to a formal speaker identity. If step 905 determines that the confirmation conditions have not yet been met, the process returns to step 904 to continue accumulating and monitoring.
[0140] Once the confirmation condition is met in step 905, the temporary identity management module triggers an identity confirmation event in step 906. This event declares that a temporary ID has been upgraded to a formal speaker identity and notifies subsequent modules to perform historical write-back and (in multi-terminal scenarios) identity synchronization. In step 907, the decision-making and write-back module officially upgrades the temporary ID to a formal ID and creates a corresponding formal entry in the speaker candidate database. Subsequently, in step 908, the decision-making and write-back module performs batch write-back processing on the identity identifiers of all historical voice segments corresponding to the temporary ID before the upgrade: it iterates through voice segments within a specified time interval, replaces the marked temporary IDs with the confirmed formal IDs, and can attach metadata tags such as "speaker-updated" to these voice segments. The updated results are permanently stored locally on the terminal to support post-meeting playback and version rollback. Through the above batch write-back, it is ensured that identity labeling in long sessions will not introduce inconsistencies due to temporary identities.
[0141] In step 909, the output module outputs the speech recognition transcription result with stable identity tags after the above process as the final meeting record. This record can be displayed in real time on the display unit 205, archived, or jointly analyzed with multimodal information. The global post-processing module preferably performs consistency checks and final cleanup on the aforementioned real-time output speaker attribution results at the end of audio stream processing or at the end of a predetermined time window. For example, it smooths out extremely short-term, sudden erroneous identity switching and merges adjacent segments belonging to the same speaker, thereby further improving the stability and accuracy of identity tagging in long-duration conversation scenarios.
[0142] Preferably, the entire process described above is executed locally online by the terminal processing unit 203, and the overall processing latency can be controlled within 200ms. Simultaneously, the terminal's Bluetooth communication unit 204 can synchronize the speaker candidate database and identity verification events in a peer-to-peer ad hoc network: when a terminal triggers an identity verification event in step 906, it can generate an event message containing fields such as event type, original temporary ID, new official ID, affected time interval, and candidate database version number, and broadcast it to other terminals via the Bluetooth network. Upon receiving the event, other terminals update the identity identifiers of their local candidate database and related historical voice segments accordingly, thereby maintaining consistency in speaker identity labeling in multi-terminal collaborative scenarios.
[0143] Based on the aforementioned pure audio architecture, this invention can be extended to an online speaker attribution method based on a combination of voiceprint and face recognition, applicable to audio-video synchronization scenarios. In this embodiment, the system further adds a video acquisition and face trajectory module, a lip-sync calculation and audio-video correlation matching module, and a face selection and binding management module: The video acquisition and face trajectory module acquires continuous video streams and outputs stable face trajectories. The lip-sync calculation and audio-video correlation matching module calculates the lip-sync activity of the face region and performs correlation analysis with voiceprint-based speech activity timing, sound source location, etc., to assist in speaker identification and error correction. The face selection and binding management module selects representative avatars from the corresponding face trajectories when preset quality and timing conditions are met, establishes a binding relationship with voiceprint identity, and maintains and corrects the results based on subsequent recognition. Through the collaboration of the above modules, this invention can achieve more robust and intuitive speaker attribution and identity display under multimodal conditions.
[0144] In terms of audio-visual joint binding and facial image generation, in a preferred embodiment, the present invention sets up an audio-visual joint binding and facial image generation module, which serves as a fusion link in a multimodal perception distributed conferencing system to achieve precise correspondence and dynamic maintenance between the speaker's identity on the audio side and visual elements. As shown in Figure 10, this module may include a video acquisition and facial trajectory submodule, a lip movement activity calculation and audio-visual correlation matching submodule, and a facial selection and binding management submodule. These submodules work together to form a complete process.
[0145] First, in step 1000, the system acquires video and audio frames synchronized with the audio stream. The terminal precisely aligns the acquired video frames and audio sampling points using a time synchronization protocol or a unified timestamp mechanism to ensure that there are no significant timing deviations in subsequent cross-modal processing.
[0146] On the video side, the video acquisition and face trajectory submodule performs face trajectory detection and quality assessment: In step 1001, real-time face detection and target tracking are performed on each video frame. Face detection boxes belonging to the same individual are concatenated through inter-frame correlation to form stable face trajectories. Simultaneously, a quality score is calculated for each detection result, such as a comprehensive quality score based on indicators like sharpness, attitude angles (yaw angle, pitch angle), lighting conditions, and occlusion, for subsequent optimization. Then, in step 1002, the system calculates the lip movement activity time series on the mouth region corresponding to each face trajectory. This can be achieved by tracking lip opening and closing, area changes, or shape features to extract a time series representing pronunciation dynamics. This time series characterizes the "speech activity level" of each face trajectory at different time periods.
[0147] On the audio side, in step 1003, the system performs speech activity / energy timing calculations in parallel based on the audio stream, estimating speech activity markers and energy intensity for each time slice to obtain a sequence of speech activity or energy changes over time. In step 1004, based on the aforementioned speaker attribution results based on voiceprint recognition, corresponding voiceprint identity information is provided for each time slice, including temporary identity and confirmed formal identity, thereby forming an "identified audio timeline".
[0148] After obtaining the aforementioned video and audio temporal information, the lip-sync activity calculation and audio-video correlation matching submodule performs a temporal correlation evaluation on the two in step 1005. Specifically, for a given identity-bearing speech time interval, within the allowable time lag range (considering end-to-end delay and buffer differences), the candidate face trajectory lip-sync activity sequence overlapping with its time is selected, and its correlation or similarity measure is calculated with the speech activity / energy sequence corresponding to that speech interval. When the obtained temporal correlation is greater than a preset correlation threshold, it is considered that the voiceprint identity and the face trajectory match each other within the current time window, and the process proceeds to step 1006. If the correlation is lower than the threshold, no binding relationship is established, only the audio identity result is retained, and matching can continue to be attempted in subsequent time windows.
[0149] When step 1005 determines that the relevance meets the threshold requirement, step 1006 establishes a three-way binding relationship between identity and facial trajectory: that is, the current voiceprint identity ID, its corresponding voice time interval, and the successfully matched facial trajectory ID are associated and stored to form a structured mapping entry, and the initial binding confidence level is recorded. This binding can distinguish between "temporary binding" and "formal binding": for short-term speech or scenarios with low confidence, it can be saved in the form of temporary binding first, and upgraded to formal binding after more temporal evidence is accumulated later.
[0150] In the face selection and binding management submodule, when a valid identity-trajectory binding relationship exists, the system performs candidate face screening within a preset capture time window: In step 1007, candidate face frames are selected from the target face trajectory within the time interval corresponding to the current speaking event. A comprehensive score is given to each candidate frame based on its clarity, pose angle, directness (degree of eye orientation towards the camera), and position in the image, filtering out frames with low quality or excessive pose deviation. Subsequently, in step 1008, representative avatars are further selected based on continuous frame stability and quality scores: frames with the highest comprehensive quality score and a number of consecutive frames exceeding a preset threshold (e.g., several consecutive frames) are selected, and these are generated and saved as the speaker's representative avatar. The generated representative avatar, along with the corresponding voiceprint identity and face trajectory ID, is written into a binding mapping table for subsequent display of the speaker's avatar on the interface or synchronous display across multiple terminals.
[0151] During long-term operation, to address video signal fluctuations, perspective changes, or initial binding errors, the face selection and binding management submodule also performs asynchronous correction and binding updates: In step 1009, the system periodically checks historical data through a background process. When higher-quality face images, more stable consistency evidence, or obvious errors in existing bindings are detected through statistical analysis, the binding relationship and representative avatar can be updated asynchronously without affecting real-time transcription and display on the front end. For example, the representative avatar in the original binding can be replaced, the binding confidence can be adjusted, or the identity-trajectory mapping can be reconstructed.
[0152] Furthermore, when the video signal is temporarily missing or its quality is below the available threshold, this module maintains the speaker attribution result on the audio side unchanged and prioritizes reusing the most recent valid face binding and representative avatar as the display prior. When the video signal recovers and again meets the relevance judgment condition in step 1005, the audio-video association matching is re-executed to correct or strengthen the binding relationship. Through this asynchronous correction and prior reuse mechanism, the present invention can maintain the continuity and consistency of interface display and identity labeling in weak network or complex scenarios.
[0153] In summary, the process shown in Figure 10 tightly integrates video acquisition with face trajectory, face quality assessment with lip movement activity time-series calculation, voice activity with voiceprint identity time-series, time-series correlation matching, and face selection and binding management through steps 1000–1009. This not only achieves consistent display of audio identity and visual elements in real-time conferencing scenarios, but also provides automatic error correction and adaptive update capabilities, significantly improving the robustness and user experience of the multimodal perception distributed conferencing system in practical applications.
[0154] In terms of system deployment, in a typical embodiment, two conference terminals can be arranged in conjunction with one conference display device in a small to medium-sized offline conference room. This arrangement maintains a simple deployment while achieving complete multimodal acquisition, fusion processing, and unified display effects. In this embodiment, the two terminals form a serverless peer-to-peer collaborative structure via Bluetooth self-organizing network and access the conference room wireless network to obtain a unified time reference when local area network conditions permit. Specifically, terminal A can be placed at the front of the conference room near the main presentation area, ensuring its camera provides good coverage of the front row and the main presentation area, and using a microphone array to achieve high-quality audio pickup from the front sound sources. Terminal B can be arranged on the side of the conference table, allowing it to supplement the acquisition of video and high signal-to-noise ratio audio information from lateral participants, covering the speeches from different seats during the discussion. The display device D is a large-size conference screen, connected to terminal A via HDMI or IP streaming. Terminal A is responsible for uniformly outputting the fused image, real-time subtitles, and identity display. The complementary camera perspectives of the two terminals and the different orientations of the microphone arrays enable the system to more accurately locate and assess the quality of multiple sound sources in the space, thus providing multi-dimensional information sources for optimal node selection.
[0155] After the meeting begins, Terminal A and Terminal B each perform local quality assessments on their respective acquired images and audio, calculating image quality scores and audio signal-to-noise ratio (SNR) metrics. The system periodically exchanges this quality data, along with node capability information such as current computing power, battery level, and network status, based on a peer-to-peer negotiation mechanism. This dynamically elects the optimal image and audio nodes for the current moment. For example, when the speaker is located at the front of the meeting room, Terminal A is typically selected as the optimal image node due to its larger target area and higher face detection confidence. When a participant speaks from the side, Terminal B, being closer to the sound source, has a higher speech SNR from its microphone array and is therefore selected as the optimal audio node to perform key audio processing tasks such as noise reduction, directional enhancement, and voiceprint recognition. After processing, the optimal audio node can broadcast the enhanced audio stream to another terminal, ensuring consistent audio output quality across the entire system. Furthermore, temporary identity verification events related to voiceprint matching are synchronized between the two terminals, ensuring that identity tags remain consistent across multiple terminals and preventing display confusion due to node switching.
[0156] On the video side, when either terminal detects that a speaker has started speaking, their facial trajectory and lip movements show a significant correlation with the timing of speech activity on the audio side. Based on this multimodal information, the system binds the voiceprint identity to the corresponding facial trajectory, and terminal A displays the speaker's avatar, name, and subtitles on the large screen. This binding relationship is updated synchronously between terminals, ensuring that regardless of how the optimal image or audio node dynamically switches between the two parties, the interface output by display device D maintains clear identity labeling and consistent avatars, thus guaranteeing a good user experience. In weak network or fluctuating network scenarios, the entire two-terminal architecture can still rely on Bluetooth self-organizing networking to maintain key collaborative logic. For example, when the LAN signal weakens, terminal A and terminal B can still exchange key information such as audio quality, identity events, and task instructions. If a terminal temporarily exits the task due to malfunction, obstruction, or low battery, the system will automatically select new optimal audio and image nodes based on the node capability information of the remaining terminals to maintain meeting continuity. In the event of a temporary loss of video signal, the system can prioritize maintaining the identity results on the audio side and reuse the most recently valid representative avatar. Once the video signal is restored, the matching and correction process is repeated, ensuring the continuity of the interface display. Through this deployment method, only two terminals are needed to achieve multi-view coverage, multi-array audio pickup, and multi-modal fusion capabilities for different areas of the conference room. The system presents highly consistent and robust meeting footage, subtitles, and identity display on the real-time display device. This typical embodiment demonstrates that the present invention not only achieves complete multi-modal distributed conferencing capabilities without an external server but also features lightweight deployment, low latency, and high fault tolerance, making it suitable for corporate meetings, small discussion rooms, teaching seminars, and daily hybrid meeting scenarios for small and medium-sized organizations.
[0157] In larger meeting scenarios or with more complex seating arrangements, this invention can further enhance coverage and multimodal perception capabilities by increasing the number of meeting terminals and placing them in different locations. In one optional embodiment, three or more meeting terminals can be deployed in long-distance meeting rooms, lecture halls, or auditoriums, positioned at the front, side aisles, and back rows, respectively. This allows each terminal's camera to capture panoramic and partial supplementary images of the meeting space from multiple heights and angles, while the microphone array covers different sound fields, such as the front row guest area, the general audience area, and the question area. All terminals still form a peer-to-peer self-organizing network via Bluetooth and, where possible, jointly access a local area network to obtain a unified time reference. Each terminal periodically reports its image quality score, audio quality score, node computing power, and current load information. Based on this, the system dynamically selects one or more optimal image nodes and optimal audio nodes to output the main screen and main audio stream. The remaining terminals serve as supplementary perspectives and candidate sound sources, participating in speaker localization, sound source separation, and multi-view stitching. When a speaker's position changes (e.g., from the podium to the audience Q&A area) or multiple points alternate speaking, the system can automatically reselect the optimal node and migrate tasks based on changes in image and audio quality at terminals in each area. This enables a smooth switch of the main perspective between front-end and rear-end terminals, and synchronizes voiceprint identity, audio-visual binding relationships, and representative avatars across all terminals, ensuring that the identity labeling and viewing angle displayed on the screen always remain consistent with the actual speaking position. Through this multi-terminal, multi-location arrangement, the invention can be flexibly extended to large and medium-sized lecture halls, training centers, or multifunctional conference spaces with mixed layouts. Without changing the core protocols and processing flow, increasing the number of terminals and optimizing physical locations can achieve multimodal perception and distributed collaborative processing over a larger spatial range, further enhancing the system's applicability and robustness in conference environments of different sizes and formats.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A distributed conferencing method for peer-to-peer ad hoc networks based on multimodal awareness, characterized in that, The method includes: multiple conference terminals establishing a serverless peer-to-peer network via Bluetooth ad hoc networking protocol, each conference terminal possessing local data processing and decision-making capabilities; each conference terminal in the peer-to-peer network collecting local image data and array microphone audio data, performing visual speaker detection and array microphone-based sound source localization locally, and preprocessing and quality assessment of the image and audio data to obtain image quality evaluation results and audio quality evaluation results; each conference terminal negotiating based on its local quality evaluation results, selecting the terminal with the best image quality as the optimal image node and the terminal with the best audio quality as the optimal audio node in the peer-to-peer network, and performing audio-video fitting and synchronized output between the optimal image node and the optimal audio node; performing speaker tracking and directional speech enhancement based on visual and auditory information, wherein the conference terminal determines the speaker's position in the image frame based on visual detection and estimates the speaker's spatial azimuth angle accordingly; using the spatial azimuth angle as guidance information, controlling the array microphone to perform directional sound source localization and beamforming to directionally enhance the speaker's speech signal.
2. The method according to claim 1, characterized in that, The method further includes an identity attribution step, which includes: combining lip movement activity information from the visual side and speech energy timing from the auditory side to determine the current speaker and establish a binding relationship between the current speaker's identity and its corresponding facial trajectory; comparing the voiceprint features extracted from the current speech segment with a dynamically maintained speaker candidate library, which includes voiceprint features of confirmed identities and voiceprint features of unconfirmed temporary identities; when the similarity with all identities in the candidate library is lower than a preset similarity threshold, creating a temporary identity for the current speech segment and outputting the temporary identity tag in real time, while continuously monitoring the accumulated speech data corresponding to the temporary identity; when the temporary identity meets preset confirmation conditions and is upgraded to a formal identity, triggering an identity confirmation event in the peer-to-peer network and collaboratively updating the historical speech segment identifier corresponding to the identity.
3. The method according to claim 1, characterized in that, The peer-to-peer negotiation and task allocation includes: each conference terminal independently calculating its own image quality score and audio quality score; the image quality score is calculated based on a preset conference mode, including at least a weighted evaluation of face detection confidence, the proportion of the target in the image, and the stability of the image composition; the audio quality score is calculated based on a preset noise reduction mode, including at least an evaluation of the signal-to-noise ratio after patterned noise reduction processing, consistency with the azimuth locked according to the spatial azimuth angle, and residual interference; based on the above scoring results, the terminal with the highest image quality score is selected as the optimal image node in the peer-to-peer network, and the terminal with the highest audio quality score is selected as the optimal audio node, so as to combine the data of the optimal image node and the data of the optimal audio node for fitting and output.
4. The method according to claim 1, characterized in that, The multimodal fusion and switching rules are as follows: when the speaker's mouth movement is detected on the visual side, and / or the face detection confidence reaches the first threshold, and / or the target's proportion in the image exceeds the second threshold and continues to exceed the first time threshold, a new candidate speaker is determined to have appeared; within the time window corresponding to the first time threshold, if the audio side's orientation estimation based on the array microphone, beamforming output energy, or signal-to-noise ratio after noise reduction is consistent with the spatial position of the candidate speaker and reaches the third threshold, then the switch to the candidate speaker's image is confirmed, and its corresponding terminal is designated as the speaker's associated node.
5. The method according to claim 4, characterized in that, The speech mouth movement detection on the visual side adopts a judgment method based on the correlation of lip shape activity, including: calculating the lip shape activity time sequence of the mouth region of the face trajectory on the visual side; measuring the correlation between the lip shape activity time sequence and the corresponding speaker's speech activity time sequence within an allowable time lag range; and establishing a three-way binding relationship between the current speaker's identity, face trajectory, and spatial location only when the correlation is greater than a first correlation threshold and the deviation between the visually estimated speaker position and the spatial orientation obtained by the auditory sound source localization is less than a preset angle threshold.
6. The method according to claim 2, characterized in that, The triggering conditions for identity verification include: Condition 1: When the cumulative speaking time of a temporary identity reaches a first duration threshold, the temporary identity is automatically upgraded to a formal identity and a global history write-back is triggered; Condition 2: When the cumulative speaking time is between the first duration threshold and the second duration threshold, the identity verification event is triggered during the silence period after the current round of speaking ends; Condition 3: When the cumulative speaking time is lower than the second duration threshold, the voice segment corresponding to the temporary identity is determined to be an invalid segment, or when the identity verification event is triggered, the voice segment is merged and assigned to an existing speaker with the highest similarity and a similarity greater than or equal to a preset merging similarity threshold; wherein, the first duration threshold is greater than the second duration threshold, the triggering logic is executed at the optimal audio node or the current computing node, and the result is synchronized to other terminals through the peer-to-peer network.
7. The method according to claim 2, characterized in that, The voiceprint feature comparison employs a dual discrimination logic, including: when the voiceprint similarity of the current speech segment is greater than or equal to a first similarity threshold, the speech segment is directly assigned to the corresponding existing identity; when the voiceprint similarity of the current speech segment is between a second similarity threshold and the first similarity threshold, and the current speech segment is temporally adjacent to the previous speech segment and the previous speech segment has already been assigned to a certain speaker identity, the current speech segment is determined to belong to the same speaker as the previous speech segment based on the contextual adjacency relationship; only when the voiceprint similarity of the current speech segment is lower than the second similarity threshold, and there is no contextual basis that satisfies the adjacency rule, is it determined that a new speaker has appeared and a corresponding temporary identity is created; wherein, the first similarity threshold is greater than the second similarity threshold.
8. The method according to claim 1, characterized in that, The conference terminal also uses a single camera with a wide field of view to intelligently crop and correct the perspective of the current speaker, so that the speaker is located in the central area of the output screen and forms a stable viewfinder; the audio and video data streams are aligned according to timestamps and time series fusion and smoothing are performed using Kalman filtering and hysteresis mechanisms to suppress frequent shaking and switching of the speaker's image.
9. The method according to claim 3, characterized in that, The audio-side noise reduction employs a patterned noise reduction strategy, including at least one of the following working modes: directional mode: locking the beam based on the target orientation obtained by visual detection and performing directional enhancement at the target orientation; voiceprint mode: under pre-registration or password triggering conditions, performing noise reduction and enhancement on the speech components based on the voiceprint characteristics of the target speaker; general mode: performing general noise reduction on the omnidirectionally acquired speech signal when no stable visual or voiceprint cues are obtained.
10. The method according to claim 9, characterized in that, The noise reduction task can be collaboratively executed by multiple conferencing terminals within the peer-to-peer network, including: dynamically allocating noise reduction tasks to each terminal based on a weighted score composed of frame energy, signal-to-noise ratio, terminal processing capability, and task priority of the speech frame; the task negotiation adopts a point-to-point communication protocol and carries task description, device status, priority, and threshold parameter fields; the terminals synchronize speech frames using a unified time base or timestamp alignment, and fuse the temporal energy trajectories from multiple terminals to constrain the consistency of the noise reduction output; the noise reduction model uses a joint loss function including a temporal signal-to-noise ratio term, a speech component term, and a noise component term during training, and assigns higher weights to low-frequency subbands and subbands with prominent non-steady-state noise; the end-to-end processing latency of the noise reduction front-end does not exceed a preset latency threshold and is deployed between the audio input and speech recognition modules.
11. The method according to claim 10, characterized in that, When the audio-side noise reduction is in general mode, real-time noise reduction processing is performed on the acquired speech signal. The noise reduction processing method includes: performing a short-time Fourier transform on the noisy speech to obtain complex spectral features; inputting the complex spectral features into a neural network containing encoding, enhancement, and decoding and with skip connections to generate an initial spectral mask; optimizing the initial spectral mask based on causal time-series weighted filtering that only uses the time-frequency information of the current frame and its preceding frames, without using subsequent frame information; correcting the optimized spectral mask according to the minimum mean square error criterion, calculating the prior signal-to-noise ratio and the posterior signal-to-noise ratio, and obtaining the final complex spectral mask according to a preset frequency point gain formula, so that the complex spectral mask weights the real and imaginary parts of the spectrum within the amplitude limit range; multiplying the final complex spectral mask with the noisy spectrum and performing an inverse short-time Fourier transform to output the denoised speech.
12. The method according to claim 11, characterized in that, The noise reduction method further includes the following training data construction process: acquiring room impulse response, clean speech, and noise data; adjusting the reverberation duration of the room impulse response to ensure the target reverberation time is between 0.10 seconds and 0.60 seconds while maintaining its decay pattern; convolving the clean speech with the adjusted room impulse response to obtain reverberant speech; selecting one or more noise sources with a probability less than 0.7; setting the signal-to-noise ratio of the training samples in the range of -10 dB to +20 dB; performing enhancement processing and amplitude consistency normalization on the reverberant speech and / or noise in the spectral domain; normalizing at least one signal level to a predetermined range; and mixing the signals according to the signal-to-noise ratio strategy to form the training input, with the clean speech as the training target.
13. The method according to claim 12, characterized in that, The noise reduction method uses a U-shaped neural network that includes encoding, enhancement, and decoding, with skip connections. The encoding and decoding parts each consist of multiple levels of convolutional and deconvolutional layers. Except for the first level of the encoding layer and the last level of the decoding layer, the kernel size of each convolutional and deconvolutional layer is 3×1, while the kernel size of the first level of the encoding layer and the last level of the decoding layer is 5×2. The enhancement part groups the layers along the channel dimension, and each group is configured with a recurrent unit for independent computation. The outputs of each group are instantaneously normalized and then fused. Skip connections use 1×1 convolutions for channel matching. The features are fused with the decoding features by feature addition, and the network outputs a complex spectral mask with the mask amplitude limited to the range of 0 to 1. The causal temporal weighted filtering applies a weighting of the number of back-look-back frames L and the neighborhood width K of the adjacent frequency points to each frequency point. The weights are non-negative and normalized frequency-by-frequency point, generated by gated convolution or attention mapping. The minimum mean square error post-processing adopts any one of Wiener gain, logarithmic amplitude minimum mean square error, or spectral amplitude minimum mean square error. The prior signal-to-noise ratio is obtained by exponentially smoothing the previous time-series estimate and the current posterior signal-to-noise ratio using a decision-oriented method. The noise power spectrum is estimated using online noise tracking or speech presence probability methods.
14. The method according to claim 2, characterized in that, The identity attribution step also includes face capture and avatar generation based on audio-video joint binding, used to generate representative avatars corresponding to the speaker's identity. Specifically, for each speaking event detected by the audio side, within a preset capture window after the start of the speech, i.e., within the time interval from the start time of the speech plus a first time offset to the start time of the speech plus a second time offset, face detection is performed on the video stream; target face trajectories are filtered based on the voiceprint identity to which the current speech segment belongs and the degree of matching with the lip movement activity of that speech segment, excluding faces of non-current speakers who have just finished speaking; among the retained faces, candidate avatars are determined based on image clarity, frontal angle, and detection confidence, and the face image with the highest comprehensive quality index and the number of consecutive frames reaching a preset frame threshold is selected as the representative avatar bound to the current speaker's identity, and is output synchronously along with the identity identifier.
15. The method according to claim 14, characterized in that, The facial image corresponding to the representative avatar is bound to the current speaker's identity and stored as a mapping relationship between the speaker identifier and the facial trajectory identifier, while the binding confidence level is recorded. When a higher quality facial image of the speaker is captured in a subsequent meeting, or when an error is found in the initial binding based on long-term clustering analysis, the mapping relationship and its corresponding facial image and binding confidence level are updated asynchronously in the background. The updated binding information is used to refresh the interface display, but does not interrupt the current real-time transcription output and speaker attribution results.
16. The method according to claim 14, characterized in that, When the video signal is missing or the current video quality is below a preset threshold, the speaker attribution result on the audio side continues to be output, and the face image corresponding to the most recent valid audio-video joint binding relationship is reused as the display prior. When the video signal is restored and the detected face meets the preset association conditions in terms of temporal correlation and similarity with the current speaker, the audio-video correlation matching is re-executed to correct the binding relationship between the speaker identity and the face image.
17. A peer-to-peer self-organizing network distributed conferencing terminal device based on multimodal sensing, characterized in that, A method for implementing any one of claims 1 to 16 includes: an array microphone assembly for acquiring local conference audio signals and supporting time difference of arrival measurement and sound source location estimation; a camera assembly for acquiring image data related to speaker detection; a Bluetooth communication unit for establishing a serverless peer-to-peer network with other conference terminals according to the Bluetooth self-organizing network protocol and conducting point-to-point communication to exchange local image quality evaluation results and audio quality evaluation results and negotiate the optimal node; a synchronization unit for providing a unified time reference and / or aligning with external timestamps to ensure the temporal comparability of audio and video data acquired by each terminal; and a processor and a memory, wherein the memory stores program instructions that run on the processor, and the processor executes... When executing the program instructions, the system is configured to: perform local preprocessing and quality assessment on the data collected by the camera component and the array microphone component to generate image quality evaluation results and audio quality evaluation results; conduct peer-to-peer negotiation with other conference terminals through the Bluetooth communication unit, select the terminal with the best image quality as the optimal image node and the terminal with the best audio quality as the optimal audio node in the peer network, and complete audio and video fitting and synchronous output between the optimal image node and the optimal audio node; determine the position of the current speaker in the image frame based on visual detection and estimate its spatial azimuth angle, use the spatial azimuth angle as guidance information to control the array microphone to perform directional sound source localization and beamforming, and perform directional enhancement of the speaker's voice signal.
18. The device according to claim 17, characterized in that, When executing the program instructions, the processor is configured to implement the following functional modules: an audio acquisition and processing module, used to receive a continuous audio stream and perform speech activity detection and speech segmentation; The voiceprint feature extraction module is used to extract speaker voiceprint feature vectors from the speech segment; the candidate library comparison and identity determination module is used to compare the voiceprint features with the maintained speaker candidate library, which includes confirmed speaker voiceprint features and unconfirmed temporary speaker voiceprint features, and determines whether the current speech segment belongs to an existing speaker identity or a new temporary speaker identity is created based on similarity and temporal context; the temporary identity management module is used to manage the temporary speaker identity, including recording its cumulative speech duration, monitoring its status, and triggering an identity confirmation event when preset confirmation conditions are met; The decision and write-back module is used to upgrade the corresponding temporary identity to the formal speaker identity when the identity confirmation event is triggered, and to batch update the identity identifier of the historical voice segment corresponding to the temporary identity to the confirmed identity and store it in a fixed manner. The central scheduling module is used to drive the online processing pipeline in a sliding window manner, serially traversing each processing stage and triggering corresponding processing and identity verification events based on conditions. The output module is used to output the speech recognition transcription results with speaker identification information. The global post-processing module is used to perform consistency checks and final cleanup on the real-time speaker attribution results at the end of audio stream processing or at the end of a predetermined time window, so as to improve the consistency and accuracy of the results.
19. The device according to claim 18, characterized in that, When executing the program instructions, the processor is configured to further implement the following functional modules: a video acquisition and face trajectory module, used to acquire a video stream synchronized with the audio stream, perform face detection on the video frames, and output the face trajectory and its quality score; The lip-sync activity calculation and audio-visual correlation matching module is used to calculate the lip-sync activity time series of the mouth region of the face trajectory, and to measure the correlation and threshold determination of the lip-sync activity time series with the speech activity or speech energy time series corresponding to each speaker identity within an allowable time lag range, so as to establish a binding relationship between speaker identity and face trajectory; the face selection and binding management module is used to screen candidate faces within a preset time window, complete the binding of speaker identity and face capture, record and update the binding confidence, and when a higher quality face image is subsequently obtained or an error is found in the initial binding through long-term statistical analysis, the binding relationship and its face capture are asynchronously corrected and updated. When the video signal is missing or the quality is substandard, the speaker attribution result on the audio side is maintained and the most recent valid face binding is reused as the display prior. When the video signal is restored and the preset correlation conditions are met, the audio-visual correlation matching is re-executed to correct the binding relationship.
20. A peer-to-peer self-organizing network distributed conferencing system based on multimodal sensing, characterized in that, The method for performing any one of claims 1 to 16 includes multiple conference terminal devices and at least one conference display device, wherein: the multiple conference terminal devices are deployed in the same offline physical conference room, each conference terminal device establishes a serverless local peer-to-peer network through its own Bluetooth communication unit according to the Bluetooth self-organizing network protocol, for collecting audio and / or video data from its own end, and performing control processing corresponding to the method locally; the conference display device is located in the physical conference room and connected to at least one conference terminal device in the local peer-to-peer network, for receiving conference processing results generated by the conference terminal devices.
Citation Information
Patent Citations
Techniques to automatically identify participants for a multimedia conference event
CN101952852A
Ultrasonic-assisted microphone array speech enhancement device
CN102800325A
Video conferencing devices and methods
CN104349112B
Voice control method and system based on distributed multiple microphones and Bluetooth Mesh
CN110310637A
Methods for capturing and playing conference audio
CN111586523B