A conference picture output method, system, storage medium and program product

CN122534191APending Publication Date: 2026-08-07YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
Filing Date
2026-07-03
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]提供一种会议画面输出方法、系统、存储介质及程序产品,旨在解决目前会议画面稳定性不佳的问题

Benefits of technology

[0018]上述技术方案中的一个技术方案具有如下优点或有益效果:本技术方案中,通过对声源进行定位以确定追踪区域和对应的定位状态,该定位状态并非仅区分"有无发言人"的二元结果,而是划分为"定位到具体发言人""未定位到具体发言人但定位到发言区域""具体发言人和发言区域均未定位到"三种粒度,由此形成了从特写到区域到预设位的分级输出机制:“当定位精度高时输出发言人特写画面,定位精度中等时输出发言区域画面,定位精度不足时输出预设位画面”,从而在任何精度下均能输出合理的画面内容,避免了因定位精度波动导致的画面空白或错误框选。同时,获取与声源对应的发言数据,以该数据作为是否触发画面切换的量化判断依据:“只有当发言数据(如持续累计发言时长、发言频率等)满足条件时才进行画面切换”,从而将短暂的插话、附和性应答、语句间的自然停顿等非实质性发言与真正的持续发言区分开,避免了因短时声源触发画面频繁跳转。此外,获取追踪区域的画面显示模式,根据该模式确定目标画面的构成形式(如单画面模式或组合画面模式),使得画面输出策略与会议类型相匹配。综合上述三个维度:定位状态确保画面输出粒度与定位精度相适应、发言数据确保画面切换仅针对持续性发言而非瞬时声源、画面显示模式确保画面输出形式与会议场景相匹配,在外界因素干扰导致声源定位不精准、声源信号短暂波动等情形下仍能输出稳定合理的会议画面,显著提升了会议画面输出的稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122534191A_ABST
    Figure CN122534191A_ABST
Patent Text Reader

Abstract

The application discloses a conference picture output method and system, a storage medium and a program product, and belongs to the technical field of electronics and cameras. In the case that a sound source is detected, the sound source is positioned, a tracking area where the sound source is located and a positioning state corresponding to the sound source are determined, wherein the positioning state comprises one of the following: positioning to a specific speaker, not positioning to the specific speaker but positioning to a speaking area, and neither the specific speaker nor the speaking area being positioned; speaking data corresponding to the sound source is acquired; a picture display mode of the tracking area is acquired; and a target conference picture is output based on the positioning state, the speaking data and the picture display mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of electronics and photography, specifically to a method, system, storage medium, and program product for outputting conference images. Background Technology

[0002] Currently, in conference systems, a common approach to speaker tracking is to use sound source localization to drive screen switching. For example, a microphone array detects the direction of the sound source in real time to obtain the speaker's spatial coordinates. Once the speaker is located, the system immediately drives the pan-tilt camera to switch the view from a panoramic view to a close-up of the speaker; when the speaker stops speaking or the sound source disappears, the view immediately switches back to a panoramic view. In this approach, the screen state completely follows the changes in sound source detection results; detection is an immediate response, and the screen reverts to its previous state when the sound source disappears.

[0003] However, in actual meeting scenarios, under the aforementioned "sound-on-demand" mechanism, it is easy for a problem to occur where even if a sound source is detected, the specific location of the sound source may not be accurately identified due to external interference. If the corresponding close-up image is still output, the speaker may be selected incorrectly due to the inaccurate identification of the sound source, making it difficult to guarantee the stability of the meeting image output.

[0004] Therefore, improving the stability of conference video output has become an urgent problem to be solved. Summary of the Invention

[0005] This invention provides a method, system, storage medium, and program product for outputting conference video, aiming to solve the problem of poor stability of current conference video output.

[0006] Firstly, a method for outputting conference video is provided, including the following steps: When a sound source is detected, the sound source is located to determine the tracking area where the sound source is located and the corresponding location status of the sound source. The location status includes one of the following: the specific speaker is located, the specific speaker is not located but the speaking area is located, and neither the specific speaker nor the speaking area is located. Obtain the speech data corresponding to the sound source; Obtain the screen display mode of the tracking area; Based on the positioning status, the speaking data, and the screen display mode, the target conference screen is output.

[0007] Secondly, a conference video output system is provided, including a processing device, a microphone, and a camera; The processing device is used to locate the sound source through the microphone when a sound source is detected, and to determine the tracking area where the sound source is located and the corresponding positioning state of the sound source. The positioning state includes one of the following: the specific speaker is located, the specific speaker is not located but the speaking area is located, and neither the specific speaker nor the speaking area is located. The processing device is used to acquire speech data corresponding to the sound source; The processing device is used to acquire the screen display mode configured in the tracking area; The processing device is used to control the camera to output the target conference image based on the positioning status, the speaking data, and the screen display mode.

[0008] Thirdly, the present application provides a conference screen output device, including: a processor configured to perform the method of any of the above-mentioned aspects.

[0009] Optionally, the device may further include the memory and / or the communication interface.

[0010] The communication interface is coupled to the processor and is used for inputting and / or outputting information.

[0011] The memory is used to store computer programs. The processor is configured to execute any of the methods described above, and can be implemented as: executing the computer program stored in the memory to execute any of the methods described above. Alternatively, the processor can be a hardware-implemented circuit, such as an artificial intelligence (AI) processor, to improve operating speed. This application does not limit the specific implementation of the processor.

[0012] Optionally, the device can be a complete machine or a module within a machine, such as a chip.

[0013] Fourthly, the present application provides a computer-readable storage medium including computer instructions that, when executed on a device, cause the device to perform any of the possible designs described above.

[0014] Fifthly, the present application provides a computer program product that, when run on a device, causes the device to execute the method in any possible design of any of the above aspects.

[0015] Sixthly, this application provides a circuit system including a processing circuit configured to perform the methods in any possible design of any of the above aspects. The processing circuit can be implemented as a corresponding circuit component, such as one or more processors. Alternatively, it can be implemented as a processor and a memory. Yet another example is a processor and a transceiver.

[0016] In a seventh aspect, this application provides a chip system including at least one processor and at least one interface circuit, wherein the at least one interface circuit is used to perform transceiver functions and send instructions to the at least one processor, and when the at least one processor executes instructions, the at least one processor performs the method described in the first aspect and any of the designs therein.

[0017] Eighthly, this application provides a conference screen output device, including a functional module, unit, or means for performing the methods in any possible design of any aspect of this application described above. The module may be implemented by software or hardware, or by a combination of software and hardware. The inclusion of a processing unit and a communication unit is not limited.

[0018] One of the above technical solutions has the following advantages or beneficial effects: In this technical solution, the sound source is located to determine the tracking area and the corresponding positioning state. This positioning state is not simply a binary result of "whether there is a speaker" but is divided into three granularities: "located to a specific speaker", "not located to a specific speaker but located to a speaking area", and "neither the specific speaker nor the speaking area is located". This forms a hierarchical output mechanism from close-up to area to preset position: "When the positioning accuracy is high, a close-up of the speaker is output; when the positioning accuracy is medium, a speaking area is output; and when the positioning accuracy is insufficient, a preset position is output". This ensures that reasonable image content can be output at any accuracy, avoiding blank screens or incorrect selections caused by fluctuations in positioning accuracy. At the same time, speech data corresponding to the sound source is acquired and used as a quantitative basis for determining whether to trigger a screen switch: "Screen switch is only performed when the speech data (such as the cumulative speaking time, speaking frequency, etc.) meets the conditions". This distinguishes between brief interruptions, echoing responses, natural pauses between sentences, and other non-substantive speech from true continuous speech, avoiding frequent screen jumps triggered by short-term sound sources. Furthermore, the system acquires the display mode of the tracking area and determines the composition of the target screen (such as single-screen mode or combined-screen mode) based on this mode, ensuring that the screen output strategy matches the meeting type. Combining these three dimensions—positioning status ensuring that the granularity of the screen output matches the positioning accuracy, speaking data ensuring that screen switching only applies to continuous speaking rather than instantaneous sound sources, and screen display mode ensuring that the screen output format matches the meeting scenario—it can still output a stable and reasonable meeting screen even when external factors cause inaccurate sound source positioning or temporary fluctuations in the sound source signal, significantly improving the stability of the meeting screen output. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic diagram of the structure of the apparatus provided in an exemplary embodiment of this disclosure; Figure 2 This is a flowchart illustrating a conference screen output method provided by an exemplary embodiment of this disclosure; Figure 3 , Figure 4 This is a schematic diagram of a scenario for a conference screen output method provided by an exemplary embodiment of this disclosure; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of this disclosure; Figure 6 This is yet another structural schematic diagram of the electronic device provided in the exemplary embodiments of this disclosure; Figure 7 This is a schematic diagram of the structure of a chip system provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0023] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.

[0024] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not exclude the applicability to or configuration to devices performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated.

[0025] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to implement and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be implemented without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0026] Currently, in conference room facial recognition systems, a simple responsive solution of sound source localization → direct image switching is commonly used. Specifically, a microphone array detects the direction of the sound source in real time to obtain the speaker's spatial coordinates. Once the speaker is located, the pan-tilt camera is immediately activated, switching the view from a panoramic shot to a close-up of the speaker. When the speaker stops speaking or the sound source disappears, the view immediately switches back to a panoramic shot. When a new speaker is detected, the view immediately switches to a close-up of the new speaker. In this solution, detection is immediate response, sound source disappearance results in reversal, and the image state directly follows the sound source detection result.

[0027] In response, the inventors of this application, after in-depth research, discovered that those skilled in the art have long been limited by the conventional technical approach of "detection equals response, sound source disappearance equals rewind," consistently viewing the relationship between sound source detection results and image output as a direct mapping: "detecting a sound source triggers a close-up, the sound source disappears triggers a full view," failing to address the unstable positioning accuracy issue this direct mapping faces in real-world meeting scenarios. Specifically, in meeting room portrait framing systems, a responsive solution based on sound source localization → direct image switching is typically adopted. This solution is logically simple and direct: once the microphone array detects the direction of the sound source, it immediately drives the gimbal camera to turn in that direction and outputs a close-up image. This design approach is intuitive and easy to implement, and therefore has long been considered a standard paradigm by those skilled in the art, rarely questioned.

[0028] However, through observation and data accumulation of numerous real-world meeting scenarios, the inventors discovered a fundamental problem that had been long overlooked in this seemingly reasonable solution: sound source localization results are unstable in real-world meeting environments, yet existing solutions directly use them as stable and reliable information. This manifests itself in the following scenarios: Scenario 1: Insufficient sound source localization accuracy leads to incorrect frame selection even when outputting close-up shots. When a speaker turns their head, multiple speakers overlap, causing sound source interference, or when echoes and reverberation exist in the environment, the microphone array can still detect the presence of the sound source, but the accuracy of its precise location will significantly decrease. For example, the sound source localization system may identify the general area where the sound source is located as the "left side area," but it cannot pinpoint the specific speaker within that area. In this situation, if the system still outputs a close-up shot according to the logic of "detection equals close-up," it will face a very high risk of frame selection errors—the person framed in the shot may not be the actual speaker, or the wrong body part may be framed. Since the sound source localization system "does provide some directional information" when it detects a sound source (albeit with insufficient accuracy), the system is not completely faulty, so this problem is easily attributed to "insufficient localization accuracy" rather than a defect in the solution itself. However, the inventors realized that the root of the problem is not the localization accuracy itself, but that the solution equates "detecting a sound source" with "having sufficient conditions for outputting a close-up shot," lacking a graded discrimination and differentiated output strategy for localization accuracy.

[0029] Scenario 2: Brief Loss of Sound Source Localization Causes Image Shaking During Rewind. During a speech, when the speaker turns their head, gets up and moves around, or is briefly obstructed by others, sound source localization may be momentarily lost (typically lasting hundreds of milliseconds to several seconds). Traditional solutions immediately rewind the image from a close-up to a wide shot once the sound source is no longer detected; when the sound source is detected again, they immediately switch back to a close-up. This process seems reasonable at the level of a single event, but it occurs repeatedly in continuous speaking scenarios, causing frequent image shaking. Because each loss of localization is transient, it is not easily reproduced in live demonstrations or short-term tests, and this problem has long been overlooked.

[0030] Scenario 3: Short-duration sound source events trigger invalid switching. Brief interruptions, echoing responses, coughs, and other non-speaking sound source events are common in meetings. Each short-duration sound source triggers a complete switch from wide shot to close-up to wide shot again. However, these sound source events are unrelated to the close-up footage that actually needs to be output, and the invalid switching severely disrupts the viewing experience for remote participants. Since each switch itself does not "error" (there is indeed a sound source), this problem has long been attributed to "normal system behavior" rather than "design flaws."

[0031] Although the specific manifestations of the above scenarios differ, they all point to the same fundamental cognitive limitation: those skilled in the art have long equated "the instantaneous detection result of the sound source" with "a sufficient condition for screen switching," while ignoring the instability of positioning accuracy, the non-substantial nature of sound source events, and the matching relationship between screen output strategy and positioning quality. In other words, the core assumption of existing solutions—"as long as a sound source is detected, a close-up image should be output"—does not hold true in real meeting scenarios. However, this systemic flaw has long gone unrecognized by those skilled in the art, even when this assumption aligns with intuition and the solution "works" (under ideal conditions).

[0032] Based on this, the inventors broke away from the conventional thinking framework of "sound source detection → direct mapping → screen output" and proposed a screen output method based on a comprehensive decision-making process of three dimensions: positioning status, speaking data, and screen display mode, in order to improve the stability of conference screen output.

[0033] The following description, in conjunction with specific embodiments, illustrates this method of outputting the image. This image output method can be executed by a processing device. The processing device can be a standalone host, a processing chip within a standalone host, an integrated device, or a processing chip within an integrated device. The integrated device can integrate a processor and a camera. For example, the integrated device can be a conference terminal device.

[0034] Figure 1 An example of the structure of the processing device is shown. For example... Figure 1 The processing device may include: a speech data acquisition module, a positioning data acquisition module, a timer management module, and a screen switching decision engine.

[0035] The speech data acquisition module can obtain one or more of the following speech data for each speaker through Voice Activity Detection (VAD) and sound source separation: speech status, cumulative speech duration (also known as cumulative speaking time or continuous cumulative speaking time), speech interval duration, number of speakers and the speaking percentage of each speaker, and speaking frequency within a preset time window. The speech status indicates whether a speaker is currently speaking.

[0036] Location data acquisition module: This module acquires the speaker's location data through microphone array positioning and / or camera visual positioning. The location data can be categorized into three levels of accuracy: Level 1 Precision: It can accurately locate the specific speaker and output the precise speaker's position.

[0037] Second-level accuracy: can roughly locate the speaking area. For example, if multiple people speak at the same time, causing sound source interference or the speaker's position is at the edge of the array coverage, it may not be possible to pinpoint the specific individual's location, but it can determine the high-confidence speaking area where the sound source is located.

[0038] Level 3 accuracy: Unable to locate. For example, the sound source signal is too weak, the ambient noise is too loud, or the sound source is outside the effective range of the microphone array. There is noise near the camera's built-in microphone, or the person is not facing the camera directly. The direction of arrival (DOA) measured by the camera's built-in microphone cannot match the angle of the person in the image, making it impossible to locate.

[0039] Timer Management Module: Maintains an independent speaking timer (cumulative speaking duration) and an interval timer (interval between the last speaking) for each detected speaker or high-confidence speaking area, and maintains a region-level cumulative speaking timer for each tracking area.

[0040] Screen switching decision engine: Based on speech data, positioning data, thresholds (T0, T1), and configured screen display modes (such as focus / relay / ensemble), it comprehensively determines the currently output screen and controls the execution of screen switching. For example, the output screen can be a panorama, an automatically framed view (Auto Framing, AF) at a preset position, a close-up of a person, a close-up of an area, or a combination of panorama and multi-view.

[0041] Among them, the preset position is a pre-set combination of camera shooting position parameters (such as PTZ angle, zoom magnification, etc.), which corresponds to the optimal shooting parameters for a certain tracking area.

[0042] Figure 2 A flowchart example of the conference screen output method of this disclosure is shown, which can be executed by a processing device. For example... Figure 2 The method may include: S101. When a sound source is detected, the sound source is located to determine the tracking area where the sound source is located and the corresponding location status of the sound source.

[0043] The positioning status includes one of the following: the specific speaker has been located, the specific speaker has not been located but the speaking area has been located, and neither the specific speaker nor the speaking area has been located. These three positioning statuses correspond to the three levels of positioning accuracy mentioned above.

[0044] As one possible implementation, the processing device can use the multi-channel audio from the microphone array, combined with the VAD algorithm, to determine whether there is a valid human voice present, and thus calculate the speaking duration. For example, the VAD algorithm outputs sound source activity information every 200ms.

[0045] As one possible implementation, the processing device can locate the sound source based on the positioning data from the microphone. For example, the positioning result can be determined by combining the positioning data from the ceiling microphone and the built-in microphone of the camera.

[0046] For example, the current space (such as a conference room) can be pre-divided into multiple tracking zones, each with a camera attached. The processing device can first determine the tracking zone where the sound source is located based on the sound source localization result from the ceiling microphone, and then determine the sound source localization result based on the camera in that tracking zone.

[0047] For example, the camera's built-in microphone outputs a sound source positioning angle, which is referenced to the camera's panoramic lens. The processing unit then determines whether this positioning angle, within a preset range (e.g., ±3°), can match the image of a person captured by the camera's panoramic lens. If a match is found, this positioning angle is used as the speaker's position.

[0048] As one possible implementation, the processing device can separate the audio of different speakers into independent audio streams based on the sound source localization results, and maintain the following attributes for each speaker: •speaker_id: A unique identifier for the speaker; •total_speech_duration: The cumulative speech duration in the current speaking round; • current_speech_start_time: The start time of the current speech segment; • last_speech_end_time: The end time of the last speech segment; •speech_ratio: The percentage of the total speaking time for this speaker; •tracking_zone_id: The tracking zone where the speaker is currently located.

[0049] S102. Obtain the speech data corresponding to the sound source.

[0050] Among them, speech data can reflect the activity level of the sound source. The types of speech data can be referred to above, and will not be repeated here.

[0051] For example, the processing device can acquire the speech data through the speech data acquisition module.

[0052] S103. Obtain the screen display mode of the tracking area.

[0053] The display mode represents the preset output strategy, such as single-screen mode, combined-screen mode, presentation mode, discussion mode, single-speaker mode, or multi-speaker mode. It can be used to determine the composition of the target meeting screen (such as single-screen or combined-screen).

[0054] For example, the screen display mode is a user's or system's preset display preference. For instance, "speech mode" tends to only show close-ups of the speaker, while "discussion mode" tends to show a discussion area screen containing multiple people.

[0055] S104. Based on the positioning status, speaking data, and screen display mode, output the target conference screen.

[0056] For example, in a conference room, the system detects a speech in area A. The system can pinpoint the sound source to the front left area (tracking area) using sound source localization. Then, through the microphone array and camera linkage, the system determines that the sound source originates from a specific speaker, Zhang San (localization status: located specific speaker). Furthermore, the system begins accumulating Zhang San's speaking time (speaking data). The system can also determine that the current meeting's display mode is presentation mode. Based on "located specific speaker," "Zhang San's speaking time has reached 5 seconds," and "presentation mode," the system can output a close-up image of Zhang San as the target meeting screen.

[0057] Compared to the traditional "send as soon as there is sound" approach, the solution provided in this embodiment makes comprehensive decisions based on three dimensions: location status, speech data, and screen display mode. This enables delayed judgment and hierarchical decision-making capabilities for screen switching. Specifically: Differentiated screen output granularity is achieved through three levels of location status, avoiding erroneous screen output when precise location is unavailable; speech data (such as cumulative speech duration and frequency) serves as a quantitative basis for switching, effectively filtering out invalid switching caused by brief interruptions and pauses between sentences; and preset screen display modes ensure that screen output conforms to the expected meeting type. This comprehensive decision-making across three dimensions effectively suppresses invalid switching caused by short speeches, pauses, and temporary loss of location, significantly improving the stability and intelligence of meeting screen output.

[0058] In some embodiments, the target conference screen includes a single screen and a combined screen; based on the positioning status, speaking data, and screen display mode, the target conference screen is output, including: S104a. Based on the speech data, determine whether the conditions for screen switching are met.

[0059] Screen switching conditions are the criteria used to trigger screen switching. For example, screen switching is only allowed when the speaking data (such as cumulative speaking time) reaches a certain threshold.

[0060] S104b: If the screen switching conditions are met, determine the target screen display mode based on the currently acquired screen display mode, speech data, and / or positioning status.

[0061] The target screen display mode includes single-screen mode and combined-screen mode. The target screen display mode is the final display mode dynamically determined by the system based on the current state after the screen switching conditions are met.

[0062] S104c: Based on the positioning status and target screen display mode, output the corresponding single screen or combination screen.

[0063] For example, the current screen display mode is single-screen mode, showing a close-up of speaker Zhang San. After Zhang San finishes speaking, the system detects that Li Si begins speaking and starts accumulating Li Si's speaking time. When Li Si's accumulated speaking time reaches a first time threshold (e.g., 3 seconds), the system determines that the screen switching condition is met. In this case, the system can determine the target screen display mode based on the current screen display mode (single-screen mode) and the speaking data (Li Si's speaking time threshold). For example, if the user wants to display both Zhang San and Li Si simultaneously, the target screen display mode is determined to be a combined screen mode. If only Li Si is desired to be displayed, the target screen display mode remains single-screen mode. Assuming the system determines the target screen display mode to be single-screen mode according to a preset strategy, it can output a close-up of Li Si based on the location status (located to Li Si), thereby achieving a single-screen switch from Zhang San to Li Si.

[0064] The solution provided in this embodiment clearly defines the conditions for screen switching. Screen switching is triggered only when these conditions are met, rather than responding immediately to any change in sound source. This reduces the frequency of screen switching.

[0065] In some embodiments, the target screen display mode is determined based on the currently acquired screen display mode, speech data, and / or positioning status, including: If the current screen display mode is single-screen mode, and it is determined based on the speech data and / or positioning status that at least one additional screen corresponding to another sound source needs to be displayed, then the target screen display mode is determined to be combined screen mode. This case is referred to as Case 1.

[0066] For example, in one embodiment, a single-screen to composite screen conversion is performed based on speaking duration: When the currently displayed single screen is a close-up of the first speaker, the processing device detects that the second speaker has started speaking. The processing device begins to accumulate the continuous speaking duration of the second speaker. When the continuous speaking duration of the second speaker reaches a first duration threshold (e.g., 3 seconds), the processing device determines that a screen corresponding to the second speaker needs to be added. At this time, if the first speaker is still speaking, the processing device determines that the target screen display mode is a composite screen mode and outputs a composite screen containing close-ups of the first and second speakers. If multiple new speakers (such as a third or fourth speaker) are detected during the first speaker's speaking, the continuous speaking duration of each new speaker is accumulated separately. For new speakers whose continuous speaking duration reaches the first duration threshold, a corresponding sub-screen is added to the composite screen; for new speakers whose continuous speaking duration has not reached the first duration threshold, no sub-screen is added temporarily.

[0067] In one embodiment, single-screen to composite screen conversion is based on speaking frequency: When the currently displayed single screen is a speaking area screen, the processing device detects that multiple sound sources are frequently speaking alternately within that area. The processing device counts the speaking frequency within a preset time window (e.g., 15 seconds). When the speaking frequency is greater than or equal to a frequency threshold (e.g., 6 times per minute), the processing device determines that a multi-person discussion is currently underway and sets the target screen display mode to composite screen mode. Based on the location status of each sound source, the processing device determines the content of the sub-screen corresponding to each sound source—for sound sources where a specific speaker can be located, the sub-screen is a close-up of that speaker; for sound sources where only the speaking area can be located, the sub-screen is the speaking area screen. The final output is a composite screen containing a panoramic view of the meeting and the aforementioned multiple sub-screens.

[0068] In one embodiment, the single-screen to combined-screen conversion is based on the location status upgrade: When the currently displayed single screen is a preset position screen of the first tracking area, the processing device detects that there are multiple sound sources in the first tracking area, and the location status of two of the sound sources upgrades from "not located to a specific speaker but located to the speaking area" to "located to a specific speaker". Based on this change in location status upgrade, and considering that the cumulative speaking time of each of the two sound sources has reached a second duration threshold, the processing device determines that both speakers need to be displayed simultaneously. The processing device determines that the target screen display mode is a combined-screen mode, generates a combined screen containing close-up shots of the first speaker and the second speaker, and retains the preset position screen of the first tracking area as a background reference in the combined screen.

[0069] Alternatively, if the current screen display mode is a combined screen mode, and based on the speech data and / or positioning status it is determined that the number of sub-screens in the combined screen needs to be reduced to one, then the target screen display mode is determined to be a single-screen mode. This case is referred to as Case Two.

[0070] For example, single-screen regression based on the end of a speech: If the currently displayed combined screen includes a first sub-screen (a close-up of speaker A) and a second sub-screen (a close-up of speaker B), and the speaking interval of speaker B is detected to be greater than or equal to a preset interval threshold (e.g., 5 seconds), it is determined that speaker B's speaking round has ended. At this time, if speaker A is still speaking and there are no other sound sources that need to be displayed, the processing device determines that the number of sub-screens in the current combined screen needs to be reduced to one. The target screen display mode is determined to be single-screen mode, and the combined screen is switched to a single close-up of speaker A.

[0071] In one embodiment, sub-screen merging is based on location status degradation: If the currently displayed combined screen contains three sub-screens—a first sub-screen (close-up of speaker A), a second sub-screen (close-up of speaker B), and a third sub-screen (close-up of speaker C)—and it is detected that the location status of two speakers (B and C) has been downgraded from "located to a specific speaker" to "not located to a specific speaker but located in the speaking area," and this downgrade continues for a preset duration (e.g., 3 seconds) without recovery. Simultaneously, the speaking intervals of speakers B and C both exceed a preset interval threshold, indicating that both have finished speaking. It is determined that only speaker A, who is still speaking and has a stable location, needs to be displayed. Therefore, the target screen display mode is set to single-screen mode, and a close-up shot of speaker A is output.

[0072] In one embodiment, single-screen retention based on priority filtering: When the currently displayed combined screen contains multiple sub-screens, the display priority of each speaker is calculated based on the speaker's speech data and image data corresponding to each sub-screen. When it is detected that all speakers except the one with the highest priority have finished speaking or the speaking interval exceeds a preset interval threshold, it is determined that the number of sub-screens in the combined screen needs to be reduced to one. The target screen display mode is determined to be single-screen mode, and the close-up of the speaker with the highest priority is retained as the output. For example, if the combined screen displays close-ups of Zhang San, Li Si, and Wang Wu, and Zhang San has the longest speaking time and is still speaking, while Li Si and Wang Wu have both stopped speaking for more than 5 seconds, then the combined screen is switched to a single close-up of Zhang San.

[0073] Alternatively, if the target screen display mode is the same as the current screen display mode, then the current screen display mode remains unchanged, and the content of the sub-screens in the single screen or combined screen is updated based on the positioning status. This case is referred to as Case 3.

[0074] For example, in single-screen mode, speaker switching occurs when the currently displayed single screen is a close-up of the first speaker. The processing device detects that the first speaker has finished speaking (the speaking interval is greater than or equal to a preset interval threshold), while the second speaker begins speaking and their cumulative speaking time reaches a first duration threshold. The processing device determines that the target screen display mode has not changed (it remains in single-screen mode), maintains the current screen display mode, but updates the screen content from a close-up of the first speaker to a close-up of the second speaker. Furthermore, if the first speaker has not finished speaking but the second speaker's speaking time percentage exceeds a second percentage threshold (e.g., 75%), the processing device can also switch the single screen from the first speaker to the second speaker, indicating that the current discussion's dominance has shifted.

[0075] In one embodiment, a sub-screen is added within the combined screen mode: when the currently displayed combined screen consists of two sub-screens (close-ups of the first speaker and the second speaker), the processing device detects that a third speaker has started speaking continuously and their cumulative speaking time has reached a first duration threshold. The processing device determines that the target screen display mode has not changed (it is still the combined screen mode), maintains the current screen display mode unchanged, and adds a sub-screen corresponding to the third speaker to the existing combined screen, forming a new combined screen containing three sub-screens. When adding a sub-screen, the processing device determines the display priority of each sub-screen based on the speaking data and rearranges the layout of the sub-screens according to priority—the sub-screen with the highest priority is displayed at the largest size in the main display area, and the remaining sub-screens are arranged at smaller sizes. For example, if Zhang San's speaking time is 8 seconds, Li Si's is 5 seconds, and Wang Wu's is 3 seconds, then Zhang San's sub-screen is located in the central main area of ​​the screen, and Li Si and Wang Wu's sub-screens are arranged at smaller sizes on both sides.

[0076] In one embodiment, the replacement and updating of sub-screens within a combined screen mode: When the currently displayed combined screen includes a first sub-screen (a close-up of speaker A) and a second sub-screen (a close-up of speaker B), the processing device detects that speaker A has finished speaking, while speaker C has started speaking and their cumulative speaking time has reached a first duration threshold. The processing device determines that the target screen display mode has not changed (it is still a combined screen mode), maintains the current screen display mode unchanged, and replaces the first sub-screen in the combined screen from a close-up of speaker A to a close-up of speaker C.

[0077] Furthermore, when the location status of the speaker corresponding to a sub-screen in the combined screen changes, the processing device can downgrade or upgrade the content of that sub-screen: Downgrade Update: If the location status of speaker B in the combined screen is downgraded from "located to specific speaker" to "not located to specific speaker but located to the speaking area", then the sub-screen will be updated from a close-up view of speaker B to a speaking area view of the area where speaker B is located, while the overall structure of the combined screen remains unchanged.

[0078] Upgrade / Update: If there is a speaking area in the combined screen, and the processing device subsequently pinpoints the specific speaker in that area, then the sub-screen will be upgraded from the speaking area screen to a close-up shot of the specific speaker.

[0079] In one embodiment, sub-screen removal within a combined screen mode: When the currently displayed combined screen contains multiple sub-screens, the processing device continuously monitors the speaking interval duration of the speaker corresponding to each sub-screen. When the processing device detects that the speaking interval duration of a speaker corresponding to a certain sub-screen is greater than or equal to a preset interval threshold, it determines that the speaker's speaking round has ended. The processing device removes the sub-screen corresponding to that speaker from the current combined screen, and the layout of the remaining sub-screens is automatically adjusted to fill the space of the removed sub-screen. Further, if the speaker corresponding to the removed sub-screen resumes speaking later and their cumulative speaking time again reaches the first duration threshold, the processing device can re-add the speaker's sub-screen to the combined screen. For example, in a combined screen containing sub-screens of Zhang San, Li Si, and Wang Wu, if Li Si stops speaking for more than 5 seconds and is removed, and then Li Si resumes speaking after half a minute and continues speaking for 3 seconds, then Li Si's sub-screen is re-added to the combined screen.

[0080] This application provides two modes of screen switching: cross-mode switching and intra-mode switching. Cross-mode switching refers to switching the screen display mode between a single-screen mode and a combined-screen mode. Intra-mode switching refers to updating the screen content while maintaining the current screen display mode. For example, switching from one speaker to another in a single-screen mode.

[0081] Taking scenario one (single screen → combined screen) as an example, the current display is a close-up of Zhang San. At this time, the system detects that Li Si also starts speaking, and his speaking data (such as cumulative duration) has reached the screen switching condition. The system determines that Zhang San and Li Si need to be displayed simultaneously, so it sets the target screen display mode to combined screen mode, and finally outputs a combined screen that includes close-up shots of Zhang San and Li Si.

[0082] Taking scenario two (combined screen → single screen) as an example: The current display is a combined screen of Zhang San and Li Si. Li Si finishes speaking, and the interval between his speeches exceeds the threshold. The system determines that only Zhang San needs to be displayed, so it sets the target screen display mode to single screen mode and finally outputs a close-up of Zhang San.

[0083] Taking scenario three (intra-mode switching) as an example: The current display shows a close-up of Zhang San. The system detects that Zhang San has finished speaking and Li Si has started speaking, and Li Si's speaking data meets the conditions for screen switching. The system determines that there is no need to change the screen mode (it remains single-screen mode), but only needs to update the screen content, switching the screen from a close-up of Zhang San to a close-up of Li Si.

[0084] In some embodiments, updating a single screen based on location status includes: If the current single screen is a close-up of the first speaker, and the location status is updated to the second speaker, and the second speaker meets the preset conditions, then the single screen will switch from the close-up of the first speaker to the close-up of the second speaker.

[0085] Among them, the preset conditions are the conditions for determining whether the second speaker needs to be shown in close-up, such as the second speaker's continuous cumulative speaking time exceeding a threshold, to ensure that the second speaker is a continuous speaker rather than a brief interruptor.

[0086] For example, the current screen is a close-up of speaker A. Speaker B then begins to speak. The system locates speaker B through sound source identification. The system begins to accumulate B's speaking time. When B's cumulative speaking time exceeds a preset threshold of 2 seconds, the system determines that B meets the preset condition and switches the single-screen view from a close-up of A to a close-up of B.

[0087] In some embodiments, updating a single screen based on location status includes: If the current single screen displays a close-up of a specific speaker, and the location status is downgraded from being located to the specific speaker to being located in the speaking area but not the specific speaker, then the single screen will switch from the close-up of the specific speaker to the corresponding speaking area screen.

[0088] Among them, the location status degradation refers to the system's location accuracy of the sound source decreasing from high to low, for example, from being able to accurately locate a specific individual to only being able to determine the general speaking area.

[0089] For example, the current screen is a close-up of speaker Zhang San. During his speech, Zhang San stands up and moves to a position where the camera's field of view is poor. The system can no longer pinpoint Zhang San's location, but can still determine that the sound is coming from the rear right area. At this point, the positioning status downgrades from "located to a specific speaker" to "not located a specific speaker but located in the speaking area." To prevent the screen from displaying an incorrect close-up, the system switches the single view from the close-up of Zhang San to the speaking area in the rear right region.

[0090] In some embodiments, updating the content of sub-screens in a combined screen based on positioning status includes: If the currently displayed composite screen includes a first sub-screen, based on the updated positioning status, perform one of the following update operations on the first sub-screen: Replace the first sub-screen with the close-up or speaking area screen corresponding to the first sound source, and replace it with the close-up or speaking area screen corresponding to the second sound source.

[0091] The first sub-screen has been downgraded from a close-up of the specific speaker to a screen showing the speaking area.

[0092] The first sub-screen is upgraded from the speaking area screen or the preset position screen to a close-up screen of the specific speaker.

[0093] Sub-pictures are the units that make up the composite picture. For example, it could be a close-up of a specific speaker or a picture of the speaking area. Replacement, downgrade, and upgrade are three operations for updating the content of sub-pictures, which could correspond to three scenarios, such as sound source change, decreased positioning accuracy, and increased positioning accuracy, respectively.

[0094] For example, a composite frame contains sub-frame A (close-up of Zhang San) and sub-frame B (close-up of Li Si). In a scene where the sub-frames are being replaced, Zhang San finishes speaking and Wang Wu begins speaking. The system replaces the close-up of Zhang San in sub-frame A with a close-up of Wang Wu.

[0095] In the scenario of sub-screen downgrading, in sub-screen B, Li Si's location status is downgraded from "located to specific speaker" to "located to speaking area". The system downgrades sub-screen B from a close-up of Li Si to a speaking area view.

[0096] In the sub-screen upgrade scenario, sub-screen C was originally the speaking area screen. The system accurately located the speaker Zhao Liu in that area, so sub-screen C was upgraded to a close-up of Zhao Liu.

[0097] In some embodiments, updating the content of sub-screens in a combined screen based on positioning status includes: If the currently displayed composite screen consists of at least one first sub-screen, based on the updated positioning status and / or speech data, at least one second sub-screen that needs to be displayed is re-determined, and the composite screen is switched as a whole to the target composite screen consisting of at least one second sub-screen.

[0098] Among them, between the set consisting of at least one second sub-picture and the set consisting of at least one first sub-picture, there is an addition, reduction, replacement or complete replacement of sub-pictures.

[0099] The above method is not limited to modifying a single sub-screen; it can also re-plan the entire set of sub-screens in the combined screen to generate a completely new set of sub-screens.

[0100] For example, the current composite screen contains sub-screen A (close-up of Zhang San) and sub-screen B (close-up of Li Si). In a scenario where additional sub-screens are added, Wang Wu begins to speak continuously, and the system determines that Wang Wu's sub-screen C (an example of the second sub-screen) needs to be displayed. In this case, the system can switch the entire composite screen to a new composite screen that includes sub-screens A, B, and C (close-up of Wang Wu).

[0101] In a scenario where the number of sub-screens is reduced, after Li Si finishes speaking, the system re-plans and generates a new combined screen that only contains sub-screen A.

[0102] In the scene where the sub-screen is replaced, Zhang San and Li Si stop speaking, Zhao Liu and Qian Qi begin to discuss, and the system re-plans and generates a new combined screen containing sub-screens D (close-up of Zhao Liu) and E (close-up of Qian Qi).

[0103] The solution provided in this embodiment offers the ability to switch the combined screen as a whole, which can more thoroughly respond to major changes in the speaking format of the meeting and make the screen layout highly matched with the current meeting progress.

[0104] For example, in one embodiment, a complete switch (addition of sub-screens) occurs when there is a significant change in the speaking layout: When the currently displayed combined screen consists of a first set of sub-screens (e.g., including close-ups of speaker A and speaker B), the processing device detects a large number of new sound sources in the meeting room—for example, at the start of the free discussion phase, five people suddenly speak simultaneously or sequentially in a previously quiet meeting room. Based on the updated positioning status and speaking data, the processing device determines that simply adding sub-screens one by one to the existing combined screen will not be sufficient to reasonably present all active speakers within the limited screen layout. In this case, the processing device no longer performs the operation of adding sub-screens one by one, but instead redetermines the second set of sub-screens to be displayed. For example, the processing device selects three speakers (C, D, E) from the five new sound sources whose cumulative speaking time is ≥ a second duration threshold. Combining this with the existing speakers A and B, the processing device redetermines the second set of sub-screens as close-up shots corresponding to speakers A, B, C, D, and E respectively, and switches the combined screen as a whole to the target combined screen consisting of these five sub-screens. In the new combined screen, the processing device recalculates the display priority based on the speaking data and positioning status of the five participants, and rearranges the size and position of the sub-screens according to the new priority. The beneficial effect of this embodiment is that when there is a significant change in the speaking pattern (such as from a two-person discussion to a multi-person debate), the overall switch can complete the reconstruction from the old layout to the new layout in one go, avoiding the problems of constantly adjusting the screen layout and difficulty in visual tracking for remote participants caused by adding each participant one by one.

[0105] In one embodiment, the overall switching (complete replacement of sub-screens) occurs when the speaking group is completely replaced: If the currently displayed combined screen consists of a first set of sub-screens (e.g., close-ups of speakers A, B, and C from the morning session), the meeting restarts after a break. The processing device detects that the speaking intervals of the original speakers A, B, and C have all exceeded a preset interval threshold, and new speakers D, E, and F have appeared in the meeting room. Based on the updated positioning status and speaking data, the processing device determines that the current set of sub-screens to be displayed is completely different from the original set. At this time, the processing device re-determines the second set of sub-screens to be displayed as close-up shots of speakers D, E, and F respectively, and switches the combined screen from the original "A+B+C" to the target combined screen of "D+E+F". Furthermore, if in the new speaking group, speaker D's location status is "located to a specific speaker," speaker E's location status is "not located to a specific speaker but located to the speaking area," and speaker F's location cannot be accurately located, then the second sub-screen set includes a close-up view of speaker D, a speaking area view of speaker E's location, and a preset position view of speaker F's tracking area. Three different granularities of images are presented simultaneously in a combined image, each adapting to the current location accuracy. The beneficial effect of this embodiment is that when the speaking group undergoes a complete change, it avoids the screen clutter caused by simultaneously retaining the old sub-screens of finished speaking and the newly added sub-screens in the combined image, allowing remote participants to clearly perceive the change in the speaking pattern.

[0106] In one embodiment, the scene transition is phased: when the speaking groups do not change simultaneously but in stages, the processing device does not immediately perform a full switch, but sets a transition observation period (e.g., 10 seconds). During the transition observation period, the processing device continuously collects updated location status and speaking data. If, at the end of the transition observation period, the overlap ratio between the first sub-screen set and the currently active sound source set is lower than a threshold (e.g., lower than 30%), then a full switch is performed; if the overlap ratio is still higher than the threshold, then the method of updating one by one continues to maintain screen stability. For example: the combined screen displays close-ups of speakers A, B, and C. Then A finishes speaking, and D joins; then B finishes speaking, and E joins. If the processing device immediately evaluates after "B finishes speaking, E joins" and finds that only C among A, C, D, and E comes from the original set (overlap rate of 25%), which is lower than the 30% threshold, then a full switch is triggered, and the combined screen is reconstructed into a target combined screen with D and E as the main focus and C as the secondary focus.

[0107] In one embodiment, scene switching based on overall change in positioning status (reduction and replacement of sub-screens): When the currently displayed combined screen consists of a first set of sub-screens (e.g., containing four sub-screens: close-up shots of speakers A and B, a speaking area shot of region X, and a preset position shot of region Y), the processing device detects the following comprehensive changes in positioning status: speaker A has finished speaking (speaking interval ≥ preset interval threshold), speaker B is still within the tracking area but the positioning status has been downgraded from "located to a specific speaker" to "located to a speaking area," and the sound source in region Y has been upgraded to be able to locate a specific speaker C. Based on the above comprehensive changes, the processing device redetermines the second set of sub-screens to be displayed: removes the sub-screen of speaker A, downgrades the close-up of speaker B to a region shot (i.e., the speaking area shot of the region where B is located, still occupying a sub-screen position), and adds a close-up shot of speaker C. The processing device does not operate one by one, but regenerates the complete set of sub-screens according to the updated positioning status, and then switches to the target combined screen as a whole. Furthermore, during the overall switch, the processing device can recalculate the layout parameters based on the new number of sub-screens. If the number of sub-screens in the new set is different from that in the original set, the display size and arrangement of the sub-screens will also be adjusted accordingly to accommodate the new number of sub-screens. For example, when the number of sub-screens changes from 4 to 3, the display area of ​​each sub-screen increases accordingly, and the arrangement changes from a 2×2 grid to a horizontal three-part arrangement.

[0108] In one embodiment: A semantic analysis-based overall topic switching (complete replacement of sub-screens): When the currently displayed composite screen consists of a first set of sub-screens, the processing device performs real-time semantic analysis on the conference audio stream and detects that the topic of the conference discussion has switched from "progress report of Project A" to "discussion of issues in Project B". Based on the relevance of the speech content to the current conference topic, combined with the location status and speech data, the processing device redetermines the second set of sub-screens to be displayed—removing sub-screens of speakers unrelated to "Project B" from the original set and replacing them with sub-screens of speakers highly relevant to "Project B". The processing device then switches the composite screen from the old set to the target composite screen composed of the new set. For example: The first half of the composite screen displays a close-up of Zhang San (reporting on the progress of Project A). When the topic switches to Project B, the system, through semantic analysis, identifies that Li Si's speech content regarding Project B is highly relevant to the current topic, and Li Si's cumulative speaking time exceeds a threshold. The newly generated second set of sub-screens includes a close-up of Li Si and sub-screens of other participants in Project B. After the overall switch, remote participants see a composite screen highly matching the current discussion topic. The beneficial effect of this embodiment is that, based on semantic content-driven screen switching, the content of the combined screen is not limited to "who is speaking", but further evolves to "who is speaking content related to the current topic", which improves the information transmission efficiency of the combined screen.

[0109] In some embodiments, the speech data includes a first continuously accumulated speech duration and / or the speech frequency within a preset time window; wherein, the first continuously accumulated speech duration represents the total duration obtained by accumulating the duration of each speech that occurs within the tracking area, starting from the first speech in any speech round, until the interval between two adjacent speeches is greater than or equal to a preset interval threshold.

[0110] Based on the speech data, determine whether the conditions for switching the screen are met, including: If the first continuous cumulative speaking duration is greater than or equal to the first duration threshold, and / or the speaking frequency is greater than or equal to the frequency threshold, then the screen switching condition is determined to be met.

[0111] If the first continuous cumulative speaking duration is less than the first duration threshold, and / or the speaking frequency is less than the frequency threshold, then the screen switching condition is determined not to be met.

[0112] The first continuous cumulative speaking time is for the entire tracking area, and does not focus on specific speakers, but only on whether anyone in the area is speaking continuously.

[0113] The first duration threshold (which can be denoted as T0) is the regional duration threshold that triggers screen switching.

[0114] The preset interval threshold (which can be denoted as T1) is used to distinguish between short pauses within a speaking round and the end of a speaking round.

[0115] In this embodiment, two consecutive speeches can be from different speakers. For example, speaker A speaks once, and speaker B speaks once. For instance, if only one speech occurs in the tracking area within a certain time period, the first continuous cumulative speech duration is the duration of that speech.

[0116] For example, within the tracking area in the left front zone, Zhang San speaks for 2 seconds, pauses for 1 second (less than T1), and Li Si then speaks for 3 seconds. Therefore, the first continuous cumulative speaking time in this area is 2 + 3 = 5 seconds. If the first duration threshold is set to 4 seconds, since 5 seconds ≥ 4 seconds, the system determines that the screen switching condition is met and can trigger a switch from the panoramic view to the preset position screen for this area.

[0117] This disclosure introduces a hierarchical delay switching model, whose core parameters include: a cumulative speaking duration threshold T0 and a speaking interval duration threshold T1. T0 includes a first duration threshold at the region level and a second duration threshold at the individual sound source level. T1 includes a preset interval threshold at the region level and a preset interval threshold at the individual sound source level. The region-level thresholds and the individual sound source-level thresholds can be set uniformly or separately, without restriction. The region-level thresholds may also include thresholds for the speaking region and the tracking region.

[0118] The processing device can maintain corresponding timers for the speaker (a single sound source), the speaking area, and the tracking area. As one possible implementation, when a microphone detects a sound source, indicating that the tracking area is activated, the timers for the tracking area, the speaker, and the speaking area all start simultaneously. Specifically, the timers for each tracking area are maintained based on the sound source localization results of the global microphone (such as a ceiling-mounted microphone). The timers for the speaker or speaking area are maintained based on the sound source localization results of microphones within the tracking area (such as the built-in microphone of a camera within the tracking area).

[0119] As one possible implementation, T1 can also be referred to as the timer reset interval threshold. When the interval between speeches exceeds T1, the timer is reset, and speeches after that interval are considered new speaking rounds, with their durations reset to zero. When the interval is less than T1, the timer is not reset, and speeches after that interval are considered continuations of the same speaking round, with their durations continuing to accumulate. T1 can be used to distinguish between pauses within a sentence and the end of a speech.

[0120] like Figure 3 In case (a), if the interval between two speeches by the speaker is less than T1, the timer is not reset, and the speech duration continues to accumulate. When the accumulated speech duration t reaches T0, the processing device switches to a close-up view of the speaker.

[0121] like Figure 3 (b) If the interval between two speeches by a speaker is less than T1, but the cumulative duration of the two speeches does not reach T0, the processing device will not switch the screen to a close-up of the speaker in order to reduce the probability of frequent screen switching.

[0122] like Figure 3 (c) If a speaker does not make a new speech after the last speech and after an interval of T1, the speaker's timer is reset and the accumulated speaking time t is cleared.

[0123] The following section describes the calculation method for the cumulative speaking time involved in this disclosure, categorized by scenario: For a tracking area, the continuous cumulative speaking duration of the tracking area refers to the total speaking duration of all speakers within the tracking area, regardless of whether they are the same speaker. As long as there are people speaking continuously in the tracking area, the continuous cumulative speaking duration accumulates continuously.

[0124] For a tracking area, as long as the interval duration between adjacent speeches in the tracking area < T1, the timer of the tracking area is not reset. For a speaker or a speaking area, as long as the interval duration between adjacent speeches of the speaker or the speaking area < T1, the timer of this speaker or speaker area is not reset.

[0125] Scenario 1: Intermittent speaking by the same speaker within the same tracking area The interval duration between speeches of the same speaker < T1: The timer is not reset, and the continuous cumulative speaking duration accumulates continuously. This indicates that the speaker is just pausing naturally, such as for breathing or thinking, and still belongs to the same speaking round.

[0126] The interval duration between speeches of the same speaker ≥ T1: The timer is reset, and the continuous cumulative speaking duration is reset to zero. This indicates that the speaking round of this speaker has ended, and subsequent speeches belong to a new speaking round.

[0127] Scenario 2: Intermittent speaking by the same speaker and a position change occurs within the same tracking area The interval duration between speeches of the same speaker < T1: The personal timer is not reset, and the continuous cumulative speaking duration accumulates continuously.

[0128] The interval duration between speeches of the same speaker ≥ T1: The personal timer is reset. If there is only one speaker in the tracking area, the timer of the tracking area will also be reset.

[0129] The speaker moves slightly within the same tracking area, such as standing up from a seat. As long as the speaker continues to speak and the interval duration between speeches does not exceed T1, it is still regarded as a continuation of the same speaking round.

[0130] Scenario 3: Intermittent speaking by the same speaker and a position change occurs in different tracking areas The interval duration between speeches of the same speaker < T1: The personal timer is not reset, and the speaking round continues. At this time, if the continuous cumulative speaking duration of the individual reaches the standard, the processing device can control the display of a close-up image of the speaker in the new tracking area. For example, based on the preset position or camera repositioning in the new tracking area to generate an image.

[0131] The interval duration between speeches of the same speaker ≥ T1: The personal timer is reset, regarded as a new speaking round. In this case, the personal speaking duration needs to be re-accumulated in the new tracking area. When the continuous cumulative speaking duration reaches T0, the processing device can trigger the control to display a close-up image of this speaker.

[0132] In this scenario, since the speaker moves to different tracking regions, at this time, whether the timers of these two tracking regions are reset depends on whether the total sound source interruption duration of the tracking region exceeds T1, without considering the specific speaker.

[0133] In this way, it can be ensured that when the speaker moves across regions, the picture can smoothly transition to the close-up of the new tracking region, rather than reverting to the panoramic picture and then entering the close-up picture again.

[0134] Scenario 4: Different speakers in the same tracking region speak successively If the speaking interval duration between the previous and the next speaker < T1: The timer of the tracking region is not reset, and the continuous cumulative speaking duration of the tracking region (the first continuous cumulative speaking duration) continues to accumulate. However, the personal timer of the new speaker starts from zero, and the personal timer of the previous speaker is not reset. The reset condition of the personal timer of the speaker is that the speaking interval duration of the speaker ≥ T1.

[0135] If the speaking interval duration between the previous and the next speaker ≥ T1: Both the tracking region timer and the personal timer are reset.

[0136] Among them, when different speakers speak in the same tracking region or different tracking regions, each speaker maintains an independent speaking timer without interference. The system starts its independent timer for the first detected speaker. Exemplarily, for a scenario with only a single speaker displayed, the picture preferentially displays the speaker who first meets the condition that the speaking duration ≥ T0.

[0137] Scenario 5: Multiple speakers all meet the condition that the speaking duration ≥ T0 For a scenario with only a single speaker displayed, determine the speaker who first meets this condition, and preferentially display the close-up picture or the speaking area picture of this speaker.

[0138] If multiple speakers all meet this condition, the multi-view mode can be enabled to simultaneously display the close-up pictures of these multiple speakers. If there is a quantity limit, then from the multiple speakers who meet the condition, select the corresponding number of speakers in the order of appearance for display.

[0139] In some embodiments, the single picture includes the close-up picture of the specific speaker, the speaking area picture, and the preset position picture corresponding to the tracking region; based on the positioning state and the target picture display mode, the corresponding single picture is output, including: In the case where the target picture display mode is the single picture mode, if the positioning state is to locate the specific speaker, the close-up picture of the specific speaker is output; if the positioning state is not to locate the specific speaker but to locate the speaking area, the speaking area picture is output; if the positioning state is that neither the specific speaker nor the speaking area is located, the preset position picture corresponding to the tracking region is output.

[0140] For example, a preset position view is a baseline image captured by the camera at a preset position in a certain tracking area, without further cropping or magnification. A close-up view is a single-person image of a specific speaker after optical or digital zoom. A speaking area view is an image of a specific speaking area.

[0141] For example, when the system is in single-screen mode: if the speaker Zhang San can be accurately located, a close-up of Zhang San's face is output. If the sound is determined to be coming from the left front area but Zhang San cannot be located, the speech area of ​​the left front area is output. If neither the left front area nor Zhang San is located, the preset position of the left front area is output.

[0142] The solution provided in this embodiment offers corresponding screen output schemes for the three positioning states in single-screen mode, forming a degraded display chain of "close-up → area → preset position", which ensures reasonable screen output under any positioning accuracy.

[0143] In some embodiments, the speech data includes a second continuously accumulated speech duration; wherein, the second continuously accumulated speech duration represents the total duration obtained by accumulating the duration of each speech of a single sound source, starting from the first speech of the current speech round of the single sound source, until the interval between two adjacent speeches is greater than or equal to a preset interval threshold.

[0144] When the location status indicates that a specific speaker has been located, output a close-up shot of the specific speaker, including: When the location status is that a specific speaker has been located, and the second continuous cumulative speaking time is greater than or equal to the second duration threshold, output a close-up shot of the specific speaker; If the location status is "no specific speaker located but the speaking area is located", output the speaking area screen, including: If the location status is that the specific speaker is not located but the speaking area is located, and the second continuous cumulative speaking time is greater than or equal to the second time threshold, the speaking area screen is output. If the location status indicates that neither the specific speaker nor the speaking area has been located, output the preset position screen corresponding to the tracking area, including: If the location status is such that neither the specific speaker nor the speaking area is located, or if the second cumulative speaking time is less than the second duration threshold, the preset position screen corresponding to the tracking area is output.

[0145] The second continuous cumulative speaking duration is for a single speaker (or a single sound source) and is used to determine whether that speaker needs to be shown in close-up. The second duration threshold (example of T0) is the individual duration threshold that triggers the close-up. The preset interval threshold is used to determine whether the speaker's speaking round has ended. In this embodiment, two adjacent speeches refer to the speeches of the same speaker, such as speaker A's first speech and speaker A's second speech. For example, if a single sound source does not speak again after its initial speech, the second continuous cumulative speaking duration is the duration of that speech.

[0146] For example, the system locates speaker Zhang San and begins accumulating his speaking time. If Zhang San's second consecutive accumulated speaking time reaches 2 seconds (the second duration threshold), the system outputs a close-up shot of him. If Zhang San speaks for only 1 second and then stops speaking, or if the system fails to locate Zhang San and his speaking area, the system does not output a close-up shot. Instead, it outputs a preset shot based on the location status to avoid switching close-up shots for shorter speeches.

[0147] The solution provided in this embodiment introduces a second, continuous, cumulative speaking time for each individual as a condition for outputting close-up shots in single-screen output, ensuring that only individuals who speak continuously can receive close-up shots. This further enhances the anti-shake capability of the output screen.

[0148] In some embodiments, the combined view includes a panoramic view of the conference and sub-views; based on the positioning status and the target view display mode, the corresponding combined view is output, including: When the location status indicates that a specific speaker has been located, the sub-screen is set to a close-up view of that speaker; when the location status indicates that a specific speaker has not been located but the speaking area has been located, the sub-screen is set to the speaking area view; when the location status indicates that neither the specific speaker nor the speaking area has been located, the sub-screen is set to the preset position view corresponding to the tracking area; finally, a combined view including the panoramic view of the meeting and the sub-screens is output.

[0149] The panoramic view of the meeting can be a full-scene shot taken with a wide-angle camera, covering most of the meeting room space. Sub-views can be magnified views superimposed on the panoramic view, used to show details of the speaker or speaking area.

[0150] For example, the system is in composite view mode. The panoramic view of the meeting is always displayed. When the system locates speaker Zhang San, a close-up sub-view of Zhang San will be overlaid in the lower right corner of the panoramic view. If the system can only locate the front left area, the sub-view will display the speaking area of ​​the front left area.

[0151] In some embodiments, when multiple sound sources are detected and each sound source meets the screen switching conditions, a combined screen including a panoramic view of the conference and sub-screens is output, including: outputting a combined screen including a panoramic view of the conference and multiple sub-screens, wherein the multiple sub-screens correspond one-to-one with the multiple sound sources.

[0152] Among them, multiple sub-views refer to the fact that a combined view can contain multiple independent magnified views at the same time, with each view corresponding to an active sound source.

[0153] For example, in a meeting, three people, Zhang San, Li Si, and Wang Wu, are speaking simultaneously, and each person's cumulative speaking time has reached their individual threshold. The system can output a combined view, including a panoramic view and three close-up sub-views corresponding to Zhang San, Li Si, and Wang Wu, respectively.

[0154] The solution provided in this embodiment supports displaying close-ups of multiple active speakers simultaneously in a combined screen, meeting the information display needs in scenarios of multi-person parallel discussions.

[0155] In some embodiments, when multiple sub-screens are displayed, the method further includes: Based on speech data from multiple sound sources, the display priority of multiple sub-screens is determined; the sub-screen with the highest display priority is displayed in the preset main display area of ​​the combined screen; or, the corresponding sub-screens are displayed in different sizes according to the order of display priority from high to low, wherein the display size of the sub-screen with higher display priority is larger than that of the sub-screen with lower display priority.

[0156] The display priority is determined by ranking the sub-screens according to the importance of speaking data (such as cumulative speaking time, speaking frequency, etc.). The preset main display area is the most visually prominent position in the combined screen, such as the center or upper left corner of the screen.

[0157] For example, the combined screen contains three sub-screens. Based on the percentage of each person's speaking time, the system determines that Zhang San has the highest priority, followed by Li Si, and then Wang Wu. The system then places a close-up of Zhang San in the main display area (such as a large window in the center of the screen), while close-ups of Li Si and Wang Wu are arranged to the sides at smaller sizes.

[0158] The solution provided in this embodiment introduces priority sorting, which makes the layout of the combined screen more hierarchical, highlights the most important speaker, and improves the efficiency of information transmission.

[0159] In some embodiments, when multiple sub-screens contain the same speaker, duplicate sub-screens are deduplicated or merged.

[0160] Deduplication refers to retaining only one sub-screen when multiple sub-screens actually point to the same speaker. Merging refers to combining multiple sub-screens pointing to the same speaker into one.

[0161] For example, two different microphone arrays might both identify the same speaker, Zhang San, as the sound source, resulting in two sub-pictures of "Zhang San in close-up" in the combined image. The system can detect this duplication and either retain only one of them or merge the two sub-pictures into one.

[0162] The solution provided in this embodiment avoids redundancy and confusion caused by repeatedly displaying the same person on the screen, ensuring the neatness of the screen and the accuracy of the information.

[0163] In some embodiments, after outputting a combined view including a panoramic view of the conference and multiple sub-views, the method further includes: if a new sound source is detected and the cumulative speaking time of the new sound source exceeds a preset duration threshold, then adding a sub-view corresponding to the new sound source to the combined view.

[0164] Among them, the preset duration threshold is the duration threshold for determining whether a new sound source needs to be added to the combined image.

[0165] For example, the current combined screen displays close-ups of Zhang San and Li Si. At this moment, Wang Wu begins to speak, and his cumulative speaking time exceeds a preset threshold of 3 seconds. The system determines that Wang Wu is also a continuous speaker, and therefore adds a close-up sub-screen of Wang Wu to the combined screen.

[0166] The solution provided in this embodiment realizes the dynamic expansion capability of the combined screen, which can add newly emerging continuous speakers to the screen in a timely manner, so that the screen content keeps pace with the meeting process.

[0167] In some embodiments, after outputting a combined view including a panoramic view of the conference and multiple sub-views, the method further includes: if there is a sub-view among the multiple sub-views with a speaking interval duration greater than or equal to a preset interval threshold, then remove the sub-view with a speaking interval duration greater than or equal to the preset interval threshold from the combined view.

[0168] The speaking interval refers to the time interval between the last time a speaker in a certain sub-screen spoke.

[0169] For example, the composite screen shows close-ups of three people: Zhang San, Li Si, and Wang Wu. After Li Si finishes speaking, the interval between his speeches exceeds 5 seconds (a preset interval threshold). The system determines that Li Si's speaking round has ended and removes the close-up sub-screen of Li Si from the composite screen.

[0170] The solution provided in this embodiment enables dynamic reduction of the combined screen, allowing for timely removal of screens of speakers who have finished speaking, thus maintaining a concise and effective screen layout.

[0171] In some embodiments, after determining the tracking area where the sound source is located, the method further includes: driving the camera bound to the tracking area to a preset position within the tracking area. Then, when outputting the target conference view, it is necessary to control the switch from a panoramic view of the conference to the target conference view while the camera is rotated to the preset position. The close-up view of a specific speaker, the speaking area view, or the preset position view in the target conference view is obtained by cropping the image captured by the camera at the preset position.

[0172] One possible implementation is to drive the camera to a preset position by physically rotating the PTZ camera to that position. For example, if the system detects speech in the left front area, it drives the PTZ camera in that area to rotate to a preset position in that area. During the camera rotation, the image still displays a panoramic view. Once the camera is in position, the system can crop a close-up or regional image from the preset position image captured by the camera, based on the positioning status, for display.

[0173] For example, if the cumulative speaking time in the tracked area is ≥ T0, the processing device can call the PTZ camera to a preset position. If the PTZ camera has been moved to the preset position, but the processing device has not yet located the specific speaker or the speaking area with high confidence, the processing device will control the display of the screen at the preset position. Once the processing device locates the specific speaker or the speaking area with high confidence, the processing device will then switch to a close-up view of the speaker or the speaking area.

[0174] For example, if the cumulative speaking time in the tracking area is ≥ T0, the processing device can call the PTZ camera to a preset position. If the PTZ camera has been moved to the preset position, and the processing device has located the specific speaker, and the speaker's cumulative speaking time is ≥ T0, then the processing device controls the display of a close-up image of the speaker.

[0175] For example, if the cumulative speaking time in the tracking area is ≥ T0, the processing device can call the PTZ camera to a preset position. If the PTZ camera has been moved to the preset position, and the processing device has located the speaking area, and the cumulative speaking time in that speaking area is ≥ T0, then the processing device controls the display of a close-up view of that speaking area.

[0176] The gimbal camera in this disclosure can be a PTZ camera or a PTZ lens module in a camera, without limitation.

[0177] The solution provided in this embodiment can decouple the physical rotation (PTZ) of the camera from the image switching, allowing the camera to rotate to the correct position before switching the image, thus achieving a smooth transition between images.

[0178] In some embodiments, when outputting a preset position image corresponding to the tracking area, the method further includes: When the location status is upgraded to locate a specific speaker, and the second continuous cumulative speaking time is greater than or equal to the second duration threshold, control the switch from the preset position screen to a close-up screen of the specific speaker; If the location status is upgraded to not locating a specific speaker but locating the speaking area, and the second continuous cumulative speaking time is greater than or equal to the second duration threshold, the control switches from the preset position screen to the corresponding speaking area screen.

[0179] Among them, the positioning status upgrade refers to the system's positioning accuracy of the sound source being improved from low to high, for example, from only being able to determine the area to being able to accurately locate a specific individual.

[0180] For example, after controlling the display of the preset position, if the processing device locates a specific speaker and the speaker's continuous cumulative speaking time exceeds T0, the screen can be switched from the preset position to a close-up of the speaker.

[0181] For example, after controlling the display of the preset position, if the processing device locates a high-confidence speaking area and the continuous cumulative speaking time of the speaking area exceeds T0, the screen can be switched from the preset position to the close-up area.

[0182] The solution provided in this embodiment offers a path to upgrade to a higher-level screen (close-up) based on the preset screen position, forming a progressive display of "preset position → close-up," which helps improve the display effect of the meeting screen. Furthermore, by determining whether the cumulative speaking time exceeds T0, invalid short-duration speeches can be filtered out, reducing frequent screen switching.

[0183] In some embodiments, when outputting a close-up view of a specific speaker, whether in single-screen mode or a sub-screen in combined-screen mode, the method further includes: If a speaker continues speaking, maintain a close-up view of that speaker. If a speaker finishes speaking, determine if a new speaker has been detected. If no new speaker is detected, switch back to the overall meeting view or the preset view corresponding to the tracking area. If a new speaker is detected, first switch back to the overall meeting view or the preset view corresponding to the tracking area where the new sound source is located. If the new speaker's cumulative speaking time meets the preset continuous speaking conditions, switch to a close-up view or the speaking area view corresponding to the new speaker, based on the new speaker's location status.

[0184] The condition for ending a speaker's speech can be defined as the interval between speakers being greater than or equal to a preset interval threshold. The condition for continuing to speak is defined as the cumulative speaking time of a new speaker reaching a second duration threshold.

[0185] For example, the current view shows a close-up of Zhang San. If Zhang San continues speaking, the view remains unchanged. If Zhang San pauses for more than T1 after saying a sentence, the system determines that his speech has ended. At this time, if no one else speaks, the view switches back to a full view or a preset position. If Li Si (the new speaker) is detected to start speaking, the system can first switch back to a full view or a preset position view of Li Si's area, and then switch to a close-up of Li Si or the speaking area view after Li Si's cumulative speaking time reaches a threshold.

[0186] The solution provided in this embodiment offers a mechanism for screen rewind after a speaker finishes speaking and for a new speaker to take over. By employing a strategy of rewinding first and then switching back, a smooth screen transition can be achieved, reducing the abruptness of direct jumps.

[0187] In some embodiments, a close-up shot of the speaker is maintained even if the speaker has not finished speaking, including: In the event that a specific speaker's movement is detected: If the specific speaker does not go beyond the tracking area, the camera bound to the tracking area will continue to output a close-up image of the specific speaker; If a specific speaker is outside the tracking area, a new tracking area is determined where the specific speaker is located. The screen is first switched to the preset position screen corresponding to the new tracking area. Then, based on the positioning status of the specific speaker by the camera bound to the new tracking area and the screen display mode configured in the new tracking area, the preset position screen corresponding to the new tracking area is switched to a close-up screen of the specific speaker or the corresponding speaking area screen.

[0188] "Outside the tracking area" means the speaker's position has moved outside the coverage area of ​​the current camera or preset position. The new tracking area is another preset area where the speaker is located after moving.

[0189] For example, when Zhang San moves from the front left area to the rear right area while speaking, if Zhang San is still within the front left area, the camera in the front left area continues to track him and output a close-up. If Zhang San moves out of the front left area, the system can switch the view to a preset position in the rear right area, and then the camera in the rear right area will reposition Zhang San. Once the conditions are met, the system will output a close-up of Zhang San in the rear right area or a view of his speaking area.

[0190] The solution provided in this embodiment enables seamless relay tracking of the speaker's image when the speaker moves across regions, avoiding the fragmented experience of "losing track → rewinding to the panoramic view → repositioning".

[0191] In some embodiments, when a speaker finishes speaking and no new speaker is detected, the system switches back to a panoramic view of the meeting or a preset view corresponding to the tracking area, including: In the event that a specific speaker's movement is detected: If the specific speaker is within the tracking area, switch back to the panoramic view of the meeting or the preset position view corresponding to the tracking area; If a specific speaker goes out of the tracking area, switch back to the panoramic view of the meeting; or, determine the new tracking area where the specific speaker is located and switch back to the preset position view corresponding to the new tracking area.

[0192] For example, after Zhang San finishes speaking, if no new speaker is detected and Zhang San has not left the current tracking area (front left area), the screen can switch back to the preset position of the current tracking area or the panoramic view. If after Zhang San finishes speaking, if no new speaker is detected and Zhang San has moved to the rear right area, the screen can switch back to the panoramic view or switch to the preset position of the rear right area.

[0193] The solution provided in this embodiment takes into account the speaker's positional movement after the speaker finishes speaking, and offers a more flexible option to go back, making the rewind screen more contextually relevant.

[0194] In some embodiments, when outputting a close-up view of a specific speaker, the method further includes: if the positioning state is downgraded from positioning to a specific speaker to not positioning a specific speaker but positioning to the speaking area, and the specific speaker is not located again for a preset time, controlling the switch to the corresponding speaking area view; and if the specific speaker is subsequently re-determined, switching the speaking area view to a close-up view of the specific speaker.

[0195] The preset duration is a waiting time window used to confirm whether the location downgrade is temporary or permanent.

[0196] For example, a close-up shot of Zhang San is currently being displayed. Zhang San turns his head to look elsewhere, and the system is temporarily unable to locate him, causing the location status to downgrade. At this point, the system will not immediately switch the view, but will wait for a preset duration (e.g., 2 seconds). If the system still cannot locate Zhang San after 2 seconds, the view will switch to the chat area where Zhang San is located. If Zhang San then turns back, and the system re-locates him, the view will switch back to the close-up shot of Zhang San.

[0197] The solution provided in this embodiment avoids immediate screen switching due to temporary location loss by downgrading and waiting for a continuously preset duration, thereby improving screen stability.

[0198] In some embodiments, the method further includes: in the same tracking session, if the location status for the same or similar sound source is downgraded from being located to a specific speaker to not being located to a specific speaker but being located to the speaking area for the second time, the current output close-up of the specific speaker is maintained and the corresponding speaking area is not switched.

[0199] In this context, "same tracking session" can refer to the process from when the system starts tracking a sound source until the sound source disappears. "Second occurrence of degradation" refers to the repeated occurrence of the degradation event for the same or nearby sound source.

[0200] For example, in a tracking session, Zhang San's close-up view has switched to a different area due to location degradation, and then returned to close-up. If the system experiences location degradation for Zhang San again, it will determine that this is edge jitter and will not switch the view, but will maintain the current close-up view.

[0201] The solution provided in this embodiment offers a shake-prevention protection mechanism. During the switching process from close-up of a person to a close-up of an area, if a second instance occurs where the specific person cannot be identified but the area can be, the system will not attempt to switch back to the close-up area view. Instead, it will maintain the current close-up of the person or the close-up view from the last successful positioning, until a substantial change in the sound source occurs, at which point the system will track and replace the person. In other words, when the same or similar sound sources repeatedly cause positioning degradation, the system will be "immune" to the second and subsequent degradation events, effectively suppressing repeated switching between close-ups and areas. For example, it can reduce repeated image switching caused by fluctuations in positioning accuracy near boundaries.

[0202] In some embodiments, when outputting a close-up shot of a specific speaker, the method further includes: If a new speaker is detected before the specific speaker has finished speaking, and the new speaker or the new speaker's speaking area can be located, and the new speaker's continuous cumulative speaking time exceeds a preset time threshold, then it is determined whether the new speaker and the specific speaker are located in the same tracking area. If they are in the same tracking area and the new speaker can be located, the screen will switch to one of the following: a single screen that includes both the specific speaker and the new speaker; a preset screen for the tracking area; a screen that is stitched together from close-up shots of the specific speaker and the new speaker; or a combined screen that includes a panoramic view of the meeting and close-up shots of the specific speaker and the new speaker. If they are in the same tracking area and the speaker’s speaking area can be located, the screen will switch to one of the following: a single screen that includes both the specific speaker and the new speaker; a preset screen of the tracking area; or a combination screen that includes a panoramic view of the meeting, a close-up view of the specific speaker, and a screen of the speaker’s corresponding speaking area. If they are not in the same tracking area, and the new speaker can be located, the screen will switch to one of the following: a screen that is stitched together with a close-up of the specific speaker and a close-up of the new speaker; or a screen that includes a panoramic view of the meeting and close-up shots of the specific speaker and the new speaker respectively. If they are not in the same tracking area, and the speaking area of ​​the new speaker can be located, then the screen will switch to one of the following: a screen that is a combination of a close-up of the specific speaker and a screen of the new speaker's speaking area, or a screen that includes a panoramic view of the meeting and a close-up of the specific speaker and a screen of the new speaker's speaking area.

[0203] For example, the system currently displays a close-up of speaker Zhang San, and Li Si begins speaking, with Li Si's cumulative speaking time reaching a preset threshold. The system can provide various screen switching options based on whether Li Si and Zhang San are in the same area and whether Li Si can be located, such as close-ups of both in the same area or spliced ​​images from different areas.

[0204] The solution provided in this embodiment offers rich and granular screen switching strategies for the complex scenario of "the current speaker has not finished speaking and a new speaker appears," enabling the system to select the optimal display solution based on the actual situation.

[0205] In some embodiments, when outputting a close-up shot of a specific speaker, the method further includes: If a new speaker is detected before a specific speaker has finished speaking, determine the percentage of speaking time for each speaker within the preset time period; Based on the speaking time percentage of each speaker, target speakers are determined from existing speakers and new speakers, where the speaking time percentage of the target speakers is greater than the first percentage threshold. When the target speaker is not a specific speaker, the camera attached to the tracking area where the target speaker is located will pre-select the target speaker; If the target speaker's speaking time exceeds the second threshold and the target speaker has not finished speaking, switch to a close-up shot of the target speaker.

[0206] Among them, the speaking time percentage refers to the proportion of a speaker's speaking time within a preset time period to the total speaking time. Pre-selection refers to driving the camera to rotate to the target position in advance without switching the output view. The first percentage threshold and the second percentage threshold are two progressive percentage thresholds.

[0207] For example, the system currently displays a close-up of speaker Zhang San, and Li Si begins to interrupt. The system calculates the percentage of speech by each speaker over the past 15 seconds. If Li Si's speaking percentage exceeds 50% (the first percentage threshold), the system can pre-select Li Si using the camera. If Li Si's speaking percentage continues to rise, exceeding 75% (the second percentage threshold), the system can switch the view from Zhang San to Li Si.

[0208] The solution provided in this embodiment pre-selects and delays switching based on the proportion of speakers, decoupling decision-making from execution. The screen is switched only after it is confirmed that the new speaker has indeed taken the lead, avoiding invalid switching caused by brief interruptions and frequent rotation of the PTZ pan-tilt unit.

[0209] In some embodiments, when multiple speakers are detected, the method further includes: Based on each speaker's speaking data and image data, determine the priority of each speaker; If the screen display mode is single speaker mode, then the speaker with the highest priority will be selected as the target speaker; If the screen display mode is multi-speaker mode, then select the number of speakers displayed in multi-speaker mode as the target speaker in descending order of priority; Output the target meeting screen, including: Output the target meeting screen corresponding to the target speaker.

[0210] Priority is a score calculated by combining spoken data and image data, used to measure the presentation value of each speaker. Single-speaker mode and multi-speaker mode are two different screen display modes, corresponding to displaying one or more speakers respectively.

[0211] For example, in a meeting, there are three people: Zhang San, Li Si, and Wang Wu. The system calculates their priority based on data such as speaking time, volume, and facial orientation: Zhang San > Li Si > Wang Wu. If the current mode is single-speaker mode, only Zhang San will be displayed. If it is multi-speaker mode (maximum of 2 speakers), both Zhang San and Li Si will be displayed.

[0212] The solution provided in this embodiment can prioritize multiple speakers, enabling the system to intelligently select the most worthy subject for display from among multiple speakers, thereby improving the display effect of the screen.

[0213] In some embodiments, the priority of each speaker is determined based on their speaking data and image data, including: Based on each speaker's speaking data, determine each speaker's speaking time percentage, voice volume, and relevance of speaking content to the meeting topic within a preset time period; Based on the image data of each speaker, determine the facial orientation and gaze focus of each speaker; The priority of each speaker is determined based on their speaking time percentage, voice volume, relevance of their speech to the meeting topic, facial orientation, and eye focus. The priority is determined by the following relationships: the higher the speaking time percentage, the higher the priority; the louder the voice volume, the higher the priority; the higher the relevance, the higher the priority; the camera with the face facing forward has a higher priority; the camera with the face facing to the side has a lower priority; the closer the gaze is to the display screen or whiteboard, the higher the priority (or, the smaller the distance between the gaze and the display screen or whiteboard, the higher the priority).

[0214] For example, when calculating priorities, the system considers the following information: A's speaking time accounts for 40%, while B's is 30%; A's volume is louder than B's; A's speech is highly relevant to the meeting topic "project progress," while B's speech deviates from the topic; A is facing the camera directly, while B is turned to the side; A's gaze is focused on the whiteboard, while B is looking at their phone. Based on this information, the system can determine that A's priority is significantly higher than B's.

[0215] The solution provided in this embodiment calculates priorities based on multiple dimensions, making the evaluation of speakers more comprehensive and avoiding misjudgments caused by relying on only a single dimension (such as last speech).

[0216] In some embodiments, obtaining the screen display mode of the tracking area includes: determining the screen display mode as a presentation mode, discussion mode, single speaker mode, or multi-speaker mode based on the gaze direction of the meeting participants, so as to correspondingly determine the output of a single speaker screen, a discussion area screen, or a combined screen containing multiple sub-screens.

[0217] Among these, gaze direction refers to the direction in which participants' eyes are directed, which can be analyzed using eye information captured by cameras. Presentation patterns, discussion patterns, etc., are meeting types inferred from participant behavior.

[0218] For example, if the system detects through the camera that most participants' gazes are focused on the speaker in front, it determines that the current mode is a presentation. The system can then automatically switch the display mode to single-speaker mode. Conversely, if the system detects that participants are exchanging glances, it determines that the current mode is a discussion mode and switches to multi-speaker mode or displays a discussion area.

[0219] The solution provided in this embodiment can automatically infer the meeting mode and set the screen display mode based on the collective behavior (gaze direction) of the participants, which can improve the intelligence level of the system.

[0220] In some embodiments, when the screen display mode is discussion mode, the target meeting screen is the discussion area screen, or a combination of close-up screens of multiple discussion participants; In cases where the target meeting screen is a combination of close-up shots of multiple participants, the method further includes: if the speaking interval of the first participant exceeds a threshold and the current discussion is not yet over, then the combination of close-up shots of all participants is maintained. Here, "the discussion is not over" means that other participants are still speaking, and the entire discussion process has not yet ended.

[0221] For example, in discussion mode, the screen displays a close-up of Zhang San, Li Si, and Wang Wu. After Li Si finishes speaking, even though his speaking interval exceeds a threshold, the system determines that the current discussion is not over because Zhang San and Wang Wu are still discussing. Therefore, Li Si's screen will not be removed, and the three-person group will continue to be displayed.

[0222] The solution provided in this embodiment ensures that, in discussion mode, even if individual members pause speaking, all members' screens remain visible as long as the discussion continues, thus guaranteeing screen stability and the integrity of the discussion atmosphere.

[0223] In some embodiments, the method further includes, before outputting a close-up view of a specific speaker: Semantic analysis is performed on the speech content of a specific speaker to predict the probability of the speaker continuing to speak and the remaining speaking time. If the probability of continuing to speak is lower than a probability threshold, and / or the remaining speaking time is less than a speaking time threshold, the current screen remains unchanged. Semantic analysis refers to using natural language processing technology to analyze the speech content to determine whether the speech is about to end.

[0224] For example, the system is preparing to switch the view to a close-up of Zhang San. However, before switching, the system analyzes Zhang San's speech and finds that he said, "Okay, that's all my opinion, thank you everyone." The system predicts that the probability of him continuing to speak is low, and the remaining time is short. Therefore, the system can maintain the current panoramic view.

[0225] For example, semantic analysis of a specific speaker's speech to predict the probability of the speaker continuing to speak and the remaining speaking time can also be achieved in the following ways: In an optional embodiment, switching suppression is based on closing remarks recognition (explicit signals): Before outputting a close-up shot of a specific speaker, the processing device performs real-time semantic analysis on the speaker's speech content. When the processing device recognizes the presence of preset closing remarks feature words or phrases in the speech content—such as "That concludes my point," "I've finished speaking, thank you everyone," "That's all I have to say," "That's all my opinion"—the processing device, based on these explicit speech end signals, predicts that the probability of the speaker continuing to speak is lower than a probability threshold (e.g., lower than 20%), and the remaining speaking time is less than a speaking time threshold (e.g., less than 5 seconds). In this case, the processing device determines that it is not appropriate to switch the close-up shot and keeps the current shot unchanged. Further, if the processing device is currently outputting a close-up shot of the speaker and recognizes the aforementioned closing remarks feature, the processing device predicts that the speaker is about to end speaking and begins preparing for the subsequent shot rewind process—for example, marking the speaker as "about to end speaking" and preloading a panoramic view of the meeting or close-up shots of other speakers as switching candidates based on whether there are other active sound sources in the meeting room, in order to reduce the delay of subsequent shot switching.

[0226] In an optional embodiment, probability prediction is based on tone and syntactic structure (implicit signals): Before outputting a close-up of a specific speaker, the processing device performs deep semantic analysis on the speaker's speech content, including but not limited to: the downward trend of sentence-end intonation, the shortening trend of sentence length, and the frequency of imperative or summarizing sentences (such as "therefore," "in conclusion," "finally"). Based on these multidimensional features, the processing device constructs a speech end probability model to predict the probability of the speaker continuing to speak. When the predicted probability of continuing to speak is lower than a probability threshold (e.g., lower than 30%), the processing device determines that the speaker's speech is about to end naturally, maintaining the current frame and avoiding switching the close-up shot just before the speaker ends. For example, the speaker begins by using long sentences (e.g., 20-30 words per sentence) with a steady tone; near the end, the sentence length shortens to 5-10 words per sentence, and summarizing conjunctions such as "therefore," "in conclusion," and "finally" appear, with a significant drop in sentence-end intonation. The processing device combines these features to predict that the probability of the speaker continuing to speak is 15%, which is below the threshold, so the current screen remains unchanged.

[0227] In an optional embodiment, a timing-based speech duration prediction and observation period mechanism is used: When the processing device detects a new speaker starting to speak, it does not immediately trigger a screen switch, but instead initiates an initial observation period (e.g., 2 seconds). During the initial observation period, the processing device accumulates the speaker's speaking time and performs semantic analysis on the content of their speech to predict the remaining speaking time. If any of the following conditions are met during the initial observation period, the current screen remains unchanged: ① the predicted remaining speaking time is less than a speaking time threshold (e.g., less than 3 seconds); ② the probability of continuing to speak is lower than a probability threshold; ③ the actual accumulated speaking time is insufficient to confirm that they are a continuous speaker. For example: The processing device detects that speaker B has started speaking and initiates an initial observation period. In the first 1.5 seconds of the observation period, speaker B says, "Excuse me, I'd like to interject, that data just now..." The processing device identifies this brief interjection signal through semantic analysis, and the subsequent content is only a brief confirmation of a data point. The processing device predicts that the probability of them continuing to speak is only 10%, and the remaining speaking time is approximately 2 seconds. Therefore, it determines that the screen switch conditions are not met, and the current screen (close-up of speaker A) remains unchanged. If speaker B finishes saying, "I need to confirm that data again—actually, according to our test results last week…", and continues to elaborate with a complete argumentative sentence structure, the processing device updates its prediction—the probability of continuing to speak increases to 70%, the remaining speaking time increases to approximately 30 seconds, and then the screen switches to a close-up of speaker B.

[0228] In an optional embodiment, personalized prediction based on historical speaking patterns is employed: the processing device maintains historical speaking data for each speaker, including the average duration of each speaker's previous speeches, the shortest effective speaking time, and typical language patterns before the end of their speech. When predicting the probability of a specific speaker continuing to speak and the remaining speaking time, the processing device adjusts its prediction based on that speaker's historical speaking patterns. For example, for two different speakers: Speaker A's historical data indicates that their speaking habit is "long paragraph statements," with an average speaking time of 45 seconds, and they typically use summarizing phrases such as "in conclusion" before ending their speech. When the processing device detects that Speaker A has spoken for 35 seconds in the current session and has not yet used a summarizing phrase, it predicts a 70% probability of them continuing to speak and approximately 10 seconds of remaining speaking time. Speaker B's historical data indicates that their speaking habit is "short interjections," with an average speaking time of only 3 seconds, and they almost never use summarizing phrases. When the processing device detects that Speaker B has spoken for 2 seconds and their sentence structure is a simple question, it predicts a 5% probability of them continuing to speak and less than 1 second of remaining speaking time, therefore no screen switching is triggered. Furthermore, the processing device can aid in prediction based on the deviation between a speaker's historical speaking patterns and the real-time characteristics of their current speech. If a speaker's current speech length significantly exceeds their historical average speaking duration (e.g., more than twice the standard deviation), the processing device predicts a higher probability that they are about to finish speaking, reducing the urgency of the screen transition. This helps avoid the inefficient switching where a close-up shot is cut to the final stage of a speaker's lengthy speech, followed immediately by a cut-out.

[0229] In an optional embodiment, prediction based on the completeness of the speech content is performed: the processing device performs semantic completeness analysis on the speaker's speech content to determine whether the current statement is a complete semantic unit. When the processing device recognizes that the speaker has completed a complete semantic unit (such as completing the exposition of an argument, answering a question, or stating a data conclusion), and no new semantic unit is initiated subsequently, the probability of the speaker continuing to speak is predicted to decrease. For example: the speaker is giving a three-point report: "The first point is about the progress of project A... (elaboration); the second point is about the budget of project B... (elaboration); the third point is about the next steps..." When the processing device recognizes that the speaker has completed the discussion of "the third point", and a pause occurs (the pause duration is less than the preset interval threshold T1, which is not enough to determine the end of the speech), based on the semantic analysis, it is determined that the speaker has completed the exposition of all the key points in his speech outline, and the probability of continuing to speak is predicted to drop below 20%, with the remaining speaking time less than 3 seconds. At this point, even if the speaker has a brief concluding statement (such as "Okay, that's all I wanted to report"), the processing device determines that it is not advisable to switch to a close-up of the speaker, because the speaker will immediately end their speech after the screen switches, resulting in an invalid switch. If, after completing the presentation of the three key points, the speaker pauses briefly and then begins a new discussion (such as "However, regarding item C, I would like to add something..."), the processing device updates its prediction—the probability of continuing to speak rises back to over 60%, and the remaining speaking time is reset.

[0230] The solution provided in this embodiment can predict before the screen switch based on semantic analysis, which can effectively avoid the scenario of "the speaker ends as soon as the screen switches" and improve the effectiveness of the switch.

[0231] In some embodiments, the preset interval threshold, the first duration threshold, and the second duration threshold are determined separately for different tracking regions, and / or the preset interval threshold, the first duration threshold, and the second duration threshold can be adaptively and dynamically adjusted based on historical speech data of the corresponding tracking regions. Here, "determined separately" means that different tracking regions can use different threshold parameters. "Adaptively and dynamically adjusted" means that the threshold parameters can be automatically optimized based on historical speech data.

[0232] For example, in key seating areas, where speeches are typically longer and more formal, the system sets higher T0 (e.g., 5 seconds) and T1 (e.g., 3 seconds). In contrast, in open discussion areas, where speeches are usually short and frequent, the system sets lower T0 (e.g., 2 seconds) and T1 (e.g., 1 second). The system also automatically adjusts these thresholds based on historical data.

[0233] For example, in an optional embodiment, thresholds based on region attributes are configured separately: a preset interval threshold (T1), a first duration threshold (T0_region), and a second duration threshold (T0_individual) are determined according to different tracking regions. Specifically, before the meeting begins or during system initialization, threshold parameters are configured independently for each tracking region based on its physical and functional attributes. For example: The meeting room is divided into three tracking areas: the podium area, the core discussion area, and the audience area. For the podium area (typically the fixed location where speakers deliver formal reports), a high threshold is configured: first duration threshold T0_area = 8 seconds, second duration threshold T0_individual = 5 seconds, and preset interval threshold T1 = 4 seconds. Because speeches from the podium are usually continuous long reports, the high threshold filters out brief confirmatory questions and audience interaction, ensuring that the screen only switches when there is indeed continuous speaking from the podium.

[0234] For the core discussion area (the area where participants discuss around a roundtable), a medium threshold is configured: first duration threshold T0_area = 4 seconds, second duration threshold T0_individual = 3 seconds, and preset interval threshold T1 = 2 seconds. Because the discussion area is characterized by multiple participants taking turns speaking with short intervals, the medium threshold can respond promptly to changes in the discussion pace while filtering out extremely brief echoing responses.

[0235] For the observation area (the area where the main audience is located), configure lower thresholds: first duration threshold T0_area = 3 seconds, second duration threshold T0_individual = 2 seconds, and preset interval threshold T1 = 3 seconds. Since observers typically do not speak voluntarily, any sound source they hear may indicate a question or supplementary comment requiring attention; lower thresholds facilitate a quicker response.

[0236] Once the threshold is configured, it is stored by the processing device in the configuration file corresponding to each tracking area, and read and used by area during the meeting.

[0237] In an optional embodiment, threshold differentiation configuration based on role tags is used: During the initialization phase, the processing device acquires the role tags (such as "speaker," "host," "participant," "observer," etc.) of attendees in each tracking area, and configures differentiated threshold parameters for each tracking area based on the role tags. Role tags can be obtained through methods such as importing meeting schedules, administrator presets, or image recognition (such as seat nameplate recognition). Specific examples: The tracking area marked as "Speaker" has the following parameters: First duration threshold T0_area = 8 seconds, Second duration threshold T0_individual = 6 seconds, and Preset interval threshold T1 = 5 seconds. The speaker's speech is usually a formal statement with relatively long pauses between sentences (such as turning pages or waiting for audience response). A higher T1 value can prevent the speaker's pauses in thought from being mistaken for the end of the speech.

[0238] The tracking area is marked as "Moderator": First duration threshold T0_area = 3 seconds, second duration threshold T0_individual = 2 seconds, preset interval threshold T1 = 1.5 seconds. The moderator's speech is characterized by being short and frequent (such as guiding discussions and calling on students to ask questions), and a lower threshold helps to respond quickly to the moderator's activities.

[0239] Tracking areas marked "participants": Use the default generic threshold.

[0240] In an optional embodiment, the threshold is dynamically adjusted adaptively based on historical data: The preset interval threshold (T1), the first duration threshold (T0_region), and the second duration threshold (T0_individual) can be adaptively and dynamically adjusted based on the historical speaking data of the corresponding tracking region. The processing device maintains a historical speaking database for each tracking region, recording statistical data including the start time, end time, speaking duration, and speaking interval duration for each speech, and updates the threshold parameters periodically (e.g., every 10 minutes or after each speaking round). The specific adjustment method is as follows: Step 1: Data Acquisition. Within a time window W (e.g., the past 30 minutes), statistically analyze the duration and interval time series of all speaking events in the tracking area.

[0241] Step 2: Adjusting the Interval Threshold T1. Perform cluster analysis on the interval time series to identify two types of intervals: ① natural pauses within sentences (typically 0.5-2 seconds); ② the actual interval between speaking rounds (typically 5 seconds or more). Use the boundary between these two types of intervals as the new T1 value. For example, if historical data shows that speakers in the tracked area typically pause within sentences for less than 1.5 seconds, while the interval between speaking rounds is typically more than 6 seconds, then T1 can be adjusted from the initial value of 3 seconds to 4 seconds to better suit the speaking rhythm of the area.

[0242] Step 3: Adjusting the duration thresholds T0_region and T0_individual. Perform statistical analysis on the speech duration sequence to calculate the average speech duration μ and standard deviation σ. Set the second duration threshold T0_individual to μ - 0.5σ (but not lower than the preset lower limit of 2 seconds) to ensure that only speeches exceeding this threshold are considered "continuous speech". Set the first duration threshold T0_region to 1.2 to 1.5 times that of T0_individual, and ensure that T0_region ≥ T0_individual.

[0243] Step 4: Boundary Protection. Set preset upper and lower limits for each threshold to prevent adaptive adjustments from causing the thresholds to exceed reasonable ranges. For example, set the upper limit for region T0 to 10 seconds and the lower limit to 2 seconds; set the upper limit for T1 to 8 seconds and the lower limit to 1 second.

[0244] In an optional embodiment, thresholds are dynamically switched based on real-time meeting phase adaptation: The processing device dynamically switches the threshold configuration of each tracking area based on different stages of the meeting process. Meeting stage identification can be achieved through one or more of the following methods: ① based on the meeting schedule; ② based on pattern recognition based on the number of current speakers and speaking frequency; ③ based on meeting modes manually switched by the user (such as "presentation mode," "discussion mode," and "Q&A mode"). Specific examples: In speech mode (single speaker speaking for a long time), all tracking areas are uniformly configured with high thresholds: T0_area = 8 seconds, T0_individual = 5 seconds, T1 = 4 seconds, to avoid unnecessary screen switching triggered by short coughs, turning pages, or other sounds from the audience.

[0245] In discussion mode (where multiple people frequently take turns speaking), the core discussion area switches to a low-threshold configuration: T0_area = 3 seconds, T0_individual = 2 seconds, T1 = 1.5 seconds, in order to quickly respond to speaker switching during the discussion; while the observer area maintains a high-threshold configuration.

[0246] In Q&A mode, the threshold for the audience seating area is temporarily lowered: T0_area = 2 seconds, T0_individual = 1.5 seconds, T1 = 1 second, in order to quickly capture audience questions and comments; while the threshold for the podium area is raised to prevent short pauses during the Q&A session from being mistakenly interpreted as the end of the speech.

[0247] In an optional embodiment, online threshold fine-tuning is based on real-time feedback: In addition to adaptive threshold adjustment, the processing device further introduces a real-time feedback mechanism. Online threshold fine-tuning is triggered when the processing device detects the following events: Trigger Event A – Excessive Invalid Switching Frequency: If, within a preset monitoring time window (e.g., 5 minutes), the number of screen switching times in the tracked area exceeds a preset frequency threshold (e.g., more than 3 times per minute), and the proportion of speakers immediately ending their speeches after a switch exceeds 50%, then the current threshold is determined to be too low, leading to oversensitivity. The processing device automatically increases T0_area and T0_person by 0.5 seconds, and increases T1 by 0.3 seconds, until the invalid switching frequency drops to an acceptable range.

[0248] Triggering Event B – Excessive Response Delay: If a speaker has been speaking for more than 1.5 times the duration of the T0_ area without triggering a screen switch, and the speaker's location status is "located to specific speaker," then the current threshold is determined to be too high, causing a response delay. The processing device automatically lowers the T0_ area and T0_ individual thresholds by 0.5 seconds to speed up the response to continuous speaking.

[0249] Trigger Event C – Change in Positioning Accuracy: When the positioning accuracy of the sound source in the tracking area changes continuously (e.g., the positioning accuracy drops from Level 1 to Level 2 for more than 2 minutes due to increased noise in the venue), the processing device automatically lowers the second duration threshold T0_person for that area (by 0.5 seconds) to reduce the impact of speech detection delay caused by the decrease in positioning accuracy. When the positioning accuracy recovers, the threshold automatically returns to its original configuration.

[0250] Furthermore, the online fine-tuning employs a gradual adjustment strategy, with each adjustment step not exceeding 0.5 seconds and a minimum interval of 30 seconds between adjacent adjustments to avoid drastic fluctuations in the threshold within a short period. Each adjustment is recorded in a log for subsequent manual review.

[0251] In an optional embodiment, personalized thresholds are based on the speaker's personal profile: the processing device maintains an independent personal profile for each speaker, recording the speaker's historical speaking characteristics, including average speaking duration, speaking interval distribution, and pause patterns. When a speaker is first identified by the system, the default threshold for their tracking area is used; after the system accumulates sufficient historical data (e.g., more than 5 speaking rounds), personalized threshold parameters are generated for the speaker, and the personal threshold is used preferentially over the area threshold when the speaker speaks. For example, the default T0_personal threshold within the tracking area is 3 seconds, and T1 is 2 seconds. However, the historical data of speaker X shows that their speaking style is "short and frequent"—the average speaking duration is only 2.5 seconds, but the interval between two speaking sessions is usually no more than 1 second. If the default threshold is used, most of speaker X's speaking sessions will not trigger a screen switch (because the average speaking duration of 2.5 seconds < 3 seconds). Based on speaker X's personal profile, the processing device adjusts their personalized thresholds to T0_personal = 1.5 seconds and T1 = 1 second. This ensures that their continuous speaking behavior can be captured promptly, while T1 = 1 second still filters out normal pauses between sentences (approximately 0.5 seconds). When speaker X leaves the current tracking area and enters another tracking area, their personalized threshold is transferred to the new area and cross-compared with the area-level threshold of the new area—the lower of the two is taken as the actual threshold used to ensure a reasonable response to different speaking styles.

[0252] The solution provided in this embodiment realizes the personalization and adaptation of the threshold parameter, enabling the system to better adapt to the speaking characteristics of different regions and improve the system's generalization ability and robustness.

[0253] In some embodiments, determining the localization state corresponding to the sound source includes: Within the tracking area, at least one candidate object is identified, wherein at least one candidate object is a candidate speaker or a candidate speaking region; the correlation degree between each candidate object and the sound source is determined; based on the correlation degree between each candidate object and the sound source, the localization state of the sound source is determined; wherein the correlation degree is determined by at least one of the following: the degree of positional matching between the candidate object and the sound source; the degree of consistency between the lip movement corresponding to the candidate object and the current speech activity of the sound source; the degree of matching between the posture corresponding to the candidate object and the current speaking state of the sound source; the continuity between the candidate object and the localization result of the sound source at the previous moment. The correlation degree is a comprehensive score used to measure the degree of matching between the candidate object and the current sound source. The candidate object can be a potential speaker or speaking region.

[0254] For example, the system detects a sound source in the left front region. Candidates A and B are located within this region. The system calculates that A has a high degree of positional match with the sound source (A is in the direction of the sound source), and A's lip movements are consistent with the speech activity of the sound source, while B's lip movements are inconsistent. Therefore, A has a high correlation with the sound source, and the system determines the localization status as having located the specific speaker A.

[0255] For example, in an optional embodiment, the correlation degree is calculated based on the degree of position matching: the processing device determines at least one candidate object (candidate speaker or candidate speaking area) within the tracking area and calculates the degree of position matching between each candidate object and the sound source as a factor of the correlation degree. Specifically, the processing device obtains the direction of arrival (DOA) angle of the sound source output by the microphone array, and simultaneously obtains the spatial coordinates of each candidate object in the image determined by the camera through facial recognition. The processing device maps the DOA angle to the camera image coordinate system and calculates the deviation between the spatial position of each candidate object and the DOA angle. The degree of position matching is determined as follows: if the deviation between the spatial position of the candidate speaker and the DOA angle is less than a first angle threshold (e.g., ±3°), the degree of position matching is high (assigned a value of 0.9); if the deviation is between the first angle threshold and a second angle threshold (e.g., ±3°~±8°), the degree of position matching is medium (assigned a value of 0.5); if the deviation is greater than the second angle threshold (e.g., ±8°), the degree of position matching is low (assigned a value of 0.1). Furthermore, the processing device combines the matching of the candidate speaking area (the entire spatial range of the area) with the DOA angle: if the DOA angle points into the interior of the candidate speaking area, the position matching degree at the region level is high; if the DOA angle points to the edge range of the candidate speaking area (such as a buffer zone of ±2° of the region boundary), the position matching degree at the region level is medium; if the DOA angle is completely deviated from the candidate speaking area, the position matching degree at the region level is zero.

[0256] In an optional embodiment, the correlation between lip movement and speech consistency is calculated: the processing device captures a sequence of lip movement images of each candidate using a camera, and simultaneously acquires the audio signal of the current speech activity using a microphone. The processing device performs lip movement detection on the lip movement image sequence, extracts lip movement features (such as lip opening and closing frequency, and mouth opening amplitude), and performs time alignment and correlation analysis with the speech activity features of the audio signal (such as speech energy envelope and fundamental frequency variation). Specifically, the processing device calculates the cross-correlation function between lip movement features and speech activity features within a time window T (e.g., 500ms). When the peak value of the cross-correlation function is greater than a preset lip-reading consistency threshold (e.g., 0.7), and the time delay is within a reasonable range (e.g., lip movement leading speech by no more than 100ms), the candidate's lip movement is determined to be consistent with the speech activity of the sound source, and the candidate is considered to have a high confidence level as the current speaker. For example, the lip movement frequency of candidate A is highly correlated with the syllable rhythm of the speech signal (cross-correlation peak value 0.85), while candidate B has no obvious lip movement during the same time period (cross-correlation peak value 0.15). The processing device determines that candidate A has a much higher correlation with the sound source than candidate B, and based on this, determines the localization status as "localized to specific speaker A". If the lip movements of multiple candidate objects are consistent with the speech activity to a certain extent (such as in a scenario where multiple people are speaking at the same time), the processing device combines sound source separation technology to decompose the audio stream into multiple independent speech audio streams, and then performs a one-to-one matching with the lip movement features of each candidate object, selecting the candidate object with the highest matching degree as the localization result.

[0257] In an optional embodiment, the correlation between posture and speaking state is calculated: the processing device acquires posture information of each candidate through a camera, including but not limited to: body orientation, head orientation, hand gestures, whether holding a microphone or flipping through documents, etc. The processing device matches and analyzes the above posture information with the speaking state of the current sound source. The matching rules are as follows: if the candidate's body orientation and head orientation are both facing the center of the meeting (such as the conference table or projection screen), and there are accompanying actions such as holding a microphone or flipping through documents, then the candidate is determined to be in an "active speaking preparation" state, and the degree of matching between posture and speaking state is high. If the candidate leans back, lowers their head, or faces the laptop screen, then the candidate is determined to be in a "non-speaking" state, and the degree of matching between posture and speaking state is low. For example: in a conference room, candidate A leans forward, faces the center of the conference table, and holds a page turner; candidate B leans back in their chair and focuses their gaze on a laptop. After detecting the sound source, the processing device, combined with the posture information, determines that the degree of matching between the posture and speaking state of candidate A is significantly higher than that of candidate B. Therefore, candidate A has a higher correlation with the sound source. Furthermore, the processing device performs weighted fusion of pose matching, position matching, and lip movement matching, with the weights preset by the system or dynamically adjusted according to the scene. For example, under conditions of sufficient light and clear faces, the weight of lip movement matching is set to 0.5, the weight of position matching is set to 0.3, and the weight of pose matching is set to 0.2; under conditions of dim light and difficulty in capturing lip details, the weight of lip movement matching decreases to 0.2, the weight of position matching increases to 0.5, and the weight of pose matching increases to 0.3.

[0258] In an optional embodiment, the correlation degree is calculated based on the continuity of the positioning results: the processing device compares the current position of each candidate object with the previous position of the sound source, and calculates a continuity score as an evaluation factor of the correlation degree. The continuity score follows these principles: if the candidate object is the same person as the speaker located at the previous position, and the displacement between the current position of the candidate object and the previous position of the speaker is within a reasonable range (e.g., movement distance ≤ 2 meters), the continuity score is high; if the candidate object is the speaker at the previous position but has moved across regions, the continuity score is medium; if the candidate object is different from the speaker at the previous position, the continuity score is low. For example: the positioning result at the previous position was "located to specific speaker A," and speaker A is located in tracking area Z1. At the current position, the sound source is still located near tracking area Z1. The processing device determines that the candidate objects within tracking area Z1 include speaker A and speaker B. Since candidate A's location result is consistent with the previous moment (same speaker, continuous location), its continuity score is 0.9; while candidate B's location result is inconsistent with the previous moment, its continuity score is 0.1. After combining other factors such as location matching, lip movement consistency, and posture matching, speaker A's overall correlation is significantly higher than speaker B's, and the processing device maintains the location status as "located to specific speaker A". If speaker A was located in tracking area Z1 in the previous moment, but the current sound source appears in tracking area Z2, the processing device determines the candidate in Z2. At this time, since speaker A has moved across areas (from Z1 to Z2), if speaker A is identified as a candidate in Z2, its continuity score is medium (0.5), allowing the location result to maintain the tracking continuity of the same speaker between different tracking areas, avoiding the speaker's identity being re-identified due to the speaker moving around in the venue.

[0259] In an optional embodiment, the comprehensive determination method for multimodal correlation is as follows: the processing device comprehensively considers the degree of position matching, the degree of lip movement consistency, the degree of posture matching, and the continuity score to calculate the comprehensive correlation between each candidate and the sound source. The comprehensive correlation is calculated using a weighted summation method: Correlation = α × Position Matching Degree + β × Lip-Speech Consistency + γ × Posture Matching Degree + δ × Continuity Score, where α, β, γ, and δ are weight coefficients, satisfying α + β + γ + δ = 1. The default configuration of the weight coefficients is α = 0.3, β = 0.3, γ = 0.2, and δ = 0.2. The processing device can dynamically adjust the weight coefficients according to the current meeting environment and detection conditions. The triggering conditions for dynamic adjustment include: Location environment detection: When the ambient noise is higher than the preset noise threshold, the processing device determines that the sound source positioning accuracy may decrease, reduces the weight α of the position matching degree (e.g., from 0.3 to 0.15), and increases the weight β of lip-reading consistency (e.g., from 0.3 to 0.45), because lip visual information is not affected by ambient noise.

[0260] Lighting condition detection: When the ambient light intensity is lower than the preset light intensity threshold, the processing device determines that the reliability of visual detection (lip movement, posture) is reduced, and the weights of lip consistency β and posture matching degree γ are reduced by 50% each. The released weights are then allocated to position matching degree α and continuity score δ.

[0261] Speaking phase detection: In the initial speaking phase (cumulative speaking time < 3 seconds), increase the weight δ of continuity score (e.g., from 0.2 to 0.4) to utilize historical information to assist in quick speaking attribution judgment; in the continuous speaking phase (cumulative speaking time ≥ 5 seconds), increase the weight of lip-reading consistency and position matching to obtain a more accurate positioning status.

[0262] The processing device compares the overall correlation with a preset correlation threshold: if the overall correlation is ≥ 0.7, the positioning status is "the specific speaker has been located"; if 0.4 ≤ overall correlation < 0.7, the positioning status is "the specific speaker has not been located but the speaking area has been located"; if the overall correlation is < 0.4, the positioning status is "neither the specific speaker nor the speaking area has been located".

[0263] In an optional embodiment, the multi-candidate correlation degree is determined based on a weighted fusion and elimination mechanism: when there are multiple candidate objects in the tracking area, the processing device uses a "multi-round elimination-weighted fusion" method to determine the final correlation degree.

[0264] The first round – coarse screening and elimination: The processing device first calculates the degree of positional matching between each candidate object and the sound source, retaining candidates whose positional deviation is within a second angle threshold (e.g., ±8°), and eliminating candidates with excessive positional deviation. If only one candidate object remains after the first round of screening, the positioning status is determined directly based on the continuous correlation evaluation of that candidate object.

[0265] The second round—lip-reading verification—calculates the consistency between lip movements and speech activity for the candidates retained from the first round. Candidates with significantly low lip-reading consistency (e.g., cross-correlation peak <0.3) are eliminated. If only one candidate remains after the second round of screening, the localization status is directly determined.

[0266] The third round—comprehensive weighting: If multiple candidate objects still exist after the first two rounds of screening, the comprehensive correlation degree of each candidate object is calculated according to the weighted fusion method in Example 7E, and the candidate object with the highest score is used as the positioning result. If multiple candidate objects have similar comprehensive correlation degrees (difference < 0.1), then the historical information in the continuity score is combined—prioritizing the candidate object that is continuous with the positioning result of the previous moment, in order to maintain the stability of the positioning state.

[0267] The beneficial effects of this embodiment are as follows: by using multiple rounds of hierarchical screening, obviously mismatched candidates are eliminated first, which reduces the computational overhead of comprehensive weighted fusion and also improves the reliability of the positioning results. This is because each round of screening uses independent information from different dimensions, and misjudgment in a single dimension will not directly affect the final positioning results.

[0268] The solution provided in this embodiment can combine multimodal information such as position, lip movements, posture, and historical information to calculate the correlation degree in order to determine the positioning status, thereby improving the accuracy and robustness of positioning.

[0269] In some embodiments, the location status of the sound source is determined based on the correlation between each candidate object and the sound source, including: if the correlation corresponding to the candidate speaker is greater than or equal to the correlation threshold, the location status is determined to be that the specific speaker has been located; if the correlation corresponding to the candidate speaker is less than the correlation threshold, and the correlation corresponding to the candidate speaking area is greater than or equal to the correlation threshold, the location status is determined to be that the specific speaker has not been located but the speaking area has been located; if the correlation corresponding to the candidate speaker is less than the correlation threshold, and the correlation corresponding to the candidate speaking area is less than the correlation threshold, the location status is determined to be that neither the specific speaker nor the speaking area has been located.

[0270] For example, the relevance threshold is set to 0.8. If the relevance of candidate speaker Zhang San is 0.9, then the specific speaker is located. If Zhang San's relevance is 0.5, but the relevance of the candidate speaking area in the left front zone is 0.85, then the speaking area is located. If both relevances are below 0.8, then no location can be made.

[0271] The solution provided in this embodiment compares the correlation degree with the threshold, providing clear boundaries for the three positioning states, making the determination of positioning results clearer and more reliable.

[0272] The following is an introduction in conjunction with the scenario: Scenario A: In this scenario, the system switches between panoramic and preset position views, without involving close-up views of specific speakers or speaking areas.

[0273] As one possible implementation, when a valid sound source is detected and the cumulative speaking time in the tracking area where the sound source is located is ≥ T0, the processing device can trigger the activation of the preset position of the gimbal camera. When the preset position of the gimbal camera is activated, that is, when the gimbal camera rotates and completes the selection of the preset position image of the tracking area, the gimbal camera reports the activation information to the processing device. The processing device digitally zooms the selected preset position image and controls the display device to display the digitally zoomed preset position image. Thus, the processing device switches the panoramic image to the preset position image.

[0274] When all sound sources have finished speaking and the interval is ≥T1, it means that the speaking round has ended and is not a short pause. In this case, the processing device can switch the display screen from the preset position screen back to the panoramic screen.

[0275] Scenario B: In this scenario, the system can switch between three progressively larger views: panoramic view → preset position view → close-up view.

[0276] Level 1: Panoramic View When no one speaks or all speakers have finished their speaking rounds (interval duration ≥ T1), the processing device controls the display of the panoramic view.

[0277] Level 2A: Panoramic View → Preset Position View If a valid sound source exists and the cumulative speaking time in the tracking area where the sound source is located is ≥ T0, the processing device can trigger the preset position call of the pan-tilt camera in that tracking area.

[0278] When the preset position of the gimbal camera is activated, the camera reports the position information to the processing device. At this time, if the processing device has not yet obtained the position of the speaker or speaking area, Avhub electronically crops (digitally zooms) the preset position image selected by the gimbal camera and controls the display of the processed preset position image.

[0279] Level 2B: Panoramic view → Close-up view, or Panoramic view → Preset position view In some scenarios, if there is a valid sound source and the cumulative speaking time in the tracking area where the sound source is located is ≥ T0, the processing device can trigger the preset position call of the PTZ camera in that tracking area.

[0280] When the preset position of the gimbal camera is activated, the activation information is reported to Avhub.

[0281] In some examples, if the processing device has obtained the speaker's location, Avhub performs optical zoom on the preset position framed by the gimbal camera and controls the display of the processed close-up image of the speaker (optical close-up image).

[0282] In some examples, if the processing device has obtained the location of the speaking area, Avhub performs optical zoom or digital zoom on the preset position framed by the gimbal camera and controls the display of the processed close-up image of the speaking area (optical close-up image or digital close-up (electronic cropping) image).

[0283] In some examples, if the processing device fails to obtain the speaker's location and the location of the speaking area when the preset position of the gimbal camera is activated, the processing device can control the switching from the panoramic view to the preset position view.

[0284] Subsequently, if a close-up of a person / area is located shortly after the preset position screen is displayed, a fade-in / fade-out or jump transition can be used to switch the close-up screen to avoid a rapid transition from the preset position screen to the close-up screen. In addition, there should be at least a certain interval between the display of the preset position screen and the display of the close-up screen. This interval is configurable, for example, a value range of 0s-5s.

[0285] Level 3: Preset screen → Close-up screen In some scenarios, when a preset position screen is already displayed, if the positioning accuracy reaches the level of precisely locating a specific speaker (accuracy level 1) and the speaker's cumulative speaking time is ≥ T0, the processing device will switch the screen to a close-up view of that speaker. Alternatively, if the approximate speaking area can be located (accuracy level 2) and the cumulative speaking time of that area is ≥ T0, the processing device will switch the screen to a view of that speaking area.

[0286] For example, the processing device performs optical close-up on the preset position image, non-electronic cropping, and outputs a close-up image of the person / area.

[0287] In some scenarios, when a preset screen is already displayed, if the cumulative speaking time of a specific speaker / speaker area has not reached the target, the preset screen remains unchanged until the specific speaker's speaking time reaches the target, at which point it switches to a close-up screen. This ensures that the close-up screen outputs to the person who is actually speaking continuously, rather than a brief interruptor.

[0288] Scene C: Panorama → Panorama + Multiple Views (Combined Views) In this scenario, both panoramic and multiple close-up views of people / areas can be displayed simultaneously.

[0289] For example, such as Figure 4 The combined view may include a panoramic view (Room View) and multiple speaker views / speaking area views; there is no limit to the layout style and number of speaker / speaking area views.

[0290] Exemplarily, when no one is speaking, only the panoramic view can be displayed. When someone is speaking and the positioning is effective, the panoramic view and multiple close-up person / close-up area views can be displayed.

[0291] In this scenario, the processing device can also trigger the call of the preset position of the camera, but does not display the preset position picture of the camera. The preset position picture is used to frame the close-up person / close-up area. In some examples, if the preset position of the camera has been called and is in place, but the specific speaker / speaking area cannot be located, the panoramic view can be displayed. For example, the approximate direction of the sound source can be marked in the panoramic view, but the sub-view of the sound source is not displayed.

[0292] As follows, the maintenance and switching of the close-up person picture involved in the present disclosure will be introduced by scenarios: In some scenarios, the current close-up person speaks intermittently and the position remains unchanged.

[0293] For example, if the speaking interval duration of the close-up person < T1: Keep the current close-up person picture unchanged. At this time, even if the speaker has a short pause, it will not go back, ensuring the continuity of the picture.

[0294] For another example, if the speaking interval duration of the close-up person ≥ T1: Determine that the speaking round has ended and there is no other sound source, and the picture is switched back to the panoramic or preset position picture.

[0295] In some scenarios, the position of the current close-up person changes, but the speech has not ended.

[0296] Judge whether the current position of the speaker exceeds the current tracking area. If it does not exceed the current tracking area, the camera continuously tracks the speaker and outputs the corresponding close-up picture. If it exceeds the current tracking area, it means that the camera in the current tracking area can no longer track the speaker. First, switch the original close-up picture to the preset position picture corresponding to the new tracking area where the speaker is located for a smooth transition. Then, use the camera in the new tracking area to re-locate the speaker. Similarly, when the conditions for picture switching are met, switch the preset position picture corresponding to the new tracking area to the corresponding close-up picture, such as a close-up person / close-up area.

[0297] In some scenarios, the positioning accuracy of the current close-up person degrades from level one to level two.

[0298] For example, if the close-up person changes position within the current tracking area and the camera in the current tracking area cannot locate the specific position of the close-up person after the change, but only the general speaking area can be located, then switch from the original close-up person picture to the close-up speaking area picture.

[0299] In some scenarios, during the speech of the current close-up person, a new speaker is detected. This scenario can include the following sub-scenarios: Sub-scenario A (Priority of Speech Ratio): Comprehensively consider the speech situations of the current close-up person and the new speaker. Finally, frame and display the person with the highest speech ratio who is still speaking currently.

[0300] If the old sound source continues to speak and a new speaker / new speech area is detected, then within a certain period of time since the detection of the new sound source, the speech duration ratio corresponding to each speaker / speech area can be counted. When the ratio of a certain speaker / speech area exceeds 50%, re-frame the new speaker / area. When the ratio of the new speaker / area exceeds 75% and is currently in the speaking state, display the picture of the new speaker / area.

[0301] Sub-scenario B: The current close-up person is still speaking and the speech duration of the new speaker ≥ T0: If the new speaker and the current close-up person are in the same tracking area: The picture containing both people can be switched for display. For example, the preset position picture or the combined picture. The combined picture contains the close-up views of the two people respectively.

[0302] If the new speaker and the current close-up person are in different tracking areas: Display the combined picture. Or, first switch to the panoramic picture and then switch to the combined picture to reduce the abruptness of the direct jump.

[0303] Sub-scenario C: The current close-up person is still speaking and the speech duration of the new speaker < T0: The picture remains unchanged, still being the current close-up person, to avoid triggering picture switching due to short interruptions.

[0304] Sub-scenario D: During the speech of the new speaker, the current close-up person has finished speaking and the interval ≥ T1: The picture first returns to the panoramic picture or the preset position picture of the tracking area where the new speaker is located. After the continuous cumulative speech duration of the new speaker ≥ T0, it is switched to the close-up picture of the new speaker. In this way, compared with directly jumping from the old close-up to the new close-up, the two-step transition of first retreating and then switching to the new close-up can make the picture switching smoother.

[0305] In some scenarios, the current close-up person interrupts the speech for more than T1 and no new sound source is continuously detected. In this case, the picture retreats from the close-up person picture to the panoramic picture or the camera preset position picture corresponding to the tracking area where the last sound source was located.

[0306] As follows, the maintenance and switching of the speech area pictures involved in the present disclosure are introduced: In some scenarios, a specific speaker / new speech area is detected and its speech duration ≥ T0: If the current close-up area picture contains this specific speaker / new speech area: The picture remains unchanged.

[0307] If the current close-up area picture does not contain this specific speaker / new speech area: Switch to the close-up picture of this speaker / switch to the picture of the new speech area.

[0308] In some scenarios, if the current close-up area is interrupted for more than T1 and no new sound source is detected, the view will revert from the close-up area to the panoramic view or the camera preset position view corresponding to the tracking area where the last sound source was located.

[0309] This application also provides a conference video output system, including a processing device, a microphone, and a camera; The processing device is used to locate the sound source through a microphone when a sound source is detected, and to determine the tracking area where the sound source is located and the corresponding positioning status of the sound source. The positioning status includes one of the following: the specific speaker is located, the specific speaker is not located but the speaking area is located, and neither the specific speaker nor the speaking area is located. Processing device, used to acquire speech data corresponding to the sound source; Processing device for acquiring the screen display mode configured in the tracking area; The processing unit is used to control the camera to output the target conference image based on the positioning status, speaking data, and screen display mode.

[0310] It should be noted that the processing device, microphone, and camera in the conference screen output system of this embodiment can also execute other methods in the above embodiments to achieve the same effect, which will not be elaborated here.

[0311] This application also provides a computer-readable storage medium storing a computer program or instructions thereon, which, when executed by a processor, implements a conference screen output method as described in any of the present disclosures.

[0312] This application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement a conference screen output method as described in any of the present disclosures.

[0313] The above primarily describes the solutions provided by the embodiments of this application from a methodological perspective. It is understood that, in order to achieve the above functions, the electronic device includes hardware structures and / or software modules corresponding to the execution of each function. Based on the units and algorithm steps of the various examples described in the embodiments disclosed in this application, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by a computer driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solutions of the embodiments of this application.

[0314] This application embodiment can divide the electronic device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional module. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0315] For example, Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown.

[0316] like Figure 5 As shown, electronic device 500 may include processor 510.

[0317] Optionally, the electronic device 500 may also include a memory 520.

[0318] Optionally, the electronic device 500 may also include a camera 530. Optionally, the electronic device 500 may also include a microphone array. Optionally, the electronic device 500 may also include a speaker. Thus, by integrating audio and video modules, the electronic device 500 can achieve an "audio-video integrated" processing effect.

[0319] For example, when the electronic device 500 is a standalone host, it may not include the camera 530, which is set up separately. When the electronic device 500 is an integrated device, it may integrate the camera 530.

[0320] Processor 510 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0321] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0322] The processor 510 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 510 is a cache memory. This memory can store instructions or data that the processor 510 has just used or that are used repeatedly. If the processor 510 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 510, and thus improves the efficiency of the system.

[0323] In some embodiments, processor 510 may include one or more interfaces. These one or more interfaces may be used to connect processor 510 to memory 520, etc.

[0324] The memory 520 can be used to store computer executable program code, which includes instructions. The memory 520 may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc. The data storage area may store data created during the use of the electronic device 500, etc. The processor 510 executes various functional applications and data processing of the electronic device 500 by running instructions stored in the memory 520 and / or instructions stored in memory disposed within the processor.

[0325] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include... Figure 5 The diagram shows more or fewer components, or combinations of components, or separate components, or different arrangements of components. The components shown can be implemented in hardware, software, or a combination of both.

[0326] like Figure 6 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. This electronic device 2200 can be used to implement the methods described in the above method embodiments. For example, the electronic device 2200 may specifically include a processing unit 2201. The processing unit 2201 is used to support the electronic device 2200 in performing operations. Figure 1 to Figure 5 The processing function described in any one of the following.

[0327] Optional, Figure 6 The electronic device 2200 shown may also include a communication unit ( Figure 6 (Not shown in the image), this communication unit is used to support electronic device 2200 in performing the steps of communication between electronic device and other electronic devices in the embodiments of this application.

[0328] Optional, Figure 6The illustrated electronic device 2200 may further include a storage unit 2203 that stores programs or instructions. When the processing unit 2201 executes the program or instructions, it causes... Figure 6 The electronic device 2200 shown can perform the method described in the above-described method embodiments.

[0329] Figure 6 The technical effects of the electronic device 2200 shown can be referred to the technical effects of the method shown in the above method embodiments, and will not be repeated here. Figure 6 The processing unit 2201 involved in the illustrated electronic device 2200 can be implemented by a processor or processor-related circuit components, and can be a processor or a processing module. The communication unit can be implemented by a transceiver or transceiver-related circuit components, and can be a transceiver or a transceiver module.

[0330] This application also provides a chip system, such as... Figure 7 As shown, the chip system includes at least one processor 2301 and at least one interface circuit 2302. The processor 2301 and the interface circuit 2302 are interconnected via lines. For example, the interface circuit 2302 can be used to receive signals from other devices. As another example, the interface circuit 2302 can be used to send signals to other devices (e.g., the processor 2301). Exemplarily, the interface circuit 2302 can read instructions stored in memory and send those instructions to the processor 2301. When the instructions are executed by the processor 2301, the electronic device can perform the various steps performed by the electronic device in the above embodiments. Of course, the chip system may also include other discrete devices, and this application embodiment does not specifically limit this.

[0331] Optionally, the chip system may contain one or more processors. These processors can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor, implemented by reading software code stored in memory.

[0332] Optionally, the chip system may contain one or more memories. The memory may be integrated with the processor or disposed separately from it; this application does not limit this. For example, the memory may be a non-transient processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or disposed separately on different chips. This application does not specifically limit the type of memory or the arrangement of the memory and processor.

[0333] For example, the chip system can be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a micro controller unit (MCU), a programmable logic device (PLD), or other integrated chips.

[0334] It should be understood that each step in the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The method steps disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0335] It should be noted that the electronic device provided in this application embodiment belongs to the same concept as the method in the above embodiments. Any of the methods provided in the method embodiments can be run on the electronic device, and the specific implementation process is detailed in the method embodiments, which will not be repeated here. For example, the processor in the electronic device can execute the steps in the method. The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.

[0336] This application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in any of the above embodiments.

[0337] In the embodiments of this application, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0338] It should be noted that, for the methods of the embodiments of this application, those skilled in the art will understand that all or part of the processes of the methods of the embodiments of this application can be implemented by a computer program controlling related hardware. This computer program can be stored in a computer-readable storage medium, such as in the memory of an electronic device, and executed by at least one processor within the electronic device. During execution, it can include the processes of the embodiments of the method. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, etc.

[0339] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0340] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although this application has disclosed preferred embodiments as above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. A method for outputting conference video, characterized in that, The method includes: When a sound source is detected, the sound source is located to determine the tracking area where the sound source is located and the corresponding location status of the sound source. The location status includes one of the following: the specific speaker is located, the specific speaker is not located but the speaking area is located, and neither the specific speaker nor the speaking area is located. Obtain the speech data corresponding to the sound source; Obtain the screen display mode of the tracking area; Based on the positioning status, the speaking data, and the screen display mode, the target conference screen is output.

2. The method according to claim 1, characterized in that, The target conference screen includes a single screen and a combined screen; the output of the target conference screen based on the positioning status, the speaking data, and the screen display mode includes: Based on the spoken data, determine whether the conditions for screen switching are met; If the screen switching conditions are met, a target screen display mode is determined based on the currently acquired screen display mode, the speech data, and / or the positioning status, wherein the target screen display mode includes a single screen mode and a combined screen mode; Based on the positioning status and the target screen display mode, the corresponding single screen or the combined screen is output.

3. The method according to claim 2, characterized in that, Determining the target screen display mode based on the currently acquired screen display mode, the speech data, and / or the positioning status includes: If the current screen display mode is single-screen mode, and based on the speech data and / or the positioning status it is determined that at least one additional screen corresponding to another sound source needs to be displayed, then the target screen display mode is determined to be a combined screen mode; or, If the current screen display mode is a combined screen mode, and based on the speech data and / or the positioning status it is determined that the number of sub-screens in the combined screen needs to be reduced to one, then the target screen display mode is determined to be a single-screen mode; or If the target screen display mode is the same as the current screen display mode, then the current screen display mode remains unchanged, and the sub-screen content in the single screen or the combined screen is updated based on the positioning status.

4. The method according to claim 3, characterized in that, Updating the single screen based on the positioning status includes: If the current single screen is a close-up of the first speaker, and the location status is updated to locate the second speaker, and the second speaker meets the preset conditions, then the single screen will be switched from the close-up of the first speaker to the close-up of the second speaker.

5. The method according to claim 3, characterized in that, Updating the single screen based on the positioning status includes: If the current single screen is a close-up of a specific speaker, and the location status is downgraded from being located to the specific speaker to not being located to the specific speaker but being located to the speaking area, then the single screen will be switched from the close-up of the specific speaker to the corresponding speaking area screen.

6. The method according to claim 3, characterized in that, Updating the sub-screen content in the combined screen based on the positioning status includes: If the currently displayed combined screen includes a first sub-screen, based on the updated positioning state, perform one of the following update operations on the first sub-screen: Replace the first sub-screen with the close-up or speaking area screen corresponding to the first sound source, and replace it with the close-up or speaking area screen corresponding to the second sound source. The first sub-screen is downgraded from a close-up of the specific speaker to a screen showing the speaking area; Upgrade the first sub-screen from the speaking area screen or the preset position screen to a close-up screen of the specific speaker.

7. The method according to claim 3, characterized in that, Updating the sub-screen content in the combined screen based on the positioning status includes: If the currently displayed combined screen consists of at least one first sub-screen, based on the updated positioning status and / or the speech data, at least one second sub-screen that needs to be displayed is re-determined, and the combined screen is switched as a whole to a target combined screen consisting of the at least one second sub-screen. Among them, the set consisting of at least one second sub-picture and the set consisting of at least one first sub-picture may have sub-pictures added, reduced, replaced, or completely replaced.

8. The method according to claim 2, characterized in that, The speech data includes a first continuously accumulated speech duration and / or the speech frequency within a preset time window; wherein, the first continuously accumulated speech duration represents the total duration obtained by accumulating the duration of each speech that occurs within the tracking area, starting from the first speech in any speech round, until the interval between two adjacent speeches is greater than or equal to a preset interval threshold. The step of determining whether the screen switching conditions are met based on the speech data includes: If the first continuous cumulative speaking duration is greater than or equal to the first duration threshold, and / or the speaking frequency is greater than or equal to the frequency threshold, then the screen switching condition is determined to be met. If the first continuous cumulative speaking duration is less than the first duration threshold, and / or the speaking frequency is less than the frequency threshold, then it is determined that the screen switching condition is not met.

9. The method according to claim 2, characterized in that, The single screen includes a close-up view of the specific speaker, a view of the speaking area, and a preset position view corresponding to the tracking area; based on the positioning status and the target screen display mode, the corresponding single screen is output, including: When the target screen display mode is single-screen mode. If the location status indicates that a specific speaker has been located, then a close-up shot of that specific speaker will be output. If the positioning status is that a specific speaker has not been located but the speaking area has been located, then the speaking area screen will be output. If the positioning status indicates that neither the specific speaker nor the speaking area has been located, then the preset position image corresponding to the tracking area will be output.

10. The method according to claim 9, characterized in that, The speech data includes a second continuous cumulative speech duration; wherein, the second continuous cumulative speech duration represents the total duration obtained by accumulating the duration of each speech of a single sound source, starting from the first speech of the current speech round of the single sound source, until the interval between two adjacent speeches is greater than or equal to a preset interval threshold. When the positioning status indicates that a specific speaker has been located, outputting a close-up image of the specific speaker includes: When the positioning status is that a specific speaker has been located, and the second continuous cumulative speaking time is greater than or equal to the second duration threshold, a close-up shot of the specific speaker is output. When the positioning status is that a specific speaker has not been located but the speaking area has been located, the speaking area screen is output, including: When the positioning status is that a specific speaker has not been located but the speaking area has been located, and the second continuous cumulative speaking time is greater than or equal to the second time threshold, the speaking area screen is output; When the positioning status indicates that neither the specific speaker nor the speaking area has been located, the preset position image corresponding to the tracking area is output, including: If the location status is such that neither the specific speaker nor the speaking area is located, or if the second cumulative speaking time is less than the second duration threshold, the preset position image corresponding to the tracking area is output.

11. The method according to claim 2, characterized in that, The combined view includes a panoramic view of the conference and sub-views; based on the positioning status and the target view display mode, the corresponding combined view is output, including: When the positioning status is that a specific speaker has been located, the sub-screen is determined to be a close-up shot of the specific speaker; When the positioning status is that a specific speaker has not been located but the speaking area has been located, the sub-screen is determined to be the speaking area screen; If the positioning status indicates that neither the specific speaker nor the speaking area has been located, the sub-screen is determined to be the preset position screen corresponding to the tracking area; The output includes a combined view of the conference panoramic view and the sub-views.

12. The method according to claim 11, characterized in that, When multiple sound sources are detected, and each sound source meets the screen switching conditions, the output includes a combined image of the conference panoramic view and the sub-views, including: The output includes a combined image of the conference panoramic view and multiple sub-views, with each sub-view corresponding to a specific sound source.

13. The method according to claim 12, characterized in that, The method further includes: Based on the speech data from the multiple sound sources, the display priority of the multiple sub-screens is determined; The highest priority sub-screen will be displayed in the preset main display area of ​​the combined screen; or, The corresponding sub-screens are displayed in different sizes according to their display priority from high to low. The display size of the sub-screen with higher display priority is larger than that of the sub-screen with lower display priority.

14. The method according to claim 12, characterized in that, The method further includes: If the same speaker is present in multiple sub-screens, the duplicate sub-screens are deduplicated or merged.

15. The method according to claim 12, characterized in that, After outputting a combined view including the panoramic view of the conference and the multiple sub-views, the method further includes: If a new sound source is detected, and the cumulative speaking time of the new sound source exceeds a preset duration threshold, a sub-screen corresponding to the new sound source is added to the combined screen.

16. The method according to claim 12, characterized in that, After outputting a combined view including the panoramic view of the conference and the multiple sub-views, the method further includes: If any of the multiple sub-screens has a speaking interval duration greater than or equal to a preset interval threshold, then the sub-screen with a speaking interval duration greater than or equal to the preset interval threshold is removed from the combined screen.

17. The method according to claim 1, characterized in that, After determining the tracking area where the sound source is located, the method further includes: Drive the camera attached to the tracking area to turn to a preset position in the tracking area; The output target conference screen includes: When the camera is rotated to the preset position, the display is switched from a panoramic view of the meeting to the target meeting view. The target meeting view includes a close-up of a specific speaker, a view of the speaking area, or the preset position view, which is obtained by cropping the image taken by the camera at the preset position.

18. The method according to claim 10, characterized in that, When outputting a preset position image corresponding to the tracking area, the method further includes: When the positioning status is upgraded to locate a specific speaker, and the second continuous cumulative speaking time is greater than or equal to the second duration threshold, the system controls the switching from the preset position screen to a close-up screen of the specific speaker. When the positioning status is upgraded to not locating a specific speaker but locating a speaking area, and the second continuous cumulative speaking duration is greater than or equal to the second duration threshold, the control switches from the preset position screen to the corresponding speaking area screen.

19. The method according to claim 10, characterized in that, When outputting a close-up shot of the specific speaker, the method further includes: While the speaker is still speaking, continue to maintain a close-up shot of the speaker. If the specific speaker finishes speaking, determine whether a new speaker is detected; If no new speaker is detected, switch back to the panoramic view of the meeting or the preset position view corresponding to the tracking area; If a new speaker is detected, the system first switches back to the panoramic view of the meeting or the preset position view corresponding to the tracking area where the new sound source is located. If the cumulative speaking time of the new speaker meets the preset continuous speaking conditions, the system switches to a close-up view or speaking area view corresponding to the new speaker based on the positioning status of the new speaker.

20. The method according to claim 19, characterized in that, Maintaining a close-up shot of the specific speaker while they are still speaking includes: In the event that movement of the specific speaker is detected. If the specific speaker does not exceed the tracking area, the camera bound to the tracking area will continue to output a close-up image of the specific speaker; If the specific speaker is outside the tracking area, a new tracking area is determined, and the screen is first switched to the preset position screen corresponding to the new tracking area. Then, based on the positioning status of the specific speaker by the camera bound to the new tracking area and the screen display mode configured in the new tracking area, the preset position screen corresponding to the new tracking area is switched to a close-up screen of the specific speaker or the corresponding speaking area screen.

21. The method according to claim 19, characterized in that, When a specific speaker finishes speaking and no new speaker is detected, switching back to the panoramic view of the meeting or the preset position view corresponding to the tracking area includes: In the event that movement of the specific speaker is detected. If the specific speaker does not exceed the tracking area, switch back to the panoramic view of the meeting or the preset position view corresponding to the tracking area; If the specific speaker is outside the tracking area, switch back to the panoramic view of the meeting; or, determine the new tracking area where the specific speaker is located and switch back to the preset position view corresponding to the new tracking area.

22. The method according to claim 10, characterized in that, When outputting a close-up shot of the specific speaker, the method further includes: If the location status is downgraded from locating a specific speaker to not locating a specific speaker but locating the speaking area, and the specific speaker is not located again for a preset time, the control switches to the corresponding speaking area screen. If the specific speaker is subsequently re-identified, the speaking area screen is switched to a close-up view of the specific speaker.

23. The method according to claim 22, characterized in that, The method further includes: In the same tracking session, if the location status for the same or similar sound source is downgraded from locating a specific speaker to locating the speaking area but not a specific speaker, the current output of the specific speaker's close-up view will continue to be maintained, and the corresponding speaking area view will not be switched.

24. The method according to claim 10, characterized in that, When outputting a close-up shot of the specific speaker, the method further includes: If the specific speaker has not finished speaking and a new speaker is detected, and if the new speaker or the new speaker's speaking area can be located, and the new speaker's continuous cumulative speaking time exceeds a preset time threshold, then it is determined whether the new speaker and the specific speaker are located in the same tracking area. If they are in the same tracking area and the new speaker can be located, the screen will switch to one of the following: a single screen that includes both the specific speaker and the new speaker; a preset screen of the tracking area; a screen that is stitched together from close-up shots of the specific speaker and the new speaker respectively; or a combined screen that includes a panoramic view of the meeting and close-up shots of the specific speaker and the new speaker. If they are in the same tracking area and the speaking area of ​​the new speaker can be located, then the screen will switch to one of the following: a single screen that includes both the specific speaker and the new speaker, a preset position screen of the tracking area, or a combined screen that includes a panoramic view of the meeting, a close-up view of the specific speaker, and a screen of the speaking area corresponding to the new speaker. If they are not in the same tracking area and the new speaker can be located, the screen will switch to one of the following: a screen that is spliced ​​together with a close-up of the specific speaker and a close-up of the new speaker; or a combined screen that includes a panoramic view of the meeting and close-up shots of the specific speaker and the new speaker respectively. If they are not in the same tracking area, and the speaking area of ​​the new speaker can be located, then the screen will switch to one of the following: a screen that is a combination of a close-up of the specific speaker and a screen showing the speaking area of ​​the new speaker, or a screen that includes a panoramic view of the meeting and a combination of a close-up of the specific speaker and a screen showing the speaking area of ​​the new speaker.

25. The method according to claim 10, characterized in that, When outputting a close-up shot of the specific speaker, the method further includes: If a specific speaker has not finished speaking and a new speaker is detected, determine the percentage of speaking time for each speaker within a preset time period; Based on the speaking time percentage of each speaker, a target speaker is determined from the specific speakers and the new speakers, wherein the speaking time percentage of the target speaker is greater than a first percentage threshold. If the target speaker is not the specific speaker, the camera attached to the tracking area where the target speaker is located is driven to pre-select the target speaker; If the speaking time of the target speaker exceeds the second percentage threshold and the target speaker has not finished speaking, switch to a close-up shot of the target speaker.

26. The method according to claim 1, characterized in that, In the case of detecting multiple speakers, the method further includes: Based on each speaker's speaking data and image data, determine the priority of each speaker; If the screen display mode is single speaker mode, then the speaker with the highest priority will be identified as the target speaker; If the screen display mode is multi-speaker mode, then select the number of speakers displayed in the multi-speaker mode from the multiple speakers in descending order of priority as the target speaker; The output target conference screen includes: Output the target meeting screen corresponding to the target speaker.

27. The method according to claim 26, characterized in that, The process of determining the priority of each speaker based on their speaking data and image data includes: Based on each speaker's speaking data, determine the percentage of speaking time, volume, and relevance of speaking content to the meeting topic for each speaker within a preset time period; Based on the image data of each speaker, determine the facial orientation and gaze focus of each speaker; The priority of each speaker is determined based on the percentage of speaking time, the volume of their voice, the relevance of their speech to the meeting topic, their facial orientation, and their gaze focus. The priority satisfies the following relationship: The higher the percentage of speaking time a speaker spends, the higher their priority. The louder the sound, the higher the priority. The higher the relevance, the higher the priority. Cameras that face the subject directly have higher priority; cameras that face the subject from the side have lower priority. The closer the eye is to the display screen or whiteboard, the higher the priority.

28. The method according to claim 1, characterized in that, The process of acquiring the screen display mode of the tracking area includes: Based on the gaze direction of the meeting participants, the screen display mode is determined to be one of the following: presentation mode, discussion mode, single speaker mode, or multiple speaker mode, so as to correspondingly determine the output of a single speaker screen, a discussion area screen, or a combination screen containing multiple sub-screens.

29. The method according to claim 28, characterized in that, When the screen display mode is discussion mode, the target meeting screen is the discussion area screen, or a combination of close-up screens of multiple discussion participants; In cases where the target meeting screen is a combination of close-up shots of multiple participants, the method further includes: If the speaking interval of the first participant among multiple participants exceeds the interval threshold, and the current discussion has not ended, then the combined close-up shots of the multiple participants will continue to be displayed.

30. The method according to claim 10, characterized in that, Before outputting a close-up shot of the specific speaker, the method further includes: Semantic analysis is performed on the speech content of the specific speaker to predict the probability of the specific speaker continuing to speak and the remaining speaking time; If the probability of continuing to speak is lower than a probability threshold, and / or the remaining speaking time is less than a speaking time threshold, the current screen remains unchanged.

31. The method according to claim 10, characterized in that, The preset interval threshold, the first duration threshold, and the second duration threshold are determined according to different tracking regions, and / or the preset interval threshold, the first duration threshold, and the second duration threshold can be adaptively and dynamically adjusted based on the historical speech data of the corresponding tracking region.

32. The method according to claim 1, characterized in that, Determining the localization state corresponding to the sound source includes: At least one candidate object is identified within the tracking area, wherein the at least one candidate object is a candidate speaker or a candidate speaking area; Determine the correlation between each candidate object and the sound source; Based on the correlation between each candidate object and the sound source, the localization status of the sound source is determined; The degree of correlation is determined by at least one of the following: The degree of positional matching between the candidate object and the sound source; The degree of consistency between the lip movement corresponding to the candidate object and the current speech activity of the sound source; The degree of matching between the posture corresponding to the candidate object and the current speaking state of the sound source; The continuity between the candidate object and the localization result of the sound source at the previous moment.

33. The method according to claim 32, characterized in that, Determining the localization state of the sound source based on the correlation between each candidate object and the sound source includes: If the relevance of a candidate speaker is greater than or equal to the relevance threshold, then the positioning status is determined to be that a specific speaker has been located. If the relevance of a candidate speaker is less than the relevance threshold, and the relevance of a candidate speaking area is greater than or equal to the relevance threshold, then the positioning status is determined to be that a specific speaker has not been located but a speaking area has been located. If the relevance of a candidate speaker is less than the relevance threshold, and the relevance of a candidate speaking area is less than the relevance threshold, then the positioning status is determined to be that neither the specific speaker nor the speaking area has been located.

34. A conference video output system, characterized in that, Includes processing devices, microphones, and cameras; The processing device is used to locate the sound source through the microphone when a sound source is detected, and to determine the tracking area where the sound source is located and the corresponding positioning state of the sound source. The positioning state includes one of the following: the specific speaker is located, the specific speaker is not located but the speaking area is located, and neither the specific speaker nor the speaking area is located. The processing device is used to acquire speech data corresponding to the sound source; The processing device is used to acquire the screen display mode configured in the tracking area; The processing device is used to control the camera to output the target conference image based on the positioning status, the speaking data, and the screen display mode.

35. A computer-readable storage medium, characterized in that, It stores a computer program or instructions, which, when executed by a processor, implement the conference screen output method as described in any one of claims 1 to 33.

36. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the conference screen output method as described in any one of claims 1 to 33.