A conference picture output method and device and a storage medium

CN122534192APending Publication Date: 2026-08-07YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YEALINK (XIAMEN) NETWORK TECHNOLOGY CO LTD
Filing Date
2026-07-03
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0006]提供一种会议画面输出方法、装置及存储介质,旨在解决视频会议终端的画面输出能力的问题

Benefits of technology

本技术方案以广角相机采集的全景画面为统一空间基准建立分区模型,将第一空间划分为多个子空间并分别配置对应的特写相机——其中第一、第二子空间配置固定安装的长焦相机,第三子空间配置具备转动和变焦能力的云台相机。在此基础上,各子空间独立确定需要输出的主目标,由对应的相机生成特写画面。由于固定式长焦相机的光轴方向在安装后保持恒定,生成特写画面时无需机械转动,其输出画面为光学直出,不引入机械振动和转动延迟,在时间维度上具备天然的连续稳定性;而云台相机被配置于第三子空间,可在该子空间范围内灵活转动和变焦,覆盖固定长焦视场角无法企及的远距离席位和非预设区域,在空间维度上具备灵活的机动性。二者通过分区模型被纳入同一空间基准框架下,固定相机负责稳定输出,机动相机负责灵活覆盖,各司其职,互不干扰,从而在同一系统架构内同时获得了光学特写的稳定性和大范围调度的灵活性。与此同时,全景画面作为统一的空间基准为所有相机提供了共享的坐标参照系,各相机的输出画面均锚定于该坐标系下,固定相机与云台相机的特写画面之间的相对空间关系保持明确,当输出画面在不同相机之间切换时,目标在全景坐标系中的位置连续,特写画面在空间上保持一致的指向性,避免了切换瞬间的空间视差跳变。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122534192A_ABST
    Figure CN122534192A_ABST
Patent Text Reader

Abstract

The application discloses a conference picture output method and device and a storage medium, and belongs to the technical field of electronics and camera. A panoramic picture of a first space collected by a wide-angle camera is acquired, and a partition model is established with the panoramic picture as a space reference; the partition model divides the first space into a plurality of subspaces, and the plurality of subspaces at least include a first subspace, a second subspace and a third subspace; a main target of each subspace is determined in the interior of each subspace; a camera corresponding to each subspace is called to generate a close-up picture corresponding to each main target; wherein the first subspace and the second subspace correspond to a first long-focus camera and a second long-focus camera respectively, and the third subspace corresponds to a gimbal camera; and a target close-up picture is output according to a preset output strategy, so that scheduling conflicts of multiple cameras competing for the same target are avoided, the panoramic picture of the wide-angle camera serves as a unified space reference, and the close-up pictures output by the cameras are consistent in space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of electronics and photography, specifically to a method, apparatus, and storage medium for outputting conference images. Background Technology

[0002] Currently, with the rapid growth in demand for remote collaboration, video conferencing systems have evolved into intelligent terminal devices that need to provide immersive, multi-level visual experiences. Modern enterprise meeting room scenarios are becoming increasingly complex, with diverse layouts (long table type, U-shaped, theater-style) and varied meeting interaction modes (discussion, whiteboard presentation, remote defense), which places higher demands on the image output capabilities of video conferencing terminals.

[0003] Currently, video conferencing terminals on the market mainly adopt the following two types of visual solutions: Option A, a combination of wide-angle and PTZ cameras. The wide-angle camera captures a panoramic view, while the PTZ (Pan-Tilt-Zoom) camera tracks a specific target and outputs a close-up image through gimbal rotation and optical zoom. However, when the PTZ camera switches targets, it undergoes a physical rotation process, during which the image cannot be continuously output. This causes the position of the distant image to change before and after the switch, resulting in a visual dizzying effect.

[0004] Option B is a combination of wide-angle and fixed telephoto lenses. A fixed telephoto camera is responsible for outputting close-up images of the preset area. However, due to the fixed field of view, the telephoto camera cannot cover non-preset areas such as whiteboards and physical display stands. If relying on digital cropping of the wide-angle image, the image quality will be significantly degraded. In addition, when the target moves to the boundary between the two telephoto fields of view, a switching blind spot will occur.

[0005] It is evident that Solution A struggles to guarantee the continuity of image output, while Solution B lacks coverage flexibility—neither can simultaneously achieve "stable close-up output" and "flexible scheduling" in complex meeting scenarios. Therefore, in complex multi-person video conferencing scenarios, how to establish a unified collaborative scheduling framework among various heterogeneous cameras, using wide-angle panoramic views as the global spatial benchmark, to achieve an organic unity between "stable close-up output" and "flexible scheduling" has become an urgent problem to be solved. Summary of the Invention

[0006] This invention provides a method, apparatus, and storage medium for outputting conference video, aiming to solve the problem of video output capability in video conferencing terminals.

[0007] Firstly, a method for outputting conference video is provided, including the following steps: A panoramic image of the first space captured by a wide-angle camera is acquired, and a partitioning model is established based on the panoramic image. The partitioning model divides the first space into multiple subspaces, and the multiple subspaces include at least a first subspace, a second subspace, and a third subspace. Within each of the subspaces, determine the primary objective of each subspace; The camera corresponding to the subspace is invoked to generate a close-up image corresponding to each main target; wherein the first subspace and the second subspace correspond to the first telephoto camera and the second telephoto camera, respectively, and the third subspace corresponds to the gimbal camera; According to the preset output strategy, output a close-up image of the target.

[0008] In one possible design approach, determining the primary objective of each of the subspaces, respectively, includes: Within each of the subspaces, attention evaluation is performed on the candidate targets within each subspace, and the main target of each subspace is determined based on the evaluation results.

[0009] In one possible design approach, the step of performing attention evaluation on candidate targets within each of the subspaces and determining the primary target of each subspace based on the evaluation results includes: Obtain multidimensional feature information for each candidate target within the subspace. The multidimensional feature information includes one or more of the following: speaking state, motion state, spatial location, posture, dialogue history, meeting role label, or electronic device usage status. The attention weight for each candidate target is determined based on the multidimensional feature information; The candidate target with the highest attention weight in the subspace is determined as the main target.

[0010] In one possible design approach, the preset output strategy includes at least one of the following strategies: First strategy: In response to the target moving across subspace, switch the camera used to generate close-up images of the target and continuously output close-up images of the target; Second strategy: Output the multiple close-up shots corresponding to each main target as the target close-up shots in parallel; Third strategy: Combine the multi-channel close-up shots corresponding to each main target into a single-channel combined shot and output it as the target close-up shot; Fourth strategy: Perform cross-subspace attention comparison on each main target, and select one close-up shot as the target close-up shot for output; Fifth strategy: Pause the output of close-up images of human subjects and switch to outputting close-up images of non-human subject areas.

[0011] In one possible design approach, the first strategy includes: When the first target is detected to move from the first subspace to the third subspace, the first telephoto camera is controlled to continuously generate close-up images of the first target until the first target moves out of the field of view of the first telephoto camera, and / or enters the field of view of the gimbal camera, the gimbal camera is controlled to take over the generation of close-up images of the first target.

[0012] In one possible design approach, the first strategy further includes: During the process of switching between the first telephoto camera and the gimbal camera to generate a close-up image of the first target, the first telephoto camera is controlled to output a transitional image of the first target to the edge of the field of view of the first telephoto camera by moving the cropping window in the captured image.

[0013] In one possible design approach, the second strategy includes: Obtain the attention weights of each main target in its respective subspace; The attention weights of each main target are sorted, and close-up shots that meet the preset conditions are selected according to the sorting results and output in parallel.

[0014] In one possible design approach, the preset conditions include any of the following: The top N positions, where N is a preset positive integer; The attention weight is greater than the preset threshold.

[0015] In one possible design approach, the third strategy includes: Obtain the attention weights of each main target in its respective subspace; The attention weights of each main target are sorted, and the arrangement position and / or display size of each close-up shot in the combined shot are determined according to the sorting result.

[0016] In one possible design approach, the fourth strategy includes: Obtain the attention weights of each main target in its respective subspace; The attention weights of each main target are compared across subspaces, and the close-up image corresponding to the main target with the highest attention weight is selected as the target close-up image for output.

[0017] In one possible design approach, the fourth strategy further includes: When the attention weight of the selected main target is lower than a preset threshold, the current output screen remains unchanged.

[0018] In one possible design approach, the fifth strategy includes: The panoramic image is used for target recognition. When a non-human target area is detected and a preset trigger condition is met, the gimbal camera is controlled to rotate toward the non-human target area to generate a close-up image of the non-human target area. The non-human target area includes at least one of a whiteboard area, a projection screen area, or a physical display stand area.

[0019] In one possible design approach, the preset triggering conditions include: Content updates were detected in the non-human target area.

[0020] In one possible design approach, the preset output strategy further includes: At least two basic strategies from the first to the fifth strategy are combined to form a combined strategy to output the close-up image of the target.

[0021] In one possible design approach, the preset output strategy further includes: Obtain the current meeting mode, and determine the target strategy from the basic strategy or the combined strategy based on the meeting mode.

[0022] One possible design approach also includes: For each target detected in the panoramic image, maintain a target state record. The target state record includes one or more of the following information: target identifier, current subspace, real-time coordinates in the first space, motion vector, timestamp of entering the current subspace, or confidence level of the subspace. The subspace to which the current subspace belongs is determined by matching the real-time coordinates with the spatial range of each subspace.

[0023] In one possible design approach, within any of the said subspaces, when the main target switches from an old target to a new target, generating a close-up shot corresponding to each of the said main targets includes: The corresponding camera is controlled to move the cropping window within the captured image by a preset movement step size, generating a close-up image transitioning from the old target to the new target; the starting position of the cropping window is the location of the old target, and the ending position is the location of the new target; The attention weight of the new target is greater than the attention weight of the old target, and the difference between the two exceeds a preset weight threshold.

[0024] In one possible design, adjacent subspaces within the plurality of subspaces have transition areas for cross-zone pick-up and drop-off; the method further includes: When a target is detected moving from its current source subspace to an adjacent target subspace via the transition region, cross-regional pick-up and drop-off scheduling is performed: Predict the target position of the target entering the target subspace based on the panoramic image; If the attention weights of all other targets in the target subspace are less than or equal to the attention weight of the target, then the camera corresponding to the target subspace is controlled to pre-rotate toward the target position or take over to generate a close-up image of the target after the target enters its field of view; If there is a high-attention target being tracked in the target subspace, the camera corresponding to the target subspace is controlled to adjust its field of view to simultaneously cover the target and the high-attention target, generating a frame-to-frame or single-person close-up image of the target and the high-attention target.

[0025] In one possible design approach, when the camera corresponding to the source subspace is a fixed telephoto camera and the camera corresponding to the target subspace is a gimbal camera: The step of controlling the camera corresponding to the target subspace to pre-rotate toward the target position includes: controlling the gimbal camera to rotate toward the target position; The step of taking over the generation of a close-up image of the target after it enters its field of view includes: During the rotation of the gimbal camera, the fixed telephoto camera is controlled to continuously output close-up images of the target by moving the cropping window in the captured image until the target moves out of the field of view of the fixed telephoto camera, or until the gimbal camera completes its rotation. After the target enters the field of view of the gimbal camera, the system switches to an optical close-up image of the target generated by the gimbal camera.

[0026] In one possible design approach, when the camera corresponding to the source subspace is a first telephoto camera and the camera corresponding to the target subspace is a second telephoto camera, the execution of cross-regional pick-up and drop-off scheduling includes: The first telephoto camera is controlled to continuously output close-up images of the target by moving the cropping window in the captured image until the target moves out of the field of view of the first telephoto camera. After the target enters the field of view of the second telephoto camera, the camera switches to a close-up view of the target generated by the second telephoto camera. Within a preset duration after switching, the second telephoto camera is controlled to continuously output close-up images of the target located in the center of the image by moving the cropping window in the captured image.

[0027] In one possible design approach, the method of moving the clipping window includes: Based on the real-time position of the target in the panoramic image, determine the cropping position of the target in the image captured by the corresponding camera; The cropping window is moved from the starting position to the ending position with a preset movement step size so that the position of the target in the output screen remains continuous.

[0028] In one possible design approach, the execution of cross-regional pick-up and drop-off scheduling further includes: When a target is detected moving from the third subspace to the first subspace via the transition region, the gimbal camera is controlled to continuously track the target; and / or After confirming that the target has entered the first subspace and that the first telephoto camera has captured the target based on the panoramic view, the system switches to a close-up view of the target generated by the first telephoto camera.

[0029] In one possible design, controlling the gimbal camera to continuously track the target includes: When at least one of the following triggering conditions is met, the gimbal camera is controlled to continuously track the target: The target is a preset priority tracking target; The attention weights of other targets in the third subspace are lower than a preset threshold.

[0030] One possible design approach also includes: During the generation of a close-up image of the target by the gimbal camera, if a change in the target's posture is detected and it is predicted that the target will move beyond the current field of view of the gimbal camera, then: The gimbal camera is controlled to generate a transitional image during the adjustment of the pitch angle. The transitional image is a dynamically digitally cropped image obtained from the image currently captured by the gimbal camera. The cropped area moves with the real-time position of the target. After the gimbal camera is adjusted to the target pitch angle, an optical close-up image of the target is generated.

[0031] One possible design approach also includes: Based on the current pose of the gimbal camera, determine the virtual coaxial reference plane (VCRP) parameters of the gimbal camera; When the first telephoto camera, the second telephoto camera, or the gimbal camera generates a close-up image, a position compensation value is determined based on the differences between the VCRP parameters of each camera, and the generated close-up image is adjusted based on the position compensation value.

[0032] One possible design approach also includes: When the target in the panoramic image is detected to be concentrated in the far-end area, the gimbal camera is controlled to generate a collective image with a wide field of view; or... When the presence of two targets speaking alternately within the third subspace is detected, the gimbal camera is controlled to capture images containing the two targets with a wide field of view and digitally cropped to generate two independent single-person close-up images.

[0033] Secondly, this application provides a video conferencing terminal, including a housing and a camera assembly disposed on the housing: A dual wide-angle stitching module is used to acquire images and stitch them together to generate the panoramic image. A dual telephoto zone close-up module, comprising a first telephoto camera and a second telephoto camera; A gimbal dynamic maneuvering module, including the gimbal camera; A processor is a method for implementing any of the above-mentioned aspects of any design.

[0034] Optionally, the video conferencing terminal may further include the memory and / or the communication interface.

[0035] The communication interface is coupled to the processor and is used for inputting and / or outputting information.

[0036] The memory is used to store computer programs. The processor is configured to execute any of the methods described above, and can be implemented as: executing the computer program stored in the memory to execute any of the methods described above. Alternatively, the processor can be a hardware-implemented circuit, such as an artificial intelligence (AI) processor, to improve operating speed. This application does not limit the specific implementation of the processor.

[0037] Optionally, the device can be a complete machine or a module within a machine, such as a chip.

[0038] Thirdly, the present application provides a computer-readable storage medium including computer instructions that, when executed on a device, cause the device to perform any of the possible designs described above.

[0039] Fourthly, the present application provides a computer program product that, when run on a device, causes the device to execute any of the possible designs described above.

[0040] Fifthly, this application provides a circuit system including a processing circuit configured to perform the methods in any possible design of any of the above aspects. The processing circuit can be implemented as a corresponding circuit component, such as one or more processors. Alternatively, it can be implemented as a processor and a memory. Yet another example is a processor and a transceiver.

[0041] In a sixth aspect, this application provides a chip system including at least one processor and at least one interface circuit, wherein the at least one interface circuit is used to perform transceiver functions and send instructions to the at least one processor, and when the at least one processor executes instructions, the at least one processor performs the method described in the first aspect and any of the designs therein.

[0042] One of the above technical solutions has the following advantages or beneficial effects: This technical solution establishes a partitioned model based on panoramic images captured by a wide-angle camera, dividing the first space into multiple subspaces and configuring corresponding close-up cameras in each—the first and second subspaces are configured with fixed telephoto cameras, and the third subspace is configured with a gimbal camera capable of rotation and zoom. Based on this, each subspace independently determines the main target to be output, and the corresponding camera generates the close-up image. Since the optical axis of the fixed telephoto camera remains constant after installation, no mechanical rotation is required when generating close-up images; its output image is optically direct, without introducing mechanical vibration or rotational delay, possessing inherent continuous stability in the time dimension. The gimbal camera, configured in the third subspace, can flexibly rotate and zoom within this subspace, covering distant positions and non-preset areas that the fixed telephoto camera's field of view cannot reach, possessing flexible mobility in the spatial dimension. Both are incorporated into the same spatial reference framework through the partitioned model; the fixed camera is responsible for stable output, and the mobile camera is responsible for flexible coverage, each performing its own function without interference, thus simultaneously achieving optical close-up stability and large-scale scheduling flexibility within the same system architecture. At the same time, the panoramic view serves as a unified spatial benchmark, providing a shared coordinate reference system for all cameras. The output images of each camera are anchored to this coordinate system, and the relative spatial relationship between the close-up images of the fixed camera and the gimbal camera remains clear. When the output image switches between different cameras, the position of the target in the panoramic coordinate system is continuous, and the close-up image maintains a consistent directionality in space, avoiding spatial parallax jumps at the moment of switching. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart illustrating a conference screen output method provided by an exemplary embodiment of this disclosure; Figure 2 This is a schematic diagram of a scenario for a conference screen output method provided by an exemplary embodiment of this disclosure; Figures 3-6 This is a schematic diagram of the structure of the apparatus provided in an exemplary embodiment of this disclosure. Detailed Implementation

[0045] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0046] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0047] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.

[0048] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not exclude the applicability to or configuration to devices performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values ​​may in practice be based on additional conditions or values ​​beyond those stated.

[0049] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0050] In real-time conferencing scenarios, video conferencing terminals need to flexibly switch screen outputs between multiple targets and regions while ensuring visual comfort for remote viewers. To address this need, two main technological approaches currently exist: The first type uses a combination of a wide-angle lens and a mechanical PTZ gimbal camera. The wide-angle lens captures a panoramic view, while the PTZ camera tracks and zooms on specific targets for close-up shots through physical rotation and optical zoom. Because the PTZ camera needs to frequently start and stop rotating when tracking different speakers, the image cannot be continuously output during the rotation, causing the position of the distant image to jump before and after switching, resulting in visual dizziness.

[0051] The second type uses a combination of wide-angle and fixed telephoto cameras. Multiple fixed telephoto cameras cover a preset area, and a digital electronic gimbal is used to crop and switch the image. Due to the fixed field of view, this type of solution cannot flexibly cover non-preset areas such as whiteboards and physical display stands; when the target moves to the boundary of the two telephoto fields of view, blind spots are likely to occur.

[0052] The two aforementioned solutions represent the mainstream design approaches for current video conferencing output, each employing different technical routes—mechanical rotation and fixed field of view—but both involve trade-offs between "stable image output" and "flexible coverage scheduling." In this field, fixed telephoto lenses and mobile pan-tilt units are typically viewed as two independent, mutually exclusive options, rather than complementary means that can coexist within the same system architecture. This fixed perception has led to a long-standing focus on performance optimization within a single technical route, such as improving PTZ response speed to alleviate switching lag or increasing the number of fixed telephoto lenses to expand coverage, without considering the possibility of collaboration between fixed and mobile cameras.

[0053] The inventors recognized that the aforementioned trade-offs did not stem from irreconcilable physical constraints, but rather from the architectural limitations caused by the default assumption that "stability" and "flexibility" were placed within the single responsibility framework of the same camera. If the implicit assumption of "one camera simultaneously undertaking stable output and flexible scheduling" is broken, and the conference room is divided into spatial partitions, allowing a fixed telephoto camera to output stable optical close-ups within its corresponding subspace without rotation, while simultaneously enabling a gimbal camera to undertake flexible scheduling tasks within a specific subspace, then fixed and mobile are no longer mutually exclusive options, but rather two capabilities that can coexist within the same system architecture.

[0054] In view of this, this application proposes a concept of using panoramic images as a unified spatial benchmark, dividing the conference room into multiple subspaces, and scheduling multi-camera collaborative output in different zones.

[0055] This disclosure provides a method for outputting conference video, which can be executed by a processing device. The processing device can be a standalone host, a processing chip within a standalone host, an integrated device, or a processing chip within an integrated device. The integrated device can integrate a processor and a camera. For example, the integrated device can be a video conferencing terminal device.

[0056] like Figure 1 An example flow of the method is shown, including the following steps S101-S104: S101. Acquire the panoramic image of the first space captured by the wide-angle camera, and establish a partition model based on the panoramic image.

[0057] In this model, the partitioning divides the first space into multiple subspaces. For example, refer to... Figure 2 Multiple subspaces include at least a first subspace, a second subspace, and a third subspace.

[0058] For example, the first subspace is the left area of ​​the conference room, and the second subspace is the right area. The left and right areas can also be referred to as the telephoto area or the mid-range area. The third subspace is the far-end area, which can also be referred to as the long-range area.

[0059] In this disclosure, the processing device establishes a global spatial coordinate system using the panoramic image from the wide-angle camera as the global spatial reference. This global spatial coordinate system serves as a shared spatial reference for all camera modules. In this way, different cameras share the same spatial reference, unifying the relationships between people, objects, and space within the conference room into a single reference system, thereby providing better collaborative image capture results.

[0060] As one possible implementation, during the initialization phase, the processing device can divide the conference room into at least three spatial partitions as shown in Table 1. Each partition has a corresponding camera for target tracking within that partition.

[0061] The correspondence between the two telephoto cameras and the subspace can be determined based on the camera arrangement.

[0062] One possible implementation involves two telephoto cameras arranged in a cross-position, with the first telephoto camera corresponding to the second subspace and the second telephoto camera corresponding to the first subspace. Taking the left and right areas as an example, the telephoto camera positioned on the left side of the housing tracks the target in the right area, while the telephoto camera positioned on the right side tracks the target in the left area. With this cross-positioning, the optical axes of the telephoto cameras point to the opposite side, which is beneficial for capturing frontal images of the target.

[0063] As another possible implementation, the two telephoto cameras are arranged at both ends, with the first telephoto camera corresponding to the first subspace and the second telephoto camera corresponding to the second subspace. Taking the left and right areas as examples, the telephoto camera located on the left side of the housing is used to track targets in the left area, and the telephoto camera located on the right side of the housing is used to track targets in the right area. With this arrangement at both ends, the telephoto cameras correspond to the subspaces on the same side, resulting in higher utilization efficiency of the device width.

[0064] Table 1 shows the relevant information for each zone using cross-photography as an example.

[0065] Table 1

[0066] A target tracked by a telephoto camera can be called a mid-range target. A gimbal camera is used for target tracking in the far-field area.

[0067] In traditional multi-telephoto setups, each telephoto camera must capture both mid-range and long-range targets. Therefore, to cover a sufficiently long distance, the telephoto camera's angle of view is forced to narrow, resulting in insufficient lateral coverage at mid-range distances. Furthermore, telephoto cameras are typically prioritized for areas most likely to have people, such as areas with densely packed tables and chairs. When attendees move their chairs backward, stand up, or adjust their posture, even a displacement of tens of centimeters can push them out of the telephoto camera's field of view. In such cases, the system must immediately switch to a wide-angle lens, maintaining the image through digital cropping. Because wide-angle cameras experience pixel degradation when capturing distant targets, distant attendees will perceive a sudden drop in image quality, instantly transforming a clear optical close-up into a blurry, digitally magnified image.

[0068] In this disclosure, by delegating target tracking in distant areas (far-end areas) to the gimbal camera, the burden on the telephoto camera for long-range coverage is reduced. The telephoto camera no longer needs to cover the far-end areas, which means that, under the same field of view, the actual lateral coverage width of the telephoto camera is expanded in the mid-range. The depth margin originally used for long-range tracking is converted into the lateral margin for mid-range tracking.

[0069] Furthermore, the field of view boundary of the telephoto camera is no longer within the high-risk spillover zone of the participants' daily activity range. For example, moving the chair back several tens of centimeters will no longer cause the target to leave the telephoto coverage area, because there is still a safety margin of 1-3 meters between the far boundary of the telephoto camera and a typical mid-range seat. Thus, the dual telephoto cameras disclosed herein can provide more stable and wider effective optical coverage at mid-range distances.

[0070] S102. Within each subspace, determine the primary objective of each subspace.

[0071] Among them, the main target is the important target that needs to be tracked in each subspace.

[0072] As one possible implementation, operations such as target detection (e.g., face detection, human detection), semantic segmentation, speaker recognition, pose estimation, motion vector tracking, and meeting status classification can be performed based on panoramic images acquired by a wide-angle camera, and each detected target can be mapped to a corresponding subspace according to its spatial coordinates.

[0073] For example, an Artificial Intelligence (AI) detection module can identify each target in each panoramic image and output the center coordinates of the human detection bounding box for each target. Then, by using calibration parameters, the pixel coordinates of the targets can be mapped to the global spatial coordinate system, obtaining the global spatial coordinates of the target. The system can then match the global spatial coordinates of the target with the spatial range of each subspace or transition region to assign the target to the corresponding partition or transition region.

[0074] In this disclosure, the determination of the target's zoning is based on the global spatial coordinate system of the panoramic image, rather than a local field of view. This allows all camera modules to share the same spatial reference system, enabling each camera module to understand the target's position within the same spatial reference system, thus reducing the probability of inconsistencies in perception, such as "the telephoto camera perceives the target as being in the left area, while the gimbal camera perceives the target as being in the distance."

[0075] S103. Call the camera corresponding to the subspace to generate close-up shots corresponding to each main target.

[0076] The first subspace and the second subspace correspond to the first telephoto camera and the second telephoto camera, respectively, and the third subspace corresponds to the gimbal camera.

[0077] For example, the first subspace is the left area, the second subspace is the right area, and the third subspace is the far-end area. The correspondence between each subspace and the camera can be found in the explanation of the arrangement of the first and second telephoto cameras above. The far-end area corresponds to the gimbal camera.

[0078] As one possible implementation, the number of subspaces M can be greater than three. For example, when the conference room space is wide, a fourth subspace can be added and a corresponding third telephoto camera can be configured to cover more mid-range areas.

[0079] As another possible implementation, two adjacent subspaces can share the same telephoto camera, which covers the main target in both subspaces by switching between time-sharing modes or expanding the field of view.

[0080] As another possible implementation, the gimbal camera can dynamically adjust its corresponding subspace range. For example, when there is no important target in the far-end area, the gimbal camera can temporarily expand its corresponding subspace range to help cover part of the adjacent subspace.

[0081] It should be noted that the "zoning model" is not limited to explicitly defining geometric boundaries. Any method that establishes a correspondence between targets and cameras based on spatial location to achieve zoned scheduling falls under the specific implementation of a zoning model using the panoramic image as a spatial reference. For example, zones can be implicitly formed by defining the effective coverage area of ​​each camera in the global spatial coordinate system, without needing to predefine explicit boundaries for the subspaces. Zoning models can also naturally form based on the camera's field of view coverage, meaning that the effective shooting area of ​​each camera itself constitutes the corresponding subspace. All of the above implementations fall under the specific implementation of a zoning model using the panoramic image as a spatial reference.

[0082] As another possible implementation, the boundaries of each subspace in the partitioning model can be dynamically adjusted according to the spatial distribution of targets in the panoramic image. For example, when most targets are detected to be clustered in a certain area of ​​the first space, the processing device can dynamically adjust the spatial range of each subspace so that the subspace boundaries adapt to changes in target distribution.

[0083] S104. Output a close-up image of the target according to the preset output strategy.

[0084] The meeting room image output method disclosed herein divides the meeting room space into multiple subspaces by establishing a partitioning model based on the panoramic view of a wide-angle camera, and configuring a corresponding camera for each subspace to generate close-up images of the main target in each subspace. Each camera only needs to focus on the main target within its corresponding subspace, avoiding scheduling conflicts caused by multiple cameras competing for the same target. Furthermore, the panoramic view of the wide-angle camera serves as a unified spatial reference, ensuring spatial consistency in the close-up images output by each camera.

[0085] In some embodiments, determining the main target of each subspace can be implemented as follows: performing attention evaluation on candidate targets in each subspace and determining the main target of each subspace based on the evaluation results.

[0086] In other words, within each spatial partition, attention is independently evaluated for the target within that partition in order to determine the primary target within that partition.

[0087] In this disclosure, the primary target is determined by independently evaluating attention within a subspace, so that the camera scheduling decision of each subspace depends on the attention evaluation of the target within that subspace, without cross-space competition, thereby reducing scheduling conflicts.

[0088] In some embodiments, the step of evaluating the attention of candidate targets within each subspace and determining the main target of each subspace based on the evaluation results can be implemented as follows: obtaining multi-dimensional feature information of each candidate target within the subspace, determining the attention weight of each candidate target based on the multi-dimensional feature information, and determining the candidate target with the highest attention weight within the subspace as the main target.

[0089] The multidimensional feature information includes one or more of the following: speaking status, movement status, spatial location, posture, dialogue history, meeting role label, or electronic device usage status.

[0090] Meeting role labels include, for example, moderator / speaker / participant. Electronic device usage status includes, for example, whether the electronic device is being used.

[0091] Each subspace maintains its own attention pool, which contains only candidate targets within that subspace. The primary target of each subspace is the candidate target with the highest attention weight in that subspace.

[0092] For example, the attention pool for the left zone (denoted as Zone_L) contains only candidate targets within that zone. The processing unit sorts the attention weights of each candidate target within the left zone, and the candidate target with the highest attention weight becomes the primary target. The telephoto camera responsible for that area locks onto and tracks the primary target. The attention pool can also be referred to as the candidate target set.

[0093] The right zone (Zone_R) attention pool contains only candidate targets within the right zone. The telephoto camera responsible for tracking this area tracks the candidate target with the highest attention weight (i.e., the main target in the right zone).

[0094] The far-field zone (Zone_F) attention pool contains only candidate targets within that far-field zone. The gimbal camera tracks the candidate target with the highest attention weight within that far-field zone.

[0095] In summary, based on the calculation of the regional attention pool, under normal circumstances (no cross-regional events), each camera can operate according to Table 2: Table 2

[0096] As shown in Table 2, dual telephoto cameras are fixedly installed, each locking onto the target with the highest attention weight within its corresponding subspace for stable mid-range optical tracking. Under normal circumstances where the target does not move across subspaces, the gimbal camera is dedicated to high-definition capture in the far-end area, such as locking onto the main target in a key position at a distance, and does not participate in routine speaker tracking in the mid-range area. Each camera module independently decides to track targets within its corresponding subspace's attention pool, without generating cross-area scheduling competition.

[0097] As one possible implementation, when multiple targets with similar attention weights exist within a fixed telephoto zone (the difference is less than a threshold), a sub-region rotation mechanism can be activated. The image captured by the fixed telephoto lens can be divided into multiple sub-regions, and the fixed telephoto lens can sequentially output close-ups of each sub-region, or output an additional rotating picture-in-picture sequence in addition to the main close-up.

[0098] As one possible implementation, attention weights can be adjusted based on audio localization. For example, by detecting the direction of a sound source using a microphone array and mapping that direction to spatial coordinates in a panoramic view, attention weights can be adjusted for targets located near the sound source. This could involve increasing or decreasing the weights to achieve cross-validation between audio and video detection.

[0099] As one possible approach, changes in the location of the sound source can also trigger the camera to pre-rotate, thus enabling early prediction using audio localization to compensate for the response delay of visual tracking.

[0100] The solution disclosed herein confines the attention pool of each camera module to its spatial range. For example, a telephoto camera responsible for the left area does not concern itself with people in the right area, and a gimbal camera normally does not concern itself with people at mid-range distances, thus reducing competition between areas. Furthermore, each attention pool is computed independently and can be executed in parallel. The number of candidate targets in a single attention pool is less than the total number of targets in the entire scene. This makes the calculation of attention weights more efficient.

[0101] Furthermore, target changes within a subspace (such as a speaker in the left area changing from A to B) only affect the tracking switch of the corresponding camera and do not affect cameras in other subspaces, thus reducing the impact of the cascading effect of "target causing system-wide rescheduling".

[0102] Furthermore, the independent calculation of the attention pool in the far region ensures that key positions at the far end always receive focused visual coverage, unaffected by dynamic interference from the mid-range subspace.

[0103] In addition, attention weighting integrates multi-dimensional features such as speaking status, movement status, spatial location, posture, and dialogue history, which can accurately identify the participants in the meeting scene who are most worthy of being shown in close-up, making the switching of close-up shots more in line with the actual attention needs of meeting participants, and improving the accuracy of intelligent framing and user experience.

[0104] As one possible implementation, the processing device can weight the aforementioned multidimensional feature information to obtain the attention weights for each subspace.

[0105] For example, the feature information of each dimension in the multidimensional feature information can be quantified into a certain value, and the value range of the feature information of each dimension is shown in Table 3: Table 3

[0106] For example, the speaking status can be determined based on the sound source azimuth and duration estimated by the microphone array's DOA. The current speaker receives the highest score, the score for those who have just finished speaking decays exponentially (half-life of approximately 3 seconds), and the score for those who remain silent approaches zero.

[0107] For example, the motion state can be determined based on the magnitude and direction of the target motion vector. Large movements (such as standing up or walking towards the whiteboard) earn high scores. Stillness or small movements (such as taking notes) earn low scores.

[0108] For example, spatial location can be determined based on the target's relative position within its subspace. The central position within the subspace receives higher weight.

[0109] For example, posture can be identified based on human posture estimation. For instance, gestures (such as raising a hand), standing up, pointing, face orientation, gaze direction, head posture, body orientation, and other body movements with meeting semantics can be given higher weights.

[0110] As one possible implementation, among the multidimensional feature information, gestures, gaze direction, head posture, and body orientation can be used as feature dimensions independent of posture in the calculation of attention weights. Specifically, gestures include hand movements such as raising, pointing, and waving, and are evaluated independently of the overall body posture; gaze direction is the target's line of sight, evaluated independently of the face orientation; head posture is the target's head pitch and yaw angles; and body orientation is the direction the target's torso faces. These feature dimensions can supplement or replace speaking state, motion state, spatial location, posture, and dialogue history, and can be used individually or in combination to determine the attention weight of the candidate target.

[0111] For example, dialogue history can be based on statistics of the target's recent posts, with high-frequency active users receiving higher weight.

[0112] For example, the attention weight W_total = α·W_voice + β·W_motion + γ·W_position + δ·W_gesture + ε·W_history. Where α, β, γ, δ, and ε are the weight coefficients of each feature.

[0113] For example, the weight coefficients α, β, γ, δ, and ε of each feature are dynamically adjusted based on the current meeting mode. For instance, the weight coefficients can be adjusted with reference to Table 4: Table 4

[0114] The above are merely exemplary calculation formulas and do not constitute a limitation on the calculation method of attention weights. In practical applications, one or more feature dimensions can be selected to participate in the calculation according to the meeting scenario.

[0115] In some embodiments, the processing device may receive a manual selection instruction to lock a specified target as the current output target. During the locking period, automatic attention evaluation and cross-regional scheduling for that target are suspended. A mechanism is in place to resume automatic mode in response to an unlock instruction.

[0116] As one possible approach, during manual locking, the system can continue to update the target's attention weights and location information to achieve a seamless transition when unlocking.

[0117] In some embodiments, the preset output strategy includes at least one of the following strategies: First strategy: In response to the target moving across subspace, switch the camera used to generate close-up shots of the target, and continuously output close-up shots of the target. This first strategy improves the continuity of close-up shots of the target during cross-subspace movement.

[0118] The second strategy is to output multiple close-up shots corresponding to each main target as the target close-up shots in parallel. For example, the main target in the left area is target A, the main target in the right area is target B, and the main target in the far area is target C. The processing device controls the parallel output of three close-up shots, namely the close-up shots of target AC.

[0119] The third strategy is to combine the multiple close-up shots corresponding to each main target into a single combined shot and output it as the target's close-up shot. For example, if the main target in the left area is target A, the main target in the right area is target B, and the main target in the far area is target C, the processing device will combine the three close-up shots, i.e., the close-up shots of target A and B, into a single combined shot and output the combined shot.

[0120] The second and third strategies support multi-channel parallel output and combined screen output, respectively, to adapt to different display devices and viewing habits.

[0121] The fourth strategy involves comparing attention across subspaces for each primary target and selecting one close-up shot as the target's close-up shot for output. This fourth strategy supports attention comparison across subspaces and is applicable to scenarios requiring global focus on a single, most important speaker.

[0122] Fifth strategy: Pause the output of close-up shots of human subjects and switch to outputting close-up shots of non-human subject areas. The fifth strategy supports close-ups of non-human subjects (such as whiteboards and projection screens), enriching the information presentation dimensions of video conferencing and improving the system's scene adaptability.

[0123] As one possible implementation, the processing device can automatically identify the meeting type and match the corresponding strategy based on scene semantic understanding. For example, by analyzing the number of participants, seating arrangement, speaking frequency, interaction mode, etc., it can automatically identify whether the current meeting is a lecture, roundtable discussion, group discussion, or remote defense, and automatically match the best strategy.

[0124] The solution disclosed herein provides a variety of output strategies, including cross-regional tracking relay, multi-channel parallel output, combined image synthesis, cross-regional attention comparison, and non-human target turning output. This allows the system to flexibly select or combine applicable output strategies based on the actual needs of the meeting scenario. These strategies are designed around a partitioned architecture with a wide-angle camera panoramic view as the unified spatial benchmark. Each strategy addresses the image output requirements of different meeting scenarios, enabling the same multi-camera hardware architecture to adapt to various meeting modes such as seminars, presentations, speeches, and discussions. It outputs close-up images of the target that suit the current scenario without requiring manual switching by the user, thus improving the system's intelligence and scene adaptability.

[0125] In some embodiments, the first strategy includes: when a first target is detected moving from a first subspace to a third subspace, controlling the first telephoto camera to continuously generate close-up images of the first target until the first target moves out of the field of view of the first telephoto camera, and / or controlling the gimbal camera to take over generating close-up images of the first target when the first target moves out of the field of view of the first telephoto camera, and / or controlling the gimbal camera to take over generating close-up images of the first target when it enters the field of view of the gimbal camera.

[0126] As one possible approach, cross-regional pick-up and drop-off can be triggered if the target is detected to be moving across regions based on the panoramic view from the wide-angle camera.

[0127] For example, when a target is detected moving from the left area to the far area, the processing device controls the telephoto camera responsible for the left area to continuously generate close-up images of the target, and controls the gimbal camera to pre-rotate towards the predicted target position.

[0128] As one possible approach, the target enters the field of view of the gimbal camera while still being within the field of view of the telephoto camera, and the image switching is completed during the period when the fields of view of the two cameras overlap.

[0129] As another possible implementation, the target moves out of the telephoto camera's field of view and then enters the gimbal camera's field of view. When the target enters the gimbal camera's field of view, the camera switches to a close-up shot of the target generated by the gimbal camera.

[0130] For example, the priority and scenarios for gimbal camera scheduling are shown in Table 5: Table 5

[0131] As shown in Table 5, under normal conditions (corresponding to the highest priority), the PTZ camera remains stationary in the primary position of the remote area to track important targets within that area. For medium priority, when a target is moving across areas and there are no important targets in the remote area, the PTZ camera performs task reception or delivery. After completing the task, it can automatically return to its preset position in the remote area. For low priority, when the target moving across areas is the only globally important speaker and there are no important targets in the remote area, the PTZ camera can follow the target out of the remote area and track that important target throughout the entire process.

[0132] In some embodiments, the first strategy further includes: during the process of switching between the first telephoto camera and the gimbal camera to generate a close-up image of the first target, controlling the first telephoto camera to output a transition image of the first target to the edge of the field of view of the first telephoto camera by moving the cropping window in the image it captures.

[0133] During the transition between the first telephoto camera and the gimbal camera, the first telephoto camera continues to capture the image, and the cropping window moves within that image to output a transitional view, ensuring a smooth transition to the gimbal camera's position. Once the gimbal camera is in place, the processing unit can seamlessly switch the image from this transitional view to the optical view of the gimbal camera.

[0134] The disclosed solution achieves a smooth transition when the target moves between the telephoto camera and the gimbal camera through the camera handover mechanism in the first strategy. Before the target moves out of the telephoto field of view, the telephoto camera continuously tracks and generates close-up images, and generates transition images by moving the cropping window. This reduces the image jumps caused by target loss or hard switching, allowing remote participants to continuously and stably view close-up images of the target person, thus improving the viewing experience in cross-regional mobile scenarios.

[0135] In some embodiments, the second strategy includes: Obtain the attention weights of each main target in its respective subspace; The attention weights of each main target are sorted, and close-up shots that meet the preset conditions are selected according to the sorting results and output in parallel.

[0136] In some embodiments, the second strategy can be implemented as follows: obtaining the attention weight of each main target in its respective subspace; uniformly sorting the attention weight of each main target across subspaces; and selecting close-up shots that meet the preset conditions according to the sorting results for parallel output.

[0137] It should be noted that the above ranking is not limited to directly comparing the absolute values ​​of attention weights between subspaces.

[0138] As one possible approach, attention weights can be normalized within each subspace first, and then the normalized weight values ​​can be compared across subspaces to eliminate the influence of differences in target density within different subspaces on the absolute value of the weights.

[0139] As another possible implementation, each subspace can reserve at least one parallel output slot to ensure that each subspace has a chance to be rendered.

[0140] The disclosed solution, through the parallel output mechanism based on attention weight ranking in the second strategy, enables the system to automatically identify and simultaneously present several key objectives of the meeting that are of utmost importance. Remote participants can simultaneously obtain close-up shots of multiple important speakers without manually switching screens, effectively improving information acquisition efficiency in scenarios such as multi-party discussions and multiple speakers taking turns speaking, and reducing attention distraction caused by frequent screen switching.

[0141] In some embodiments, the preset conditions include any of the following: The top N positions, where N is a preset positive integer; The attention weight is greater than the preset threshold.

[0142] For example, the main targets in the three subspaces are target A, target B, and target C. If the preset condition is that they are ranked in the top N and N=2, then the close-up shots of the top two targets (such as A and C) after sorting by attention weight will be output in parallel, while the third-ranked target B will not be output. If N=1, only one output will be output; if N=3, all three outputs will be output. N can be a preset positive integer, or it can be determined according to the number of screens on the display device or user configuration.

[0143] By further limiting the selection quantity in the second strategy—constraining the number of targets displayed side-by-side by taking the values ​​of the top N or exceeding a preset attention weight threshold—screen fragmentation and information overload caused by indiscriminately displaying all candidate targets side-by-side are avoided. The quantity cap mechanism (top N) allows the number of targets displayed side-by-side to dynamically adapt to the screen size, resolution, or viewing preferences of the display device, maintaining a clear and readable screen layout while ensuring that key speakers are not overlooked. The weight threshold mechanism (exceeding a preset threshold) further filters out temporary low-activity targets—such as participants who briefly interrupt and then fall silent—retaining only speakers whose attention weight is reorganized to gain sustained attention value, thus improving the information density of the side-by-side screen and the stability of the viewing experience.

[0144] In some embodiments, the third strategy includes: Obtain the attention weights of each main target in its respective subspace; The attention weights of each main target are sorted, and the arrangement position and / or display size of each close-up shot in the combined shot are determined according to the sorting result.

[0145] For example, the main target in the left area is target A, the main target in the right area is target B, and the main target in the far area is target C. Targets A and C have higher attention weights, with target A having a greater attention weight than target C. The processing device can combine these two close-up shots of targets A and C into a single combined image. In this combined image, target A is positioned first and has a larger display size.

[0146] In the third strategy of this disclosed solution, the arrangement and / or display size of each close-up shot in the combined screen are determined according to the attention weight ranking, so that the layout of the combined screen can dynamically reflect the importance of each main target. Main targets with higher attention weights receive more prominent display positions and larger screen sizes, while main targets with lower weights are correspondingly reduced in size or placed in secondary positions, enabling remote participants to clearly grasp the importance distribution of each speaker in the meeting and improving the information transmission efficiency of the combined screen.

[0147] In some embodiments, the fourth strategy includes: Obtain the attention weights of each main target in its respective subspace; The attention weights of each main target are compared across subspaces, and the close-up image corresponding to the main target with the highest attention weight is selected as the target close-up image for output.

[0148] For example, the primary target in the left area is target A, the primary target in the right area is target B, and the primary target in the far area is target C. Among these, target A has the highest attention weight. The processing device can output a close-up image of target A as a target close-up image.

[0149] In some embodiments, the fourth strategy further includes: When the attention weight of the selected main target is lower than a preset threshold, the current output screen remains unchanged.

[0150] For example, although target A has the highest attention weight among all main targets, its attention weight is less than a preset threshold. In this case, the current close-up shot remains unchanged, and the close-up shot of target A is not switched to for the time being.

[0151] As one possible approach, setting an attention weight threshold can avoid frequent camera cuts when all targets in the room have low activity levels. For example, after the speaker finishes speaking, the entire room enters a silent phase for note-taking or reviewing materials. In this case, the attention weights of candidate targets within each subspace are generally low and similar. Even if the selected target has the highest attention weight, its weight is below the preset threshold, indicating that there is no substantial need for speaking or interaction. Maintaining the current view in this situation avoids frequent switching between multiple inactive targets due to minor fluctuations in attention weights, thus improving the stability of the output view. Optionally, after a certain period, the output can switch to a panoramic view to present the overall state of the meeting room.

[0152] The disclosed solution, through the attention weight comparison mechanism across subspaces in the fourth strategy, enables the system to identify the most important speaker among the main targets in multiple subspaces and provide close-up output. Simultaneously, through an attention weight threshold mechanism, it avoids frequent camera switching when all targets are at low activity levels. This strategy is adaptable to scenarios requiring focus on a single speaker, such as keynote speeches or thematic presentations, as well as scenarios requiring dynamic switching, such as multiple speakers taking turns or free discussions, balancing focus capability and output stability.

[0153] In some embodiments, the fifth strategy includes: performing target recognition on the panoramic image; and when a non-human target area is identified and a preset trigger condition is met, controlling the gimbal camera to rotate toward the non-human target area to generate a close-up image of the non-human target area. The non-human target area includes at least one of a whiteboard area, a projection screen area, or a physical display stand area.

[0154] For example, in presentation or seminar mode, the processing device can identify high-attention areas such as whiteboards and projection screens, and the PTZ camera is directed to point at that area to output high-definition close-ups so that remote participants can clearly identify the content in that area. After the task is completed, the PTZ camera can automatically return to its dominant position in the remote area.

[0155] The solution disclosed herein, through the mechanism in the fifth strategy of identifying non-human target areas and triggering the rotation of the PTZ camera, enables the video conferencing system to automatically capture non-human visual information in the meeting, such as whiteboard writing, projected presentations, and physical object displays, thereby expanding the information collection dimensions of video conferencing.

[0156] In some embodiments, the preset triggering condition includes: detecting content updates in the non-human target area.

[0157] The solution disclosed herein reduces unnecessary camera movement by detecting content updates in non-human target areas as a trigger condition. Close-up shots are only output when the content being presented changes, thereby improving the system's scheduling efficiency and the information density of the image.

[0158] In some embodiments, the preset output strategy further includes: combining at least two basic strategies from the first strategy to the fifth strategy to form a combined strategy to output the target close-up image.

[0159] As one possible implementation, different combinations of basic strategies can achieve different screen output capabilities. For example, combining the first strategy of cross-area tracking with the second strategy of parallel output can simultaneously display the screen of other subspaces while tracking a moving target. Combining the fourth strategy of cross-area attention comparison with the fifth strategy of non-human target output can focus on the speaker while also presenting whiteboard or screen content. The processing device can flexibly combine basic strategies to adapt to different output goals according to the current meeting needs and screen layout requirements.

[0160] In some embodiments, the preset output strategy further includes: obtaining the current meeting mode, and determining a target strategy from the basic strategy or the combined strategy based on the meeting mode.

[0161] As one possible implementation, the meeting mode can be manually selected by the user or automatically identified by the processing device based on information such as the number of participants, seating distribution, and speaking frequency in the panoramic view.

[0162] Different meeting modes can correspond to different target strategies. For example, in seminar mode, multiple participants take turns speaking, and the processing device can prioritize the second strategy to output close-up shots of the main target in each subspace in parallel, so that remote participants can see multiple discussants simultaneously. In presentation mode, the presenter uses a whiteboard or projection screen to explain, and the processing device can use a combination of the second and fifth strategies to output close-up shots of the presenter and the whiteboard / projection screen in parallel. In key speaker mode, a single speaker speaks continuously, and the processing device can use the fourth strategy to focus on the close-up shot of the key speaker, while reducing unnecessary camera cuts caused by the activity of targets in other subspaces. In free discussion mode, the speaking rights may frequently switch between multiple participants, and the processing device can combine the first and fourth strategies to maintain tracking continuity when the target moves across areas, while dynamically selecting the most worthy close-up shot to focus on based on cross-area attention comparison.

[0163] As one possible implementation, the output strategy can be dynamically determined based on real-time information. For example, the processing device can calculate the optimal output strategy in real time based on real-time information such as the current attention weight distribution of each main target, the number of targets, and the duration of the meeting, rather than selecting from a preset set of strategies.

[0164] The disclosed solution, through the combination of basic strategies and a strategy selection mechanism based on meeting modes, enables the system to dynamically adapt the optimal combination of output strategies according to the actual meeting scenario. The same hardware architecture can adapt to various meeting modes such as seminars, presentations, speeches, and discussions without any changes. It automatically outputs visuals suitable for the current scenario without requiring manual configuration by the user, improving the system's intelligence and ease of use, and reducing the user's operational burden.

[0165] In some embodiments, the method may further include: the processing device maintaining a target state record for each target detected in the panoramic image, the target state record including one or more of the following information: target identifier, currently assigned subspace, real-time coordinates in the first space, motion vector, and timestamp of entering the currently assigned subspace. The currently assigned subspace is determined by matching the real-time coordinates with the spatial range of each subspace.

[0166] For example, the Target State Record (TSR) and its key fields are as follows: TSR = { target_id: Target identifier spatial_zone: The current subspace to which the user belongs ∈ {Zone_L, Zone_R, Zone_F, TZ_LF, TZ_RF}. global_coord: Real-time 3D coordinates (x, y, z) in the global spatial coordinate system of the wide-angle panorama. motion_vector: Motion vector (vx, vy, vz) zone_entry_time: The timestamp of entering the current subspace. zone_confidence: Subspace ownership confidence (0-1) } The subspace attribution confidence score is determined based on the distance from the target's current coordinates to the nearest subspace boundary; the closer the distance, the lower the confidence score. When the confidence score falls below a preset threshold, the target can be determined to have entered a transition region.

[0167] TZ_LF represents the transition region between the left and far regions, and TZ_RF represents the transition region between the right and far regions. Subspace affiliation confidence indicates the confidence that the target belongs to the current subspace. When the target crosses regions, the value of spatial_zone changes accordingly.

[0168] The TSR can include more or fewer fields, without limitation. For example, it can also include: the camera currently tracking the target.

[0169] For example, after the meeting begins, each detected target is assigned a unique ID, its initial spatial location is mapped to the corresponding subspace, and the TSR is initialized. The TSR can be updated as the state changes subsequently. For example, if a target moves, its current belonging subspace is updated.

[0170] The disclosed solution provides a data foundation for subspace attribution determination and cross-regional dispatch scheduling by maintaining a target status record for each target in the panoramic image, including subspace attribution, real-time coordinates, and motion vectors. The real-time updating of the target status record enables the system to accurately track the motion trajectory and subspace attribution changes of each target, upgrading the zoning model from static spatial division to dynamic target-subspace association management, and providing an accurate and reliable information source for camera scheduling decisions.

[0171] In some embodiments, during the period when the target is occluded, the processing device can predict the target's reappearance location based on its motion vector and velocity. After the occlusion is removed, target re-identification (Re-ID) can be performed using facial recognition or clothing features to match the historical state record with the newly appearing target, thus avoiding misidentification of the reappearing target after occlusion as a new target.

[0172] In some embodiments, within any of the subspaces, when the main target switches from an old target to a new target, generating a close-up shot corresponding to each main target can be achieved by controlling the corresponding camera to move the cropping window within its captured image by a preset movement step size, thereby generating a close-up shot transitioning from the old target to the new target. The starting position of the cropping window is the location of the old target, and the ending position is the location of the new target.

[0173] The attention weight of the new target is greater than the attention weight of the old target, and the difference between the two exceeds a preset weight threshold. This weight threshold can also be called the hysteresis threshold.

[0174] In other words, when the main target in a certain subspace changes (such as the speaker in the left zone changing from A to B), the camera module corresponding to that subspace performs an intra-zone focus handoff.

[0175] For example, if the attention weight of a new target in the right area exceeds that of the currently tracked target (old target) in the right area, and the difference between their attention weights exceeds a weight threshold, then the new target is designated as the new primary target in the right area. The right telephoto camera in the right area performs a smooth tracking transition, achieving a smooth transition from the old target to the new target through the gradual movement of the digital cropping window.

[0176] The disclosed solution achieves a smooth transition between different targets within the same camera's field of view by moving the cropping window with a preset movement step size when the main target switches within the subspace. The cropping window gradually moves from the old target position to the new target position, resulting in a continuous scrolling visual effect in the output image and reducing the visual abruptness caused by hard cuts. Furthermore, by setting a weight threshold to trigger the switch, frequent ping-pong switches caused by slight fluctuations in attention weights are reduced, improving the stability and anti-interference capability of target switching. In addition, when switching tracking targets, the gimbal camera naturally involves image movement due to physical rotation, while the telephoto camera is fixed in place. If the target is switched directly without a cropping window transition, the output image will show an instantaneous jump. By configuring the same target switching transition method with a moving cropping window for the telephoto camera, the switching process of different camera modules can be visually consistent, thus providing a unified image switching experience.

[0177] As one possible implementation, a telephoto camera can directly switch from a close-up of the old target to a close-up of the new target during target switching, without going through a cropping window transition. In this case, a gimbal camera also uses electronic cropping to switch directly during target switching, and the two cameras present the same image during target switching.

[0178] In some embodiments, adjacent subspaces within the plurality of subspaces have a transition zone for cross-zone transfer. This transition zone, also referred to as a transition area, serves as a field-of-view overlap buffer between adjacent subspaces. Within the transition zone, cameras in adjacent subspaces can output valid images to ensure a smooth transition. For example, when dual sources are available simultaneously, the processing device can perform alpha blending switching of the images within the transition zone, making the transition even smoother.

[0179] For example, the left transition region (TZ_LF) is the transition region between the left region and the far region, for example, with a width of about 1-3m. The right transition region (TZ_RF) is the transition region between the right region and the far region, for example, with a width of about 1-3m.

[0180] Transition zones are one means of implementing pick-up and drop-off mechanisms, not the only one. For example, algorithms for determining switching points based on wide-angle coordinate prediction can also support pick-up and drop-off mechanisms.

[0181] In some embodiments, the method may further include performing cross-regional pick-up and drop-off scheduling when a target is detected moving from its currently assigned source subspace to an adjacent target subspace via the transition region. For example, cross-regional pick-up and drop-off scheduling is performed when a target is detected moving from the left region to an adjacent far region via the transition region.

[0182] As one possible implementation, cross-regional transport assessment is triggered when the spatial coordinates of the tracked target enter the transition region and its motion vector points to the adjacent subspace.

[0183] For example, if a target being tracked by the left telephoto camera moves from Zone_L into TZ_LF and its motion vector points towards the far zone (moving towards the far zone), it may trigger a cross-zone "connection". A cross-zone "connection" means that the target moves from the telephoto zone into the far zone, and the gimbal camera may take over target tracking from the telephoto camera.

[0184] The target moves from Zone_F into TZ_LF, and its motion vector points to the left zone (moving towards the mid-range), which may trigger a cross-zone "transfer". A cross-zone "transfer" means that the target moves from the far zone into the telephoto zone, and the gimbal camera may hand over the target to the telephoto camera for tracking.

[0185] The target enters TZ_RF from Zone_R and the motion vector points to the far zone, which may trigger a cross-zone "connection".

[0186] The target enters TZ_RF from Zone_F, and the motion vector points to the right zone, which may trigger a cross-zone "transmission".

[0187] Performing cross-regional pick-up and drop-off scheduling can be achieved by predicting the target's location when it enters the target subspace based on the panoramic image. For example, based on the target in the panoramic image, the target's movement trend or trajectory can be determined, and thus the target's location when it moves into the remote area can be predicted (e.g., predicting the target location based on motion vectors).

[0188] In some scenarios, if the attention weights of all other targets in the target subspace are less than or equal to the attention weight of the target, then the camera corresponding to the target subspace is controlled to pre-rotate toward the target position or take over to generate a close-up image of the target after the target enters its field of view.

[0189] For example, as the current target moves from the left area to the far area, the processing device checks the current state of the attention pool in the far area to determine if there are any targets with higher attention weights that require continuous coverage by the gimbal camera. If there are no targets with higher attention weights in the far area—for example, if the attention weights of other targets in the far area's attention pool are all lower than the current target's, or if the pool is empty—the gimbal camera is released to perform the "takeover" task. The gimbal camera can pre-rotate (pre-aim) towards the predicted target position. Thus, when the target enters the effective field of view of the gimbal camera, the gimbal camera can complete the takeover of the view.

[0190] As another example, the current target moves from the far zone to the left zone, and there are no other targets in the left zone with a higher attention weight than the current target. In this case, after the target enters the field of view of the left telephoto camera in the left zone, the left telephoto camera can take over and generate a close-up image of the target.

[0191] In some scenarios, if there is a high-attention target being tracked in the target subspace, the camera corresponding to the target subspace is controlled to adjust its field of view to simultaneously cover the target and the high-attention target, generating a frame-to-frame or single-person close-up shot of the target and the high-attention target.

[0192] High attention targets can refer to targets whose attention weight is higher than or equal to a threshold.

[0193] For example, the current target moves from the left area to the far area. If there is a higher attention target in the far area (such as a key person speaking), the gimbal camera in the far area adjusts its field of view to generate a close-up shot that includes the current target and the high attention target in the same frame.

[0194] As another possible implementation, the gimbal camera captures an image containing both the current target and the high-attention target with a wide field of view, and generates an independent single-person close-up image for each of the current target and the high-attention target through electronic cropping.

[0195] In this embodiment, by setting a transition area between adjacent subspaces and predicting the target's entry position into the target subspace based on the panoramic view, precise triggering and predictive execution of cross-regional pick-up and drop-off scheduling are achieved. When there is no target with higher attention weight in the target subspace, the camera corresponding to the target subspace rotates to the predicted position in advance or takes over to generate a close-up image after the target enters its field of view, reducing image loss during the target's cross-regional movement. When there is a target with high attention weight being output in the target subspace, the camera adjusts its field of view to cover both targets simultaneously, generating a close-up image of the same person or a single person, ensuring continuous output of high attention targets and achieving synchronous presentation of cross-regional targets.

[0196] In some embodiments, when the camera corresponding to the source subspace is a fixed telephoto camera and the camera corresponding to the target subspace is a gimbal camera: The control of the camera corresponding to the target subspace to pre-rotate toward the target position can be implemented by controlling the gimbal camera to rotate toward the target position.

[0197] The step of taking over the generation of a close-up image of the target after the target enters its field of view can be implemented as follows: during the rotation of the gimbal camera, the fixed telephoto camera is controlled to continuously output a close-up image of the target by moving the cropping window in the image it captures until the target moves out of the field of view of the fixed telephoto camera, or until the gimbal camera completes its rotation.

[0198] For example, the current target moves from the left area to the far area. If there is no target with higher attention weight in the far area, the gimbal camera can be released to perform the next task. The gimbal camera in the far area can pre-rotate towards the target position based on the predicted target position. During the rotation of the gimbal camera, the processing device controls the left telephoto camera in the left area to generate a virtual smooth trajectory from the current position to the target position, and continuously outputs close-up images of the target by moving a cropping window in its captured image, until the target moves out of the field of view of the left telephoto camera, or until the gimbal camera completes its rotation. In this way, the output image can present the visual effect of "smoothly moving towards the target position". After the gimbal camera is in position, the left telephoto camera releases the target and returns to the left area to track the second highest target in the left area's attention pool.

[0199] The trajectory parameters include: the starting coordinates of the cropping window (the coordinates of the current position), the target coordinates (the coordinates of the target position), and the movement speed curve. The movement speed curve, for example, is a gradual in / out curve to reduce visual abruptness. The trajectory calculation is based on a global spatial coordinate system established from the panoramic image to ensure that the spatial orientation in the image is consistent with physical space.

[0200] For example, while the telephoto camera performs a digital cropping transition, the gimbal camera physically rotates at a rate lower than its maximum rotation speed to reduce noise and mechanical shock. Since the digital cropping has already completed the visual movement of the image, the physical rotation time window of the gimbal camera is hidden within the smooth transition process. The user sees a smooth panning of the image and is unaware that the gimbal camera is rotating.

[0201] For example, if the current target moves from the left area to the far area, and there is a target with higher attention weight in the far area (such as a key person speaking), the gimbal camera will refuse to perform the follow-up task. In this case, the telephoto camera continues to track to the edge of the field of view, and then the panoramic view from the wide-angle camera is digitally cropped to maintain the output image.

[0202] After the target enters the field of view of the gimbal camera, the system switches to an optical close-up image of the target generated by the gimbal camera.

[0203] In other words, when the target enters the effective field of view of the gimbal camera, the gimbal camera can take over the view. For example, once the gimbal camera is physically positioned and the image is stable, the system can gradually switch the output source from the digitally cropped image to the optical image of the gimbal camera within 2-3 frames, such as a linear transition of alpha from 1.0 to 0.0. Since the gimbal camera is now directly facing the target area, the switched image is consistent with the transition image in terms of composition, brightness, and color, and the user is unaware of it.

[0204] In this embodiment, for scenarios where the target moves across zones from the subspace corresponding to the fixed telephoto camera to the subspace corresponding to the gimbal camera, the fixed telephoto camera continuously outputs a close-up of the target through a moving cropping window during the rotation of the gimbal camera, thus covering the mechanical rotation delay of the gimbal camera. After the target enters the field of view of the gimbal camera, the output switches to an optical close-up image of the target generated by the gimbal camera. This achieves a smooth transition from telephoto digital tracking to physical pre-rotation of the gimbal camera to digital fine-tuning of the gimbal camera to optically stable output, reducing image jumps during cross-module switching. This allows the remote user to perceive a smooth camera movement similar to that of a professional cinematographer.

[0205] In some embodiments, when the camera corresponding to the source subspace is a first telephoto camera and the camera corresponding to the target subspace is a second telephoto camera, the cross-regional pick-up and drop-off scheduling can be implemented as follows: The first telephoto camera is controlled to continuously output close-up images of the target by moving a cropping window within its captured image until the target moves out of the first telephoto camera's field of view. After the target enters the field of view of the second telephoto camera, the system switches to generating close-up images of the target using the second telephoto camera. Within a preset time period after the switch, the second telephoto camera is controlled to continuously output close-up images of the target located at the center of the image by moving a cropping window within its captured image.

[0206] As one possible implementation, taking the movement of a target from the left zone to the right zone as an example, when the processing device detects that the target is moving to the right zone via a transition area, it controls the telephoto camera responsible for the left zone to continuously output close-up images of the target. As the target approaches the boundary of the left zone, the telephoto camera responsible for the left zone moves its cropping window within its captured image, following the target's real-time position to keep the target centered in the output image until it moves out of the field of view. After the target enters the right zone and comes into the field of view of the telephoto camera responsible for the right zone, the image output source is switched to the telephoto camera responsible for the right zone. After the switch is complete, the telephoto camera responsible for the right zone moves its cropping window within its captured image, positioning the cropping area to the target's location to maintain the target's position and size in the image consistent before and after the switch.

[0207] In some embodiments, the method of moving the cropping window includes: determining the cropping position of the target in the image captured by the corresponding camera based on the real-time position of the target in the panoramic view; and moving the cropping window from the starting position to the ending position with a preset movement step size to keep the position of the target continuous in the output image.

[0208] In this embodiment, a lightweight handover process is provided for scenarios where a target moves across subspaces corresponding to two fixed telephoto cameras. Since both telephoto cameras are fixedly mounted, there is no need for mechanical rotation or waiting. The source telephoto camera tracks the target to the edge of its field of view by moving the cropping window. Once the target enters the target telephoto camera's field of view, the system immediately switches, and the target telephoto camera maintains the target in the center of the image by moving the cropping window. This solution utilizes the digital movement of the cropping window to cover the image shift during the transition between fixed telephoto cameras, improving the continuity and smoothness of cross-region tracking between adjacent fixed telephoto cameras. The implementation is simple and provides stable image quality. It can be applied to target tracking and switching between different seats within a mid-range area.

[0209] In some embodiments, the execution of cross-regional pick-up and drop-off scheduling further includes: when a target is detected moving from the third subspace to the first subspace via the transition region: The processing device controls the gimbal camera to continuously track the target. For example, when a target is detected moving from the far region through the transition region to the left region, the processing device controls the gimbal camera to continuously track the target, even if the target has left the current field of view of the gimbal camera.

[0210] And / or, after confirming from the panoramic view that the target has entered the first subspace and that the first telephoto camera has captured the target, the system switches to generating a close-up view of the target using the first telephoto camera. For example, when a target is detected moving from the far-end area through the transition area to the left area, the gimbal camera tracks the target to the telephoto area. After confirming that the left telephoto camera has stably captured the target's image, the processing device controls the system to switch to tracking the target using the left telephoto camera.

[0211] When a target is detected moving from the third subspace to the second subspace via the transition region: Control the gimbal camera to continuously track the target, and / or, after confirming from the panoramic view that the target has entered the second subspace and the second telephoto camera has captured the target, switch to a close-up view of the target generated by the second telephoto camera.

[0212] For example, the processing device switches the output image from the gimbal camera in the third subspace to the corresponding telephoto camera, and a smooth transition can be performed during the switching process. That is, the gimbal camera outputs the transition image of the target to the edge of the telephoto camera's field of view by moving the cropping window in its captured image. After the gimbal camera releases the target, it can automatically return to the preset position in the far-end area and continue to perform the normal tasks of the far-end area's attention pool.

[0213] In this embodiment, a return trip solution is provided for scenarios where the target returns from the subspace corresponding to the gimbal camera to the subspace corresponding to the telephoto camera. The gimbal camera continuously tracks the target until the telephoto camera confirms capture, ensuring that the target's image is not lost during the return trip.

[0214] In some embodiments, controlling the gimbal camera to continuously track the target can be implemented as follows: When at least one of the following triggering conditions is met, the gimbal camera is controlled to continuously track the target: The target is a preset priority tracking target; The attention weights of other targets in the third subspace are lower than a preset threshold.

[0215] As one possible implementation, priority tracking targets can be pre-specified by the user. For example, before or during a meeting, a user can mark a participant as a priority tracking target, and this mark will not expire with changes in attention weight. Once marked as a priority tracking target, the gimbal camera will continue to track the target even if its attention weight decreases temporarily due to inactivity, without releasing tracking due to fluctuations in attention weight. This mark can be manually removed by the user or automatically cleared after the meeting ends.

[0216] For example, in a keynote speech scenario, the speaker is pre-marked as the priority tracking target. When the speaker pauses briefly or someone interrupts to ask a question, the attention weight of the interrupter may temporarily increase, but because the speaker has already been marked as the priority tracking target, the gimbal camera continues to track the speaker and does not switch tracking targets due to the brief activity of others; it should be added that the telephoto camera responsible for the subspace where the interrupter is located can independently track the interrupter, and two close-up images are output in parallel, so that remote participants can see close-up images of both the speaker and the interrupter at the same time.

[0217] As another possible implementation, when the attention weights of other targets in the third subspace are all below a preset threshold, the gimbal camera is controlled to continuously track the target.

[0218] For example, if a target tracked by a gimbal camera moves from the far-end area into the telephoto area, and this target is the only or absolutely most important active speaker in the current meeting scenario (its attention weight is significantly higher than other targets across the entire room, and all other targets are inactive), and the attention pool in the far-end area currently has no other targets that need to be covered by the gimbal camera (e.g., the attention pool in the far-end area is empty or only contains non-important targets), and the meeting mode allows it (e.g., free discussion mode or single-person presentation mode), then in this case, the gimbal camera can continuously track the target without being restricted by the zoning, even if the target moves out of the far-end area (far-distance area). The gimbal camera can continue to track the target to the telephoto area (mid-distance area) or even closer areas. Thus, the two telephoto cameras are responsible for other targets within their respective subspaces, while the gimbal camera keeps track of this single important target throughout.

[0219] For example, in a single-person presentation or remote teaching scenario, if there is only one key person in the venue, the gimbal camera enters full-time tracking mode, without being restricted by zones, and continuously locks onto the target.

[0220] In this embodiment, by using preset priority tracking targets and low attention weights within the subspace as triggering conditions, the gimbal camera performs follow-up tracking only when necessary and returns to the dominant position in the far-end area in a timely manner when not necessary. This ensures both flexible maneuverability in important scenarios and stable coverage of the dominant position in the far-end area under normal conditions, balancing system resource allocation and image output quality.

[0221] In some embodiments, the method may further include: during the generation of a close-up image of the target by the gimbal camera, if a change in the target's attitude is detected and it is predicted that the target will exceed the current field of view of the gimbal camera, then: the gimbal camera is controlled to generate a transition image during the adjustment of its pitch angle, the transition image being a dynamically digitally cropped image of the currently captured image by the gimbal camera, the cropped area moving with the real-time position of the target. After the gimbal camera adjusts to the target's pitch angle, the generation of an optical close-up image of the target is resumed.

[0222] For example, if a target being tracked by the gimbal camera in the far zone (Zone_F) undergoes a significant change in posture, such as changing from a sitting to a standing posture, or a change in height of 40-60cm, causing the target's head to almost exceed the upper boundary of the current gimbal camera's image, there is a risk of losing the target's head. In this situation, the gimbal camera's adaptive pitch adjustment is triggered. Furthermore, the gimbal camera can generate transitional frames during the pitch angle adjustment process.

[0223] In this embodiment, when the target tracked by the gimbal camera undergoes a change in posture, a dynamic digital cropping transition image is generated during the adjustment of the gimbal camera's pitch angle. This achieves seamless composition adjustment when the target's posture changes (such as switching from a sitting to a standing posture). The cropped area moves with the target's real-time position, ensuring that the output image always maintains a complete and stable view of the target's head. This reduces the risk of the target's head exceeding the image boundary (target head loss) due to gimbal mechanical rotation lag. This allows remote attendees to always see a complete close-up of the speaker's head and upper body, improving the viewing experience in dynamic scenarios such as the speaker standing up or moving.

[0224] In some embodiments, the processing device can also adaptively select a matching smooth transition strategy based on the meeting status. For example, refer to Table 6: Table 6

[0225] In some embodiments, the method further includes: determining the Virtual Coaxial Reference Plane (VCRP) parameters of the gimbal camera based on its current pose. When the first telephoto camera, the second telephoto camera, or the gimbal camera generates a close-up image, a position compensation value is determined based on the differences between the VCRP parameters of each camera, and the generated close-up image is adjusted based on the position compensation value.

[0226] As one possible implementation, VCRP can be defined as the optical center of the wide-angle stitching module, parallel to the plane of the device's front panel. For each camera, the following VCRP parameters are measured and stored: The physical offset (Δx, Δy, Δz) of the lens optical center relative to the VCRP; The intrinsic parameters of the lens, such as focal length, optical center, and distortion coefficient; The deflection angle of the lens optical axis relative to the VCRP normal (θ_pan, θ_tilt).

[0227] The camera parameters can be compiled into a Spatial Mapping Lookup Table (SMLUT). When a shot switch is to be performed, the processing unit reads the VCRP parameters of the source and target shots from the SMLUT, uses the difference between them as a position compensation value, and combines this with the scene depth estimate 'd' provided by the wide-angle panorama to perform a homography warp on the target shot image. This allows for position compensation of the close-up image generated by the target shot, ensuring precise alignment of the target position in the images from the two cameras at the moment of the switch.

[0228] In this embodiment, a virtual coaxial reference plane (VCRP) based on the optical center of the wide-angle stitching module is established to quantify the physical position differences between cameras into calibrable and compensable system parameters. During cross-camera switching, position compensation is performed on the close-up image based on the VCRP parameter differences between the source and target cameras, ensuring that the spatial position of the target remains aligned in the images of different cameras. This reduces abrupt changes in target position (spatial parallax jump) during switching caused by different camera installation positions, and improves the spatial continuity of the image and the viewing experience at a distance during cross-camera switching.

[0229] In some embodiments, the method further includes: when it is detected that the targets in the panoramic image are concentrated in the far-end area, controlling the gimbal camera to generate a group image with a wide field of view. Alternatively, when it is detected that there are two targets speaking alternately in the third subspace, controlling the gimbal camera to capture an image containing the two targets with a wide field of view and digitally cropping it to generate two independent close-up images of each individual.

[0230] Wide field of view refers to the angle from which the group can be captured.

[0231] When there are multiple key positions, the gimbal camera can simultaneously capture multiple key figures with a wide field of view and generate multiple independent single-person close-ups through digital cropping, reducing the frequent camera switching caused by alternating speeches.

[0232] For example, the processing device detects that an active speaker A is located in the far-end area corresponding to the gimbal camera, while another active speaker B is present within the telephoto coverage area. B's real-time spatial coordinates are located in the transition zone (TZ_LF or TZ_RF) or near the boundary of the transition zone within the telephoto coverage area (<1m from the boundary). Speakers A and B alternate speaking for approximately N minutes (e.g., N=1-3), forming a continuous discussion mode. In this case, a cross-zone unified output mode is triggered. The gimbal camera appropriately widens its field of view, simultaneously including both speakers in the same optical frame. Since speaker B is in the transition zone, which is itself within the effective coverage area of ​​the gimbal camera, the gimbal camera only needs to fine-tune its field of view to simultaneously cover both speakers A and B. The processing device can perform real-time intelligent cropping of the gimbal camera's image, generating an independent single-person close-up sub-image for each speaker. The main output image selects to display the close-up of the speaker based on the current speaking status.

[0233] By using a gimbal camera to capture two speakers taking turns speaking with a wide field of view, and digitally cropping the footage to generate two independent close-up shots of each speaker, the system ensures that both speakers receive independent and stable close-up outputs. Furthermore, the switching between the two cropped shots is done electronically, reducing the frequency of gimbal camera rotation and improving the stability and smoothness of the footage. In addition, since both close-ups originate from the same gimbal camera's optical image, the shooting angles and visual styles are completely consistent, enhancing the uniformity of image quality.

[0234] In some embodiments, when attendees are seated in the back row of the far end area, the processing device controls the gimbal camera to capture all attendees in the back row with a suitable field of view and outputs a group image.

[0235] For example, if the attention weights of the current target individuals are all below the threshold, or if the system outputs only a group shot through user triggering or mode selection, the processing device can identify that the participants are concentrated in the back row of the far-end area through the panoramic view. It then controls the gimbal camera to capture all the participants in the back row with a suitable field of view and outputs a group shot. When the participants are concentrated in the far-end area, the gimbal camera without close-up tracking is used to output a group shot. The optical zoom capability of the gimbal camera compensates for the insufficient resolution of wide-angle cameras at long distances, allowing far-end participants to clearly identify the details of the people in the back row, thus improving the usability of the panoramic view and the flexibility of multi-camera resource configuration.

[0236] Different lenses use independent image sensors and image signal processors (ISPs), and their respective AE / AWB parameters operate independently, which may result in the same scene appearing with different brightness and color in different lens shots. For example, flickering may occur in the image during transitions.

[0237] Therefore, in some embodiments, the panoramic view from the wide-angle camera can be defined as an optical reference frame. The AE target brightness of the telephoto camera and the gimbal camera is referenced to the local brightness of the corresponding area in the panoramic view. Furthermore, the AWB of the telephoto camera and the gimbal camera is based on the white balance of the panoramic view, with a compensation offset superimposed that is only related to the lens optical characteristics.

[0238] As one possible implementation, the processing device determines the N frames before switching the output source (e.g., N=3-5), and the target lens begins to gradually adjust its ISP parameters to align with the optical reference frame. During the switching transition frames (e.g., frames 2-3), the processing device can perform weighted alpha blending on the source and target images to reduce residual inconsistencies.

[0239] In some embodiments, an Audio-Visual Fusion Coordinate Repository (AVFCR) can be established. The microphone array estimates the DOA (Direction of Arrival) and distance of each sound source in real time and stores the sound source coordinates in the AVFCR. A wide-angle camera searches for face / human targets near the acoustic locations marked by the AVFCR in the panoramic view, replacing the estimated acoustic coordinates with the coordinates of the center of the visual detection box. Thus, through both visual and audio detection, the accuracy of target recognition can be improved.

[0240] As one possible approach, before visual confirmation is complete, the gimbal camera can initiate low-speed pre-aiming based on acoustic coarse positioning to shorten response latency.

[0241] As one possible implementation, the processing device can adaptively adjust the close-up composition based on acoustic signal characteristics such as distance, duration, and volume. For example, the composition remains unchanged for close-up half-body shots, long-distance full-body shots, and shots of continuous speakers.

[0242] The above primarily describes the solutions provided by the embodiments of this application from a methodological perspective. It is understood that, in order to achieve the above functions, the electronic device includes hardware structures and / or software modules corresponding to the execution of each function. Based on the units and algorithm steps of the various examples described in the embodiments disclosed in this application, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by a computer driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solutions of the embodiments of this application.

[0243] This application embodiment can divide the electronic device into functional modules according to the above method example. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional module. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0244] This application embodiment also provides a video conferencing terminal, including a housing and a camera assembly disposed on the housing: A dual wide-angle stitching module is used to acquire images and stitch them together to generate the panoramic image. A dual telephoto zone close-up module, comprising a first telephoto camera and a second telephoto camera; A gimbal dynamic maneuvering module, including the gimbal camera; A processor for implementing the method described in any one of the embodiments of this application.

[0245] For example, Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown.

[0246] like Figure 3 As shown, electronic device 500 may include processor 510.

[0247] Optionally, the electronic device 500 may also include a memory 520.

[0248] For example, when the electronic device 500 is a standalone host, it may not include a camera, which is set up separately. When the electronic device 500 is an integrated device, it may integrate multiple cameras, which may include at least a first camera 530, a second camera 540, and an Nth camera 550.

[0249] Dual wide-angle stitching module: This module can include two wide-angle lenses with overlapping fields of view to ensure high-quality panoramic stitching. The images from the two wide-angle lenses can be stitched together to generate a panoramic image. The global spatial coordinate system is established based on this panoramic image. The dual wide-angle module can also handle close-up shots of objects at close range (such as portraits) when needed.

[0250] Dual telephoto zone close-up module: This may include multiple telephoto lenses; this embodiment uses two as an example. The two telephoto lenses are located on the same horizontal plane to ensure spatial consistency in tracking the left and right zones. For example, they are responsible for stable close-up tracking of the left and right zones of a conference room, respectively.

[0251] The gimbal dynamic maneuvering module includes a gimbal zoom lens. For example, it can be located on the lower center line of the equipment housing, initially pointing towards the far end of the conference room's central axis, such as key seats in the back row, serving as a dynamically manageable mobile visual resource. Alternatively, it can be mounted on the middle section of the housing to reduce the computational complexity of coordinate conversion.

[0252] For example, design dimensions such as camera position and number of rows can be referenced in Table 7.

[0253] Table 7

[0254] For example, all five cameras are located on the same horizontal line, arranged in a single row. This allows for a thinner device, requiring only the height to accommodate a single lens group, making it suitable for slim and lightweight video conferencing terminal designs. Furthermore, with the optical centers of all lenses located in the same horizontal plane, VCRP parallax calibration only needs to handle offsets in the horizontal (X-axis) and depth (Z-axis) dimensions, with zero parallax in the vertical (Y-axis) direction, resulting in lower calibration complexity and higher calibration accuracy. In addition, all five lenses in the single-row arrangement are fully visible on the front panel of the device, eliminating lens obstruction issues caused by stacking. For example, an example of a single-row camera arrangement can be found here. Figure 4 (a)

[0255] As another example, the five cameras are arranged in two rows. The wide-angle module is located at the top, and the telephoto module and gimbal camera are located at the bottom, as shown in the reference. Figure 4 The double-row arrangement is shown in (b). Alternatively, the wide-angle module is located in the lower row, and the telephoto module and gimbal camera are located in the upper row.

[0256] As another example, the five cameras are arranged in three rows. The wide-angle module, the telephoto module, and the gimbal camera are each in one row, as shown in the example. Figure 4 The three-row arrangement shown in (c) is correct. Of course, Figure 4 The arrangement of the five cameras shown in (a), (b) and (c) is merely an example and is not intended to limit the scope of this application.

[0257] For example, the telephoto cameras are located in the middle area of ​​the housing and are arranged opposite each other. That is, the telephoto camera responsible for the left area is placed on the right side of the housing, and the telephoto camera responsible for the right area is placed on the left side of the housing, with overlapping fields of view.

[0258] For example, the telephoto camera responsible for the left area is placed at the farthest left end of the housing, and the telephoto camera responsible for the right area is placed at the farthest right end of the housing. This arrangement is independent of the row number selection. This allows the two telephoto cameras to extend as far outwards as possible from both ends of the housing, rather than being clustered together in the center next to the wide-angle camera. On one hand, the two telephoto cameras are located at the far left and far right ends of the housing, respectively, with a large horizontal distance (baseline) between them. The longer baseline makes the telephoto camera responsible for the left area further to the left and the telephoto camera responsible for the right area further to the right. With their respective field of view remaining unchanged, the total coverage sector of the left and right zones naturally expands outwards, reducing the coverage blind spots at the mid-range left and right ends. Simultaneously, the long baseline enhances the parallax signal strength of the dual wide-angle modules in the overlapping area, which is beneficial for improving the accuracy of VCRP depth estimation. On the other hand, placing the telephoto cameras at both ends of the housing means that the lateral spacing of the telephoto cameras is not limited by the width of the wide-angle module. Even though the dual wide-angle modules require a certain amount of lateral space to ensure the stitching overlap rate (e.g., a spacing of approximately 60mm between the two wide-angle lenses), the telephoto camera can still be independently positioned further out (e.g., a spacing of 80-120mm). The two do not conflict in terms of lateral space, effectively improving the utilization efficiency of the device's width. Furthermore, in terms of image quality, a higher proportion of the subject's face is directly facing the lens, resulting in a more complete frontal view rather than a side profile.

[0259] The above examples illustrate the number and location of each camera, but these examples do not constitute a limitation on the number and location of cameras.

[0260] For large conference room scenarios (such as 15-25 people, U-shaped layout), wide-angle cameras with a larger field of view (such as 120°), telephoto cameras with a longer focal length (such as 80mm), and gimbal cameras with a larger zoom range (such as 50mm-600mm) can be used to accommodate wider conference room layouts and greater distances between key seats (such as up to 15-20m).

[0261] After powering on, the dual wide-angle modules can simultaneously acquire images and stitch them together in real time to generate an ultra-wide panoramic image. The system performs spatial calibration: using a calibration board, images from each lens are acquired at multiple different locations, and the offset parameters of each lens relative to the VCRP are calculated based on these images. Furthermore, the boundaries of the left, right, far-end, and transition zones are delineated in the global spatial coordinate system: this can be calibrated by installers through the management interface based on the actual layout of the conference room, or automatically estimated by the system based on the table and chair layout in the panoramic view from the wide-angle camera.

[0262] For example, the output of the five-camera system defaults to a picture-in-picture layout. For instance, the main screen is a close-up of the target with the highest current weight, while the sub-screen is a panoramic view from the wide-angle stitching module. For example, in the panoramic view, detection boxes for all speakers are labeled, and the box colors gradient from warm to cool based on activity level.

[0263] Exemplary Reference Figure 3 The processor 510 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). These different processing units may be independent devices or integrated into one or more processors.

[0264] The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0265] The processor 510 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 510 is a cache memory. This memory can store instructions or data that the processor 510 has just used or that are used repeatedly. If the processor 510 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 510, and thus improves the efficiency of the system.

[0266] In some embodiments, processor 510 may include one or more interfaces. These one or more interfaces may be used to connect processor 510 to memory 520, etc.

[0267] In some embodiments of this application, the processor 510 can be used to control the first camera 530, the second camera 540, and the Nth camera 550 to perform any of the methods described herein. For detailed descriptions, please refer to other relevant content herein.

[0268] For example, the processor 510 may employ a high-performance multi-core chip to support parallel processing of data from multiple cameras.

[0269] The memory 520 can be used to store computer executable program code, which includes instructions. The memory 520 may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc. The data storage area may store data created during the use of the electronic device 500, etc. The processor 510 executes various functional applications and data processing of the electronic device 500 by running instructions stored in the memory 520 and / or instructions stored in memory disposed within the processor.

[0270] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include... Figure 3 The diagram shows more or fewer components, or combinations of components, or separate components, or different arrangements of components. The components shown can be implemented in hardware, software, or a combination of both.

[0271] like Figure 5 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application. This electronic device 2200 can be used to implement the methods described in the above method embodiments. For example, the electronic device 2200 may specifically include a processing unit 2201. The processing unit 2201 is used to support the electronic device 2200 in performing operations. Figure 1 The aforementioned processing function.

[0272] Optional, Figure 5 The electronic device 2200 shown may also include a communication unit ( Figure 5 (Not shown in the image), this communication unit is used to support electronic device 2200 in performing the steps of communication between electronic device and other electronic devices in the embodiments of this application.

[0273] Optional, Figure 5 The illustrated electronic device 2200 may further include a storage unit 2203 that stores programs or instructions. When the processing unit 2201 executes the program or instructions, it causes... Figure 5 The electronic device 2200 shown can perform the method described in the above-described method embodiments.

[0274] Figure 5 The technical effects of the electronic device 2200 shown can be referred to the technical effects of the method shown in the above method embodiments, and will not be repeated here. Figure 5 The processing unit 2201 involved in the illustrated electronic device 2200 can be implemented by a processor or processor-related circuit components, and can be a processor or a processing module. The communication unit can be implemented by a transceiver or transceiver-related circuit components, and can be a transceiver or a transceiver module.

[0275] This application also provides a chip system, such as... Figure 6As shown, the chip system includes at least one processor 2301 and at least one interface circuit 2302. The processor 2301 and the interface circuit 2302 are interconnected via lines. For example, the interface circuit 2302 can be used to receive signals from other devices. As another example, the interface circuit 2302 can be used to send signals to other devices (e.g., the processor 2301). Exemplarily, the interface circuit 2302 can read instructions stored in memory and send those instructions to the processor 2301. When the instructions are executed by the processor 2301, the electronic device can perform the various steps performed by the electronic device in the above embodiments. Of course, the chip system may also include other discrete devices, and this application embodiment does not specifically limit this.

[0276] Optionally, the chip system may contain one or more processors. These processors can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented in software, the processor can be a general-purpose processor, implemented by reading software code stored in memory.

[0277] Optionally, the chip system may contain one or more memories. The memory may be integrated with the processor or disposed separately from it; this application does not limit this. For example, the memory may be a non-transient processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or disposed separately on different chips. This application does not specifically limit the type of memory or the arrangement of the memory and processor.

[0278] For example, the chip system can be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a micro controller unit (MCU), a programmable logic device (PLD), or other integrated chips.

[0279] It should be understood that each step in the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The method steps disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor.

[0280] It should be noted that the electronic device provided in this application embodiment belongs to the same concept as the method in the above embodiments. Any of the methods provided in the method embodiments can be run on the electronic device, and the specific implementation process is detailed in the method embodiments, which will not be repeated here. For example, the processor in the electronic device can execute the steps in the method. The embodiments, implementation methods, and related technical features of this application can be combined and substituted for each other without conflict.

[0281] This application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in any of the above embodiments.

[0282] In the embodiments of this application, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0283] It should be noted that, for the methods of the embodiments of this application, those skilled in the art will understand that all or part of the processes of the methods of the embodiments of this application can be implemented by a computer program controlling related hardware. This computer program can be stored in a computer-readable storage medium, such as in the memory of an electronic device, and executed by at least one processor within the electronic device. During execution, it can include the processes of the embodiments of the method. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, etc.

[0284] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0285] The above are merely preferred embodiments of this application and are not intended to limit this application in any way. Although this application has disclosed preferred embodiments as above, it is not intended to limit this application. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the technical solution of this application. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of this application without departing from the scope of the technical solution of this application shall still fall within the scope of the technical solution of this application.

Claims

1. A method for outputting conference video, characterized in that, include: Acquire a panoramic image of the first space captured by a wide-angle camera, and establish a partition model based on the panoramic image. The partitioning model divides the first space into multiple subspaces, and the multiple subspaces include at least a first subspace, a second subspace, and a third subspace. Within each of the subspaces, determine the primary objective of each subspace; The camera corresponding to the subspace is invoked to generate a close-up image corresponding to each main target; wherein the first subspace and the second subspace correspond to the first telephoto camera and the second telephoto camera, respectively, and the third subspace corresponds to the gimbal camera; According to the preset output strategy, output a close-up image of the target.

2. The method according to claim 1, characterized in that, The step of determining the main target of each of the subspaces includes: Within each of the subspaces, attention evaluation is performed on the candidate targets within each subspace, and the main target of each subspace is determined based on the evaluation results.

3. The method according to claim 2, characterized in that, The process of evaluating the attention of candidate targets within each of the subspaces and determining the primary target of each subspace based on the evaluation results includes: Obtain multidimensional feature information for each candidate target within the subspace. The multidimensional feature information includes one or more of the following: speaking state, motion state, spatial location, posture, dialogue history, meeting role label, or electronic device usage status. The attention weight for each candidate target is determined based on the multidimensional feature information; The candidate target with the highest attention weight in the subspace is determined as the main target.

4. The method according to claim 1, characterized in that, The preset output strategy includes at least one of the following strategies: First strategy: In response to the target moving across subspace, switch the camera used to generate close-up images of the target and continuously output close-up images of the target; Second strategy: Output the multiple close-up shots corresponding to each main target as the target close-up shots in parallel; Third strategy: Combine the multi-channel close-up shots corresponding to each main target into a single-channel combined shot and output it as the target close-up shot; Fourth strategy: Perform cross-subspace attention comparison on each main target, and select one close-up shot as the target close-up shot for output; Fifth strategy: Pause the output of close-up images of human subjects and switch to outputting close-up images of non-human subject areas.

5. The method according to claim 4, characterized in that, The first strategy includes: When the first target is detected to move from the first subspace to the third subspace, the first telephoto camera is controlled to continuously generate close-up images of the first target until the first target moves out of the field of view of the first telephoto camera, and / or enters the field of view of the gimbal camera, the gimbal camera is controlled to take over the generation of close-up images of the first target.

6. The method according to claim 4, characterized in that, The first strategy also includes: During the process of switching between the first telephoto camera and the gimbal camera to generate a close-up image of the first target, the first telephoto camera is controlled to output a transitional image of the first target to the edge of the field of view of the first telephoto camera by moving the cropping window in the captured image.

7. The method according to claim 4, characterized in that, The second strategy includes: Obtain the attention weights of each main target in its respective subspace; The attention weights of each main target are sorted, and close-up shots that meet the preset conditions are selected according to the sorting results and output in parallel.

8. The method according to claim 7, characterized in that, The preset conditions include any of the following: The top N positions, where N is a preset positive integer; The attention weight is greater than the preset threshold.

9. The method according to claim 4, characterized in that, The third strategy includes: Obtain the attention weights of each main target in its respective subspace; The attention weights of each main target are sorted, and the arrangement position and / or display size of each close-up shot in the combined shot are determined according to the sorting result.

10. The method according to claim 4, characterized in that, The fourth strategy includes: Obtain the attention weights of each main target in its respective subspace; The attention weights of each main target are compared across subspaces, and the close-up image corresponding to the main target with the highest attention weight is selected as the target close-up image for output.

11. The method according to claim 4, characterized in that, The fourth strategy also includes: When the attention weight of the selected main target is lower than a preset threshold, the current output screen remains unchanged.

12. The method according to claim 4, characterized in that, The fifth strategy includes: The panoramic image is used for target recognition. When a non-human target area is detected and a preset trigger condition is met, the gimbal camera is controlled to rotate toward the non-human target area to generate a close-up image of the non-human target area. The non-human target area includes at least one of a whiteboard area, a projection screen area, or a physical display stand area.

13. The method according to claim 12, characterized in that, The preset triggering conditions include: Content updates were detected in the non-human target area.

14. The method according to claim 4, characterized in that, The preset output strategy also includes: At least two basic strategies from the first to the fifth strategy are combined to form a combined strategy to output the close-up image of the target.

15. The method according to claim 14, characterized in that, The preset output strategy also includes: Obtain the current meeting mode, and determine the target strategy from the basic strategy or the combined strategy based on the meeting mode.

16. The method according to claim 1, characterized in that, Also includes: For each target detected in the panoramic image, maintain a target state record. The target state record includes one or more of the following information: target identifier, current subspace, real-time coordinates in the first space, motion vector, timestamp of entering the current subspace, or confidence level of the subspace. The subspace to which the current subspace belongs is determined by matching the real-time coordinates with the spatial range of each subspace.

17. The method according to claim 3, characterized in that, Within any of the subspaces, when the primary target changes from an old target to a new target, generating a close-up shot corresponding to each primary target includes: The corresponding camera is controlled to move the cropping window within the captured image by a preset movement step size, generating a close-up image transitioning from the old target to the new target; the starting position of the cropping window is the location of the old target, and the ending position is the location of the new target; The attention weight of the new target is greater than the attention weight of the old target, and the difference between the two exceeds a preset weight threshold.

18. The method according to claim 1, characterized in that, Among the plurality of subspaces, adjacent subspaces have transition areas for cross-regional pick-up and drop-off; the method further includes: When a target is detected moving from its current source subspace to an adjacent target subspace via the transition region, cross-regional pick-up and drop-off scheduling is performed: Predict the target position of the target entering the target subspace based on the panoramic image; If the attention weights of all other targets in the target subspace are less than or equal to the attention weight of the target, then the camera corresponding to the target subspace is controlled to pre-rotate toward the target position or take over to generate a close-up image of the target after the target enters its field of view; If there is a high-attention target being tracked in the target subspace, the camera corresponding to the target subspace is controlled to adjust its field of view to simultaneously cover the target and the high-attention target, generating a frame-to-frame or single-person close-up image of the target and the high-attention target.

19. The method according to claim 18, characterized in that, When the camera corresponding to the source subspace is a fixed telephoto camera and the camera corresponding to the target subspace is a gimbal camera: The step of controlling the camera corresponding to the target subspace to pre-rotate toward the target position includes: controlling the gimbal camera to rotate toward the target position; The step of taking over the generation of a close-up image of the target after it enters its field of view includes: During the rotation of the gimbal camera, the fixed telephoto camera is controlled to continuously output close-up images of the target by moving the cropping window in the captured image until the target moves out of the field of view of the fixed telephoto camera, or until the gimbal camera completes its rotation. After the target enters the field of view of the gimbal camera, the system switches to an optical close-up image of the target generated by the gimbal camera.

20. The method according to claim 18, characterized in that, When the camera corresponding to the source subspace is a first telephoto camera and the camera corresponding to the target subspace is a second telephoto camera, the execution of cross-regional pick-up and drop-off scheduling includes: The first telephoto camera is controlled to continuously output close-up images of the target by moving the cropping window in the captured image until the target moves out of the field of view of the first telephoto camera. After the target enters the field of view of the second telephoto camera, the camera switches to a close-up view of the target generated by the second telephoto camera. Within a preset duration after switching, the second telephoto camera is controlled to continuously output close-up images of the target located in the center of the image by moving the cropping window in the captured image.

21. The method according to claim 20, characterized in that, The methods for moving the cropping window include: Based on the real-time position of the target in the panoramic image, determine the cropping position of the target in the image captured by the corresponding camera; The cropping window is moved from the starting position to the ending position with a preset movement step size so that the position of the target in the output screen remains continuous.

22. The method according to claim 18, characterized in that, The execution of cross-regional pick-up and drop-off scheduling also includes: When a target is detected moving from the third subspace to the first subspace via the transition region, the gimbal camera is controlled to continuously track the target; and / or After confirming that the target has entered the first subspace and that the first telephoto camera has captured the target based on the panoramic view, the system switches to a close-up view of the target generated by the first telephoto camera.

23. The method according to claim 22, characterized in that, The control of the gimbal camera to continuously track the target includes: When at least one of the following triggering conditions is met, the gimbal camera is controlled to continuously track the target: The target is a preset priority tracking target; The attention weights of other targets in the third subspace are lower than a preset threshold.

24. The method according to claim 1, characterized in that, Also includes: During the generation of a close-up image of the target by the gimbal camera, if a change in the target's posture is detected and it is predicted that the target will move beyond the current field of view of the gimbal camera, then: The gimbal camera is controlled to generate a transitional image during the adjustment of the pitch angle. The transitional image is a dynamically digitally cropped image obtained from the image currently captured by the gimbal camera. The cropped area moves with the real-time position of the target. After the gimbal camera is adjusted to the target pitch angle, an optical close-up image of the target is generated.

25. The method according to claim 1, characterized in that, Also includes: Based on the current pose of the gimbal camera, determine the virtual coaxial reference plane (VCRP) parameters of the gimbal camera; When the first telephoto camera, the second telephoto camera, or the gimbal camera generates a close-up image, a position compensation value is determined based on the differences between the VCRP parameters of each camera, and the generated close-up image is adjusted based on the position compensation value.

26. The method according to claim 1, characterized in that, Also includes: When it is detected that the target in the panoramic image is concentrated in the far-end area, the gimbal camera is controlled to generate a collective image with a wide field of view. or, When the presence of two targets speaking alternately within the third subspace is detected, the gimbal camera is controlled to capture images containing the two targets with a wide field of view and digitally cropped to generate two independent single-person close-up images.

27. A video conferencing terminal, characterized in that, Includes a housing and a camera assembly mounted on the housing: A dual wide-angle stitching module is used to acquire images and stitch them together to generate the panoramic image. A dual telephoto zone close-up module, comprising a first telephoto camera and a second telephoto camera; A gimbal dynamic maneuvering module, including the gimbal camera; A processor for implementing the method of any one of claims 1 to 26.

28. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 26.