Video collection method and apparatus, and smart device and storage medium
By combining the built-in camera and microphone array of smart devices with the installation positions and parameter information of multiple pan-tilt cameras, the problem of fixed-angle cameras being unable to adapt to multi-person communication scenarios is solved, enabling rapid detection and good presentation of key communication figures and improving the communication experience of video data.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUANGZHOU SHIYUAN ELECTRONICS CO LTD
- Filing Date
- 2025-10-27
- Publication Date
- 2026-05-07
AI Technical Summary
The cameras on existing smart devices can only capture video from a fixed angle, which cannot adapt to the changing multi-person communication scenarios. The resulting technical problem is that they cannot guarantee high-quality video capture of the communication scene and cannot adapt to changes in the scene, thus affecting the communication experience.
By using the built-in camera and microphone array of smart devices for positioning and attitude detection, and combining the installation positions and parameter information of multiple pan-tilt cameras, the target camera is identified and controlled to shoot under the target control parameters, adapting to changes in the key person being communicated with and ensuring good presentation.
It enables rapid detection of changes in key communication figures in multi-person communication scenarios, improves the communication experience of video data, and ensures that key communication figures are presented well to remote users.
Smart Images

Figure CN2025130308_07052026_PF_FP_ABST
Abstract
Description
Video capture methods, devices, smart devices and storage media
[0001] This application claims priority to Chinese Patent Application No. CN2024115394071, filed on October 31, 2024, entitled "Video Acquisition Method, Apparatus, Smart Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of video technology, and more particularly to video acquisition methods, apparatus, smart devices, and storage media. Background Technology
[0003] With the continuous development of electronic technology, intelligent devices that comprehensively utilize various electronic technologies are being used more and more extensively in data-driven information transmission scenarios. For example, in scenarios involving information transmission among multiple people, such as offices and classrooms, intelligent devices (such as conference tablets and smart blackboards) are configured to present the key information during communication. Furthermore, cameras installed in these intelligent devices can capture images of the surrounding environment, allowing remote participants to see the live feed, thus enhancing their immersion and communication experience.
[0004] When the inventors carefully examined the live footage captured by existing smart devices, they discovered that the cameras fixed to the smart devices, once configured for a specific scenario, could only capture live footage from a fixed angle. The only adjustable parameter during the capture process was the focal length. In the face of multi-person communication scenarios with rich and varied live conditions, capturing footage from a fixed angle could not guarantee a high-quality reproduction of the communication scene, thus affecting the communication experience. Summary of the Invention
[0005] This application provides a video acquisition method, apparatus, smart device, and storage medium to solve the technical problem that existing video acquisition methods for communication sites cannot adapt to the rich changes in the site conditions, and the acquired video data affects the communication experience.
[0006] In a first aspect, embodiments of this application provide a video acquisition method, which is applied to a smart device. The smart device includes a first camera and a microphone array, and is also connected to multiple pan-tilt cameras located in the same target space, and stores the installation location information and device parameter information of the multiple pan-tilt cameras; the video acquisition method includes:
[0007] Acquire the first video stream captured by the first camera and the location of the sound source confirmed by the microphone array;
[0008] Based on the sound source location and the first video feed, confirm the facial orientation and spatial location of the human figure in the first video feed;
[0009] Based on the installation location information and equipment parameter information, the target camera that matches the face orientation and spatial position, as well as the corresponding target control parameters, are identified from multiple PTZ cameras.
[0010] The target camera is controlled to capture video data of the target space under the target control parameters, and the target video data is used as the video acquisition result of multiple pan-tilt cameras.
[0011] The above describes a system that uses the built-in camera and microphone array of a smart device to locate and detect the pose of a human figure in the target space. Multiple external pan-tilt cameras have already been calibrated within the target space, and the shooting range of each camera can be confirmed based on device parameter information. The system then identifies the target camera that will produce good video capture of the human figure, along with the necessary target control parameters. When capturing video using multiple pan-tilt cameras, the system directly uses the capture results from the target camera under these parameters. In multi-person communication scenarios, this system can quickly detect changes in the key person in the communication and control the pan-tilt camera that provides good shooting results for that person. This adapts to the varied changes in the key person during the communication, ensuring a good presentation of them to remote users and enhancing the video-based communication experience.
[0012] Based on the location of the sound source and the first video feed, the facial orientation and spatial location of the human figure in the first video feed were determined, including:
[0013] Identify the area to be identified in the video frame corresponding to the first video stream, pointing to the direction of the sound source.
[0014] Perform lip state detection on the face image of the region to be identified to confirm that the lips of the portrait target are in an active state.
[0015] Based on the human target and the preset parameter mapping relationship, the face orientation and spatial position of the human target are confirmed. The parameter mapping relationship is used to characterize the length ratio relationship in the imaging optical path of the first camera.
[0016] Based on the directional nature of the sound source, the human target is quickly identified in the first video stream. Then, based on the human target and the parameter relationships during the imaging process, the face orientation and the spatial position estimated based on the two-dimensional image are determined, providing a reference for accurately calling and controlling the PTZ camera.
[0017] Specifically, based on the mapping relationship between the human face target and preset parameters, the facial orientation and spatial location of the human face target are determined, including:
[0018] If the deviation between the human face target and the historical human face target when the face orientation and spatial location were most recently confirmed reaches a preset trigger threshold, the face orientation and spatial location of the human face target are confirmed according to the human face target and the preset parameter mapping relationship.
[0019] The process described above, which involves confirming whether to reconfirm the target camera and target control parameters based on changes in the human image target, can effectively reduce the repeated confirmation of the same target camera and the repeated confirmation of the same or similar target control parameters, thereby reducing unnecessary data processing. At the same time, it ensures that when the human image target changes significantly, the PTZ camera can be quickly controlled to track and shoot, ensuring the presentation of the key person in the target video data.
[0020] Among them, when the deviation between the portrait target and the historical portrait target at the time of the most recent confirmation of face orientation and spatial location reaches a preset trigger threshold, after confirming the face orientation and spatial location of the portrait target according to the portrait target and the preset parameter mapping relationship, the process also includes:
[0021] Cache portrait targets as historical portrait targets;
[0022] Correspondingly, when the deviation between the portrait target and the historical portrait target at the time of the most recent confirmation of face orientation and spatial location reaches a preset trigger threshold, the face orientation and spatial location of the portrait target are confirmed according to the portrait target and the preset parameter mapping relationship, including:
[0023] The deviation between the target human face and the cached historical target human face is confirmed. If the deviation reaches a preset trigger threshold, the face orientation and spatial position of the target human face are confirmed according to the preset parameter mapping relationship between the target human face and the target human face.
[0024] As mentioned above, whenever the human target changes, causing changes in the target camera or target control parameters, the human target is cached as a historical human target. This serves as a reference for whether to reconfirm the target camera or target control parameters during subsequent human target detection, thereby improving data processing speed.
[0025] Specifically, based on the mapping relationship between the human face target and preset parameters, the facial orientation and spatial location of the human face target are determined, including:
[0026] Facial orientation recognition is performed on human targets to confirm the orientation of their faces;
[0027] Confirm the size and position parameters of the human figure, and determine the depth parameters of the human figure based on the size parameters and parameter mapping relationship;
[0028] The spatial location of the human figure target is determined based on the position and depth parameters.
[0029] The above-mentioned method uses the ratio between the image size and the length of the imaging optical path to determine the spatial location of the human target in the real world based on the two-dimensional image. The method uses the hardware that is standard in smart devices to quickly identify the human target, effectively controlling the overall hardware cost of the solution.
[0030] Specifically, based on installation location information and device parameter information, the target camera matching the face orientation and spatial position is identified from multiple PTZ cameras, along with the corresponding target control parameters, including:
[0031] Based on the installation location information and device parameter information, from multiple PTZ cameras, identify the target camera whose installation location information and spatial location meet the preset distance conditions, and whose image size and image position corresponding to the spatial location meet the preset imaging conditions. Then, identify the device parameters required for the target camera to meet the preset imaging conditions as the corresponding target control parameters.
[0032] The above method, by considering the spatial location of the human figure within the shooting coverage area of each PTZ camera, selects the target camera that can capture the human figure corresponding to the core communication figure in the real world from the perspective of distance and angle. This results in video data showing the size and position of the core communication figure in the video frame without the need for image recognition or detection of multiple videos captured by multiple PTZ cameras, thus improving data processing speed.
[0033] Before acquiring the first video feed from the first camera and the location of the sound source confirmed by the microphone array, the process also includes:
[0034] Images are captured using multiple pan-tilt cameras to generate a three-dimensional spatial image of the target space.
[0035] Accordingly, based on the installation location information and device parameter information, from multiple PTZ cameras, the target camera whose installation location information and spatial location meet the preset distance conditions, and whose image size and image position corresponding to the spatial location meet the preset imaging conditions, is identified. The device parameters required for the target camera to meet the preset imaging conditions are then identified as the corresponding target control parameters, including:
[0036] The spatial position of the first camera in the first coordinate system is transformed to the second coordinate system according to the transformation relationship between the first coordinate system and the second coordinate system of the three-dimensional spatial image to obtain the second spatial position;
[0037] Based on the installation location information and equipment parameter information, confirm the image acquisition range of each PTZ camera in the second coordinate system;
[0038] When the candidate installation location information and the second spatial location meet the preset distance conditions, the second spatial location is within the image acquisition range corresponding to the candidate installation location information, the face is facing the PTZ camera corresponding to the candidate installation location information, and the image size and position of the second spatial location under a set of device parameters of the candidate installation location information meet the preset imaging conditions, the PTZ camera corresponding to the candidate installation location information is identified as the target camera, and a set of device parameters is identified as the corresponding target control parameters.
[0039] The above describes how the image acquisition results of different cameras in different coordinate systems on the same target space are transformed and mapped according to the coordinate system transformation relationship, so that the analysis results of the video data captured by the first camera can be accurately and quickly applied to the target video data captured by the target camera to meet the application requirements.
[0040] The video capture method also includes:
[0041] Upon receiving a target video request from an upper-layer application, the target video data is sent to the upper-layer application.
[0042] The above describes a method that uses multiple virtual camera drivers to comprehensively manage the underlying hardware consisting of multiple cameras. When an upper-layer application needs video data to present the live scene, the upper-layer application can directly request the target video data from the virtual camera drivers, thus enabling a good presentation of the local live scene for remote participants in multi-person communication.
[0043] Secondly, embodiments of this application provide a video acquisition device, which is applied to a smart device. The smart device includes a first camera and a microphone array. The smart device is also connected to multiple pan-tilt cameras located in the same target space and stores the installation location information and device parameter information of the multiple pan-tilt cameras. The video acquisition device includes:
[0044] The data acquisition unit is used to acquire the first video captured by the first camera and the location of the sound source confirmed by the microphone array;
[0045] The data processing unit is used to determine the facial orientation and spatial location of the human figure in the first video based on the location of the sound source and the first video stream.
[0046] The target confirmation unit is used to confirm the target camera that matches the face orientation and spatial position from multiple pan-tilt cameras, as well as the corresponding target control parameters, based on the installation location information and equipment parameter information.
[0047] The acquisition and control unit is used to control the target camera to capture target video data in the target space under the target control parameters, and to use the target video data as the video acquisition result of multiple pan-tilt cameras.
[0048] The data processing unit includes:
[0049] The area confirmation module is used to confirm the area to be identified in the video frame corresponding to the first video stream, pointing to the direction of the sound source.
[0050] The target confirmation module is used to detect the lip state of the face image of the region to be identified and confirm that the lips of the face target are in an active state.
[0051] The information confirmation module is used to confirm the face orientation and spatial position of the human target based on the human target and the preset parameter mapping relationship. The parameter mapping relationship is used to characterize the length ratio relationship in the imaging optical path of the first camera.
[0052] The data processing unit includes:
[0053] The trigger processing module is used to confirm the face orientation and spatial position of the portrait target when the deviation between the portrait target and the historical portrait target when the most recently confirmed face orientation and spatial position reaches a preset trigger threshold, based on the portrait target and the preset parameter mapping relationship.
[0054] The video capture device also includes:
[0055] The data caching unit is used to cache the portrait target as a historical portrait target when the deviation between the portrait target and the historical portrait target when the most recently confirmed face orientation and spatial position reaches a preset trigger threshold, based on the portrait target and the preset parameter mapping relationship, after confirming the face orientation and spatial position of the portrait target.
[0056] Correspondingly, the trigger processing module also includes:
[0057] The portrait comparison submodule is used to confirm the deviation between the portrait target and the cached historical portrait targets. When the deviation reaches a preset trigger threshold, the facial orientation and spatial position of the portrait target are confirmed according to the portrait target and the preset parameter mapping relationship.
[0058] The information confirmation module includes:
[0059] The orientation recognition submodule is used to recognize the facial orientation of a human target and confirm the orientation of the human target's face.
[0060] The depth confirmation submodule is used to confirm the size and position parameters of the human figure target, and to confirm the depth parameters of the human figure target based on the size parameters and parameter mapping relationship;
[0061] The location confirmation module is used to confirm the spatial location of the human figure target based on the location parameters and depth parameters.
[0062] The target confirmation unit includes:
[0063] The parameter matching module is used to identify, from multiple PTZ cameras, the target camera whose installation location and spatial location meet preset distance conditions, and whose image size and image position corresponding to the spatial location meet preset imaging conditions, based on the installation location information and device parameter information. The module then identifies the device parameters required for the target camera to meet the preset imaging conditions as the corresponding target control parameters.
[0064] The video capture device also includes:
[0065] The image preprocessing unit is used to acquire images through multiple pan-tilt cameras before acquiring the first video captured by the first camera and the location of the sound source confirmed based on the microphone array, so as to generate a spatial three-dimensional image of the target space.
[0066] Correspondingly, the parameter matching module includes:
[0067] The coordinate transformation submodule is used to transform the spatial position of the first camera in the first coordinate system to the second coordinate system to obtain the second spatial position, according to the transformation relationship between the first coordinate system and the second coordinate system of the three-dimensional spatial image.
[0068] The range confirmation submodule is used to confirm the image acquisition range of each PTZ camera in the second coordinate system based on the installation location information and equipment parameter information;
[0069] The range matching submodule is used to identify the pan-tilt camera corresponding to the alternative installation location information as the target camera and to identify a set of device parameters as the corresponding target control parameters when the alternative installation location information and the second spatial location meet the preset distance conditions, the second spatial location is within the image acquisition range corresponding to the alternative installation location information, the face is facing the pan-tilt camera corresponding to the alternative installation location information, and the image size and position of the second spatial location under a set of device parameters of the alternative installation location information meet the preset imaging conditions.
[0070] The video capture device also includes:
[0071] The data sending unit is used to send the target video data to the upper-layer application when it receives a target video request from the upper-layer application.
[0072] Thirdly, embodiments of this application provide a smart device, which includes:
[0073] One or more processors;
[0074] Memory, used to store one or more computer programs;
[0075] When one or more computer programs are executed by one or more processors, the smart device enables the video capture method as described in any of the first aspects.
[0076] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements a video capture method as described in any of the first aspects. Attached Figure Description
[0077] Figure 1 is a flowchart of the video acquisition method provided in an embodiment of this application.
[0078] Figure 2 is a schematic diagram of the hardware layout of the target space used in the embodiments of this application.
[0079] Figure 3 is a schematic diagram of the field of view confirmed by the microphone array in a smart device.
[0080] Figure 4 is a schematic diagram of a video frame of target video data from the second gimbal camera in Figure 2.
[0081] Figure 5 is a schematic diagram of a video frame of target video data from the third gimbal camera in Figure 2.
[0082] Figure 6 is a schematic diagram of the structure of the video acquisition device provided in the embodiment of this application.
[0083] Figure 7 is a schematic diagram of the structure of a smart device provided in an embodiment of this application. Detailed Implementation
[0084] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and not for limiting the scope of the application. Furthermore, it should be noted that, for ease of description, only the parts relevant to the present application are shown in the drawings, not the entire structure.
[0085] It should be noted that, due to space limitations, this application specification does not exhaustively list all possible implementation methods. Those skilled in the art should be able to conceive after reading this application specification that, as long as the technical features do not contradict each other, any combination of technical features can constitute an optional implementation method.
[0086] The embodiments of this application will be described in detail below.
[0087] In scenarios involving remote information transmission among multiple people, such as meetings, offices, and classrooms, smart devices (such as conference tablets and smart blackboards) can be configured to present the key information during communication. The cameras installed in these smart devices can also capture images of the scene, allowing remote participants to see the live feed, thus enhancing their immersion and communication experience.
[0088] The inventors, through detailed examination of existing smart devices' captured images, discovered that cameras fixed to these devices, once configured for a specific scenario, can only capture images from a single, fixed angle. The only adjustable parameter during capture is the focal length. In the context of diverse and dynamic multi-person communication scenarios, this fixed-angle capture cannot guarantee high-quality representation of the communication, impacting the communication experience. For example, in multi-person communication, a speaker typically acts as the core, delivering information verbally. The speaker's output serves as the focal point and is crucial for effective communication. The speaker may change position and orientation, and may even change their role, to ensure the effectiveness and progress of the communication. For participants, directly experiencing the scene through auditory reception and visual observation guarantees a superior information delivery experience. However, for participants in remote multi-person communication, they can only participate in the multi-person communication scenario through a basically fixed scene captured by a camera at a fixed angle. It is very likely that they will not be able to see the speaker's face as the speaker changes. When participants in different spaces participate in multi-person communication through video, this kind of image presentation has poor reproduction of the scene, resulting in poor immersion in remote participation and a poor multi-person communication experience.
[0089] To address the aforementioned technical issues, this application proposes using a first camera and microphone array integrated into a smart device to locate and detect the pose of a human figure in a target space. Multiple external PTZ cameras have already been calibrated in the target space, and the shooting range of each PTZ camera in the target space can be confirmed based on device parameter information. The application then identifies the target camera from among the multiple PTZ cameras that can produce good video capture of the human figure, and determines the target control parameters required for achieving good video capture. When shooting video using multiple PTZ cameras, the capture result of the target camera under the target control parameters is directly used. In multi-person communication scenarios, this approach can quickly detect changes in the key communication figure (speaker) and control the PTZ camera that provides good shooting results for the key communication figure to capture video, thereby adapting to the rich changes in the key communication figure at the communication scene, ensuring a good presentation of the key communication figure to remote users, and improving the communication experience based on video data.
[0090] Please refer to Figure 1, which is a flowchart of the video acquisition method provided in this embodiment. This video acquisition method is applied to a smart device, which can be implemented by software and / or hardware. The smart device can consist of two or more physical entities, or it can consist of a single physical entity. Currently, the smart device can be a smart blackboard used in teaching scenarios, a conference tablet or interactive tablet used in meeting scenarios, a laptop, or other smart devices. Specifically, the smart device and multiple cameras are arranged, and the physical space (e.g., a meeting room or classroom) used in this solution is defined as the target space. As shown in Figure 1, the video acquisition method includes, but is not limited to, steps S110-S140:
[0091] Step S110: Obtain the first video stream captured by the first camera and the location of the sound source confirmed based on the microphone array.
[0092] The smart device in this embodiment includes a first camera and a microphone array. The smart device is also connected to multiple PTZ cameras located in the same target space and stores the installation location information and device parameter information of the multiple PTZ cameras.
[0093] Figure 2 is a schematic diagram of the hardware layout of the target space 10 used in this embodiment of the application, including the smart device 15 described above and the associated gimbal camera. Referring to Figure 2, the target space 10 includes multiple gimbal cameras (Figure 2 exemplarily includes a first gimbal camera 11, a second gimbal camera 12, and a third gimbal camera 13, a total of three gimbal cameras) and a smart device 15. Each gimbal camera is connected to the smart device 15 via wired (e.g., USB) or wireless (e.g., WIFI) means to realize the transmission of data and commands between the smart device 15 and each gimbal camera.
[0094] Optionally, the smart device 15 may also include a display screen that displays information based on instructions from the processor. The display screen may be a Liquid Crystal Display (LCD), an LED display, an Organic Light-Emitting Diode (OLED) display, or a Flexible Light-Emitting Diode (FLED) display, etc. In one embodiment, the display screen may also integrate touch functionality; in this case, the display screen includes a display panel and a touch panel. The display panel is used to provide visual output. The touch panel may be a touch component that supports infrared touch, electromagnetic touch, capacitive touch, and / or resistive touch, etc.
[0095] The smart device 15 may also include one or more communication interfaces to enable communication with other devices. For example, the smart device 15 communicates with multiple pan-tilt cameras shown in Figure 2 via the communication interface.
[0096] In addition, the smart device 15 may also include components such as a power supply, a speaker, and physical buttons, but the embodiments do not limit this.
[0097] Based on the aforementioned hardware structure, the smart device 15 supports at least one type of operating system, such as Android, Windows, or Linux. The smart device 15 can install at least one application under this operating system. The installed application can be a built-in application of the operating system, an application downloaded from a backend server, or a third-party device. By running these applications, the smart device 15 can perform corresponding functions. In one embodiment, the smart device 15 has at least one application installed for implementing a video capture method. In this embodiment, the video data captured by multiple PTZ cameras is processed comprehensively, and one is selected as the comprehensive output of all PTZ cameras. Therefore, this application can be considered as a virtual driver for managing all PTZ cameras connected to the smart device 15.
[0098] In this embodiment, the gimbal camera can be a camera with a gimbal that allows it to rotate. The gimbal can support the camera's rotation in both horizontal and vertical directions, allowing the camera to freely adjust its angle within the gimbal's rotation range for shooting. Currently used gimbal cameras can be exemplarily PTZ cameras. PTZ cameras, in addition to controlling the overall movement of the camera (horizontal and vertical), can also adjust the distance of internal optical components, thereby enabling zooming (zoom) of the captured image. When using a PTZ camera for shooting, at least the Pan parameter, Tilt parameter, and Zoom parameter are required. The Pan parameter determines the horizontal position of the PTZ camera (which can be represented by an angle), i.e., the Pan parameter determines the position the PTZ camera should rotate to when rotating left or right in the horizontal direction; different Pan parameters correspond to different horizontal rotation positions. The Tilt parameter determines the vertical position of the PTZ camera (which can be represented by an angle), i.e., the Tilt parameter determines the position the PTZ camera should rotate to when rotating in the vertical direction; different Tilt parameters correspond to different vertical rotation positions. The Zoom parameter refers to the magnification and focal length of the PTZ camera; adjusting the Zoom parameter allows the PTZ camera to zoom in and out. Adjusting the Pan and Tilt parameters changes the shooting angle of the PTZ camera. Currently, setting appropriate Pan and Tilt parameters ensures the subject (e.g., a speaker) is positioned in the center of the video frame. Adjusting the Zoom parameter changes the size of the subject in the video frame. Setting appropriate Zoom parameters ensures the subject (e.g., a speaker) occupies a suitable size in the video frame while maintaining clarity, ensuring high-definition video (because the PTZ camera is a zoomable camera, it can still capture high-definition images after zooming). The PTZ camera can be installed in one or more directions of the target space 10; Figure 2 exemplifies installation in three directions.
[0099] In this embodiment, as shown in Figure 2, the camera also includes a first camera 151 integrated into the smart device 15. After being installed on the smart device 15, the first camera 151 typically cannot adjust its overall shooting angle; it primarily supports zoom and magnification adjustments. The first camera 151 can be a wide-angle fixed-focus camera, or other types of cameras, without limitation. The first camera 151 is generally installed on the upper edge of the user-facing side of the smart device 15. For example, in Figure 2, the first camera 151 is installed in the upper center of the user-facing surface of the smart device 15. In practical applications, the first camera 151 can also be installed in other locations, without limitation in this embodiment. The first camera 151 can typically capture video of most of the target space 10, and its video capture range at least covers the conference table 14 and the surrounding participant seating area. The video captured by the first camera 151 is defined as the first video stream.
[0100] For multiple PTZ cameras, although they capture video from different angles towards the same target space, each PTZ camera only captures the scene entering its lens. To ensure that the final captured video data can be processed under the same spatial reference, the smart device needs to know the installation location information (i.e., spatial coordinates in the target space) and device parameter information of each PTZ camera. The installation location information can be based on the smart device or a confirmed point in the target space, allowing the smart device to determine the accurate position of each PTZ camera. The device parameter information characterizes the shooting range capability supported by the PTZ camera, specifically including the rotatable angles in various directions and the adjustable range of the shooting range after the orientation is confirmed (e.g., limited by the focal length). The installation location information can be measured or calibrated during the installation of the PTZ camera; the device parameter information can be directly recorded as the hardware specifications of the PTZ camera (e.g., maximum horizontal adjustable angle, maximum vertical adjustable angle, maximum adjustable zoom range, etc.), or it can be obtained after confirming the installation location information by converting the hardware specifications to a spatial coordinate system with the target space as a reference.
[0101] As shown in Figure 2, the smart device 15 is also equipped with a microphone array 152. The microphone array 152 can collect the sound wave vibrations generated during the speaker's speech. At this time, using the Direction of Arrival (DOA) technology, based on the speaker's voice collected by the microphone array 152, the angle of the speaker relative to the microphone array 152 itself in the HFOV can be obtained. Here, HFOV refers to the horizontal field of view, which can reflect the range of sound that the microphone array 152 can collect in the horizontal direction. That is, the angle between the two edges of the range of sound that the microphone array 152 can collect in the horizontal direction can be used as the HFOV of the microphone array.
[0102] The angle of the speaker within the HFOV of the microphone array 152 can be understood as the angle between the speaker's horizontal field of view and the line connecting the edge (a pre-defined edge) within the horizontal acquisition range. For example, Figure 3 is a schematic diagram of the horizontal field of view of a microphone array 152 according to an embodiment of this application. Referring to Figure 3, the HFOV (i.e., the acquisition range in the horizontal direction where sound can be acquired) of the microphone array 152 corresponds to region 17. The speaker 16 is located within the acquisition range. When the speaker 16 speaks, the smart device 15 can determine the direction of the speaker 16 in the horizontal field of view based on the sound signal acquired by the microphone array 152, and this direction can be represented by an angle. This angle can be the angle α between the speaker 16 and the line connecting one edge (pre-defined) of region 17. This angle can then be considered as the angle of the speaker 16 within the HFOV of the microphone array 152. In this embodiment, the determined angle of the speaker 16 within the HFOV of the microphone array 152 is recorded as DOA information. Based on the DOA information, the direction of the speaker 16 relative to the microphone array 152 can be known, that is, a specific direction of the speaker 16 within the angular range corresponding to the HFOV can be determined.
[0103] For the smart device 15, it has pre-set installation position information for the microphone array 152, or pre-recorded angle range of the microphone array 152's HFOV relative to itself. Therefore, by clearly identifying a specific direction of the speaker within the angle range corresponding to the HFOV, it is also clear that the speaker has a specific direction in the target space 10 with itself as a reference. This specific direction can be used as the basis for subsequent control of the camera for tracking and shooting. In the specific implementation, there is no exclusive limitation on the specific representation method of the sound source location; either the microphone array 152 or the smart device 15 can be used as a reference, as long as the smart device 15 can confirm a clearly directional direction in the target space 10 based on the sound source location.
[0104] In the application scenario shown in Figure 2, in addition to the smart device 15 and the first PTZ camera 11 set on the first wall, the target space 10 also has electronic devices such as the second PTZ camera 12 and the third PTZ camera 13 set on the second and third walls adjacent to the first wall. The target space 10 also has non-electronic facilities such as a conference table 14. When multiple people communicate in the target space 10 based on various devices and facilities, the members participating in the multi-person communication sit around the conference table 14, and usually leave the end of the conference table 14 closest to the smart device 15 empty. When multiple people communicate, the smart device 15 displays the focus information. The members participating in the multi-person communication turn their heads to the left (e.g., members 16d and 16c in Figure 2), turn their heads to the right (e.g., members 16a and 16b in Figure 2), or look straight ahead (e.g., the members sitting at the end of the conference table 14 furthest from the smart device 15) to view the focus information displayed by the smart device 15. With the layout design corresponding to the needs of this scenario, as shown in Figure 2, the pan-tilt cameras are set on three walls, which can ensure video capture of any member participating in the multi-person communication on site. The subsequent steps will explain in detail how to confirm the shooting target and achieve accurate shooting of the shooting target.
[0105] Based on the above hardware architecture, all video data captured by the cameras can be sent to the virtual driver for centralized processing.
[0106] It should be understood that Figure 2 is merely an exemplary layout diagram of the equipment and facilities in the target space 10, used to present the general location and basic composition of the equipment and facilities in the target space 10. For example, in a real target space 10, the smart device 15 should be installed on the first wall, and the display screen of the display device 15 should be perpendicular to the tabletop of the conference table 14. However, in the layout diagram of Figure 2, to simultaneously present the first camera 151 and the microphone array 152 in the smart device 15, the display screen of the display device 15 is displayed parallel to the tabletop of the conference table 14. In addition, depending on the size of the target space 10 and the shooting coverage of the PTZ camera, if setting one PTZ camera on one wall (e.g., the second and third walls) cannot cover all participants sitting opposite the conference table 14, or cannot capture video footage of all participants facing the PTZ camera with their faces close to the PTZ camera centered on them, more PTZ cameras can be set on one wall to ensure that for any participant in the target space 10, at least one PTZ camera can capture corresponding video data centered on them.
[0107] Step S120: Based on the sound source location and the first video stream, confirm the face orientation and spatial position of the human target in the first video stream.
[0108] In the target space, the shooting range of the first camera 151 configured on the smart device can usually cover the area centered on the conference table. However, when the first camera 151 is fixedly installed on the smart device, participants sitting in different positions around the conference table can only be filmed by the first camera from different facial angles. If the first camera is used to film the participant who is currently the speaker, it may only capture their side profile. When the scene is sent to a remote location to be presented to remote participants, the remote participants cannot have the experience of the speaker conveying information to them from the front. Instead, they only feel like bystanders, resulting in a poor sense of immersion in remote multi-person communication.
[0109] In this embodiment, the first camera 151, configured based on the smart device, can cover and capture the faces of all or part of the participants sitting around the conference table. Combining the layout of the microphone array 152 shown in Figures 2 and 3, and the HFOV diagram of the microphone array 152, it can be seen that the DOA confirmed by the microphone array 152 is actually a certain direction in front of the smart device 15. Combining the layout of the first camera 151 shown in Figure 2, it can be seen that the first video captured by the first camera 151 is actually the image of most of the area in front of the smart device 15, that is, it basically covers the HFOV. Based on the above layout relationship, when the smart device 15 clearly knows the distribution positions and acquisition ranges of various electronic devices in the target space 10, it can use the spatial parameters of the target space 10 as a reference to confirm that the shooting range of the first camera 151 is exactly in the general area corresponding to the direction pointed to by the DOA in the first video. In other words, based on the first video captured by the first camera, the system further combines the sound source location confirmed by the microphone array. After confirming the speaker's general direction from the sound source location, facial recognition is performed on the first video stream in that general direction, along with location detection based on the sound source location, to obtain the current speaker's facial orientation and the spatial position of the speaker's face (or head) in the target space. The corresponding image area of the speaker in each frame of the first video stream is the image target. Specifically, after confirming the speaker in the first video stream, facial features, human shape features, shoulder features, and head features are extracted and analyzed to confirm the orientation. For example, a large number of samples with different facial orientations are pre-labeled, and these samples are used to extract features from the above dimensions, training a big data model capable of facial orientation recognition. This big data model is then used to confirm the facial orientation. Spatial location can also be identified and confirmed through a corresponding big data model; or by combining images captured by at least one PTZ camera with facial recognition, facial recognition can be used to confirm the head area in the first video stream, and a multi-camera system constructed with the first camera and at least one PTZ camera can be used to complete the spatial positioning of the heads of each participant in the target space, thus obtaining the spatial location.
[0110] In one optional implementation, because the first camera and microphone array continuously acquire data, they obtain the first video and the location of the sound source, respectively, and the first video and the location of the sound source can strictly correspond based on the same time axis. During the speaker's speech, the first video serves as a video frame (each frame) that records the faces (front or side) of all participants. Correspondingly, the speaker's image is also recorded in a certain area of the video frame. At this time, with the positional relationship of the first camera and microphone array in the same coordinate system (e.g., the first coordinate system of the first camera) pre-calibrated, the approximate area of the speaker's image in the first video frame (the area to be identified) can be determined based on the location of the sound source when the first camera acquires video data. Then, the approximate area is identified to confirm the image of the speaker whose lips are in an active state. The image of the speaker whose lips are in an active state at the sound source location can be confirmed as the image target corresponding to the speaker. The detection of whether the lips are in an active state can be achieved by comparing the lip area image with the image of normally closed lips through a pre-trained recognition model, or by dynamically recognizing several preceding frames of images. The specific method is not limited.
[0111] For the first video stream, after confirming the area to be identified in the video frame corresponding to the first video stream, the lip state of the face image in the area to be identified is detected to confirm that the lips are in an active state. Given that the face target has been identified, i.e., the area of the face target in the first video stream has been confirmed, considering that the imaging optical paths of various cameras have basic geometric relationships, and under these geometric relationships there are corresponding length ratios, these length ratios are recorded as parameter mapping relationships. Based on the face target and the preset parameter mapping relationships, the face orientation and spatial position of the face target can be confirmed. As described above, the first video stream basically covers the High Field of View (HFOV). The Direction of Attention (DOA) is used to indicate a specific location of the speaker within the HFOV. Since the smart device has already recorded various parameters of the microphone array and the first camera in the target space, the portion of the video frame corresponding to the first video stream that falls within the DOA's pointing range can be identified based on the direction indicated by the DOA. Because the speaker is necessarily in the direction indicated by the DOA, full-screen recognition of the first video stream can be avoided, improving recognition efficiency. This involves using the ratios of length indicators in the target image region to statistically based length indicators (such as the distance between the corners of the eyes, the distance between the eyeballs, and the head height) at the corresponding location on the real person, and the ratio of the distance from the camera's sensor to the lens to the distance between the real person and the camera. These two ratios must be equal. Given that the statistically based length indicators at the real person's location and the distance from the camera's sensor to the lens can be preset, and the length indicators in the target image region can be detected in real time, the distance between the real person and the camera (i.e., the depth parameter) can be quickly determined. Furthermore, once the planar position of the target image in the video frame is confirmed, adding the depth parameter confirms the target image's spatial position. Facial orientation can be determined by pre-labeling a large number of face images with different orientations, using these as training samples to train an orientation recognition model, and then inputting the target image into the model to identify its facial orientation. Alternatively, facial orientation can be based on statistics, specifically by statistically analyzing the proportions of faces with different orientations on the left and right sides of the image, and conversely, determining the facial orientation based on these proportions.
[0112] In the actual confirmation of face orientation and spatial position, recognition processing can be performed on each frame of the image. That is, for the first video obtained by encoding multiple consecutive frames of images with frames as the basic image unit, recognition and judgment can be performed on each frame of the first video, completely executing the entire process of confirming face orientation and spatial position, as well as further confirming the target camera and target control parameters. This allows the target camera and target control parameters to be adjusted to correspond to any subtle changes in the human face, ultimately ensuring that the human face is displayed at a precise size and as close to the center as possible in the video frame of the target video data.
[0113] In actual confirmation of facial orientation and spatial location, recognition processing can be performed as needed based on changes in the facial target. Specifically, for the first video stream obtained by encoding multiple consecutive frames (with frames as the basic image unit), each frame in the first video stream can be identified and judged. First, the facial target is confirmed. Then, the facial target is compared with the state at the most recent confirmation of facial orientation and spatial location—that is, the facial target in the current frame is compared with historical facial targets. Only when the deviation between the facial target and historical facial targets reaches a preset trigger threshold is the facial orientation and spatial location confirmed according to the preset parameter mapping relationship between the facial target and historical facial targets. If the deviation between the facial target and historical facial targets does not reach the preset trigger threshold, the facial target can be considered relatively close to historical facial targets. The video data captured by the previous target camera and target control parameters can then be used as the target video data, as this meets the display requirements for the target facial image.
[0114] Comparing a human target with historical human targets can be done by comparing the entire image frames corresponding to the two targets to determine the degree of overlap. If the non-overlapping area reaches a set trigger threshold relative to the size of the human target (e.g., 20% or 30% of the human target size), then a change in the target camera and target control parameters is considered necessary, requiring reconfirmation of the human target's orientation and spatial position as the basis for subsequent adjustments. Alternatively, comparing a human target with historical targets can involve recording only the absolute size and relative position of the human target within the image frame. When confirming deviations, the absolute sizes are compared first, then the overlap of the relative positions is compared, and finally, a comprehensive judgment is made to determine whether a preset trigger threshold has been reached. For example, if the speaker moves a certain distance away from the target camera, the target image will appear smaller when filmed using the original target control parameters. Similarly, if the speaker stands up or moves a certain distance to the side, the target image will deviate upwards or to the left or right from the center of the video frame when filmed using the original target control parameters. These deviations can be quantified by comparing the target image with historical target images. Correspondingly, the target control parameters can be adjusted by zooming in and out to ensure the target image's size within the target video data. Alternatively, the horizontal or vertical angle of the target camera can be adjusted, or even a different pan-tilt camera can be used as the target camera to film video data using the reconfirmed target control parameters to ensure the target image's position within the target video data. It should be understood that to ensure the accuracy of target image tracking, the face orientation and spatial position must be continuously acquired and compared. Furthermore, if either state changes to reach the corresponding trigger threshold, the new face orientation and spatial position must be confirmed.
[0115] Based on changes in the human subject, the process of confirming whether to reconfirm the target camera and target control parameters can effectively reduce the repeated confirmation of the same target camera and the repeated confirmation of the same or similar target control parameters, reduce unnecessary data processing volume, and at the same time ensure that when the human subject changes significantly, the pan-tilt camera can be quickly controlled to track and shoot, ensuring the presentation effect of the key person in the target video data.
[0116] To confirm each deviation of the human target, whenever the human target changes, causing changes in the target camera or target control parameters, the human target can be cached as a historical human target. This serves as a reference for whether to reconfirm the target camera or target control parameters during subsequent human target detection. When confirming a deviation, the deviation between the human target and the cached historical human target is directly confirmed. If the deviation reaches a preset trigger threshold, the face orientation and spatial position of the human target are confirmed according to the preset parameter mapping relationship, which can effectively improve the data processing speed.
[0117] In one optional implementation, the facial orientation and spatial position of the target image are determined based on a pre-defined parameter mapping relationship. This can be achieved by determining the facial orientation based on the target image itself, or by determining the depth parameters of the target image based on a relatively fixed size parameter and its associated parameter mapping relationship. Facial orientation recognition is performed on the target image to confirm its orientation, for example, through a pre-trained orientation recognition model. The size and position parameters of the target image are then confirmed, and the depth parameters are determined based on these size parameters and their mapping relationship. For example, considering that individual differences in head height and width are relatively small, and different target images use the same static index with minimal deviation, the depth parameters can be determined based on the basic length relationship of the imaging optical path. In the target space referenced by the coordinate system of the first camera, the spatial position of the target image can be determined based on the position and depth parameters. It should be understood that the depth parameters confirmed in this way are equivalent to an estimation of the depth parameters. However, considering that the estimation deviation is acceptable, and that the target camera covers the spatial location when shooting, it will also shoot a certain range of areas outside the spatial location, so it will at least not affect the basic recording requirements of the human image target in the final target video data.
[0118] Step S130: Based on the installation location information and device parameter information, identify the target camera that matches the face orientation and spatial position from multiple PTZ cameras, as well as the corresponding target control parameters.
[0119] After confirming the relevant information of the human target, the target camera and the target control parameters required for the target camera to capture the human target can be identified based on this information. For cameras in a target space, not all cameras can capture a clear frontal image of the speaker, but the orientation of the human target is clear, and the orientation of each camera is also clear. Therefore, the camera opposite the orientation of the human target can be identified as the target camera. The target control parameters required for the target camera to capture the human target (i.e., the speaker) may include Pan parameters and Tilt parameters, or Pan parameters, Tilt parameters, and Zoom parameters. In one embodiment, the target control parameters include Pan parameters, Tilt parameters, and Zoom parameters as an example.
[0120] The precise orientation of the target camera is primarily related to the Pan and Tilt parameters, and the shooting angle will vary depending on the camera's rotation position. When filming a speaker, the speaker needs to be positioned as centrally as possible within the video frame. However, the speaker cannot simultaneously maintain verbal expression while actively focusing on and adjusting their position within the video frame. Therefore, the required shooting angle is mainly determined by the speaker's location. In other words, the target camera needs to adjust its Pan and Tilt parameters accordingly to ensure the speaker is centered within the video frame.
[0121] In practical implementation, the process of multi-person communication is dynamic. For example, a speaker may adjust their posture or move their position, or even change speakers. During real-time video data acquisition, as the video data is processed, the target camera may change, and the target control parameters may need to be reconfirmed accordingly. However, the overall processing remains the same. Essentially, this solution can be a process of real-time confirmation of human targets during video data acquisition, and reconfirmation of the target camera and target control parameters when a new human target is identified. Matching target cameras are characterized by corresponding device identifiers and have corresponding communication addresses. The target control parameters to be sent to the target camera are transmitted according to the communication address assigned to the corresponding device identifier to complete the control.
[0122] In one optional implementation, the process of identifying a target camera that matches the face orientation and spatial position from multiple PTZ cameras involves confirming whether the PTZ camera can directly or nearly directly face the human face target, given that the PTZ camera's image acquisition range coverage (confirmed by device parameter information) and installation position (confirmed by installation position information) have been determined. This process ensures the human face target is displayed in the center or near the center of the video image. For a PTZ camera, a single control parameter is equivalent to capturing an image of a cone-shaped (or cone-like) space with itself as the apex. By adjusting the horizontal and vertical angles of the PTZ camera, it's equivalent to expanding the cone-shaped (or cone-like) space outwards according to the maximum adjustable angle (part of the device parameter information). This confirms the PTZ camera's image acquisition range directly within the target space. The space outside the image acquisition range is the image acquisition blind zone.
[0123] Since the image acquisition range of a PTZ camera can be digitally represented in a space, and a location in that space is known, it can be confirmed whether the PTZ camera can capture an object at that location. It can also be precisely determined whether that location is at the edge or center of the image acquisition range. If the location is at the edge of the image acquisition range or directly in the image acquisition blind zone, it is virtually impossible to guarantee that the human image will be in the center or near the center of the video frame in the final image. In this case, it can be confirmed that the PTZ camera is not the target camera. The relationship between the image acquisition range of other PTZ cameras and this location is then assessed. If the layout of multiple PTZ cameras can comprehensively cover the target space, the target camera can generally be identified. Finally, based on the device parameter information of the PTZ cameras, it can be determined at what angle and zoom level that the face at that location will appear appropriately positioned and sized in the video frame. The specific angle and zoom level parameters are the target control parameters. The above process, from determining the facial orientation and spatial location of the target person, to the device parameters and installation location of the PTZ camera, and finally confirming the PTZ camera and target control parameters that can capture the target person's size and position in the video frame, involves identifying the target camera from multiple PTZ cameras based on the installation location and device parameters. This means selecting the target camera whose installation location and spatial location meet preset distance conditions, and whose corresponding image size and position meet preset imaging conditions. The device parameters required for the target camera to meet these preset imaging conditions are then identified as the corresponding target control parameters. Given a known digitally represented spatial range and a location, the relative relationship between the location and the spatial range is confirmed through conventional digital comparisons. By examining the presence of the target person's spatial location within the shooting coverage area of each PTZ camera, the target camera that can capture the target person's corresponding key communication figure in the real world from the appropriate distance and angle, obtaining video data showing the key communication figure's size and position in the video frame, eliminates the need for image recognition or detection of multiple video streams from multiple PTZ cameras, thus improving data processing speed.
[0124] In another alternative implementation, when constructing the target space for multi-person communication—that is, before acquiring the first video feed from the first camera and the sound source location confirmed by the microphone array—multiple PTZ cameras can be used to capture images to generate a three-dimensional spatial image of the target space. Within this three-dimensional spatial image, the image capture ranges corresponding to each PTZ camera can be virtually marked, and additionally, the image capture ranges that allow a human figure to be displayed in the video frame can be virtually marked. Capturing images in the same space using multiple cameras with determined installation positions and generating a corresponding three-dimensional spatial image is a conventional multi-camera image processing approach, which will not be elaborated upon here. When confirming the target camera and target control parameters based on a spatial 3D image, the coordinate system used for the spatial position confirmed by the first camera and microphone array may be different from the coordinate system of the spatial 3D image. However, both represent the same target space. Therefore, it is necessary to first transform the coordinate systems of the two. That is, the spatial position of the first camera in the first coordinate system is transformed to the second coordinate system according to the transformation relationship between the first coordinate system and the second coordinate system of the spatial 3D image to obtain the second spatial position. Then, based on the installation position information and equipment parameter information, the image acquisition range of each pan-tilt camera in the second coordinate system is confirmed. When the candidate installation position information and the second spatial position meet the preset distance conditions, the second spatial position is within the image acquisition range corresponding to the candidate installation position information, the face is facing the pan-tilt camera corresponding to the candidate installation position information, and the image size and position of the second spatial position under a set of equipment parameters of the candidate installation position information meet the preset imaging conditions, the pan-tilt camera corresponding to the candidate installation position information is confirmed as the target camera, and a set of equipment parameters is confirmed as the corresponding target control parameters. In the virtual marker-based judgment, for the second spatial position in the second coordinate system that has been transformed into a 3D spatial image, the second spatial position is directly mapped onto the 3D spatial image. The pan-tilt camera associated with the corresponding virtual marker is the target camera. Then, the target control parameters can be confirmed based on the device parameter information. The image acquisition results of different cameras in different coordinate systems of the same target space are transformed and mapped according to the coordinate system transformation relationship. This allows the analysis results of the video data captured by the first camera to be accurately and quickly applied to the target video data captured by the target camera to meet the application requirements.
[0125] It should be understood that for multiple PTZ cameras in a target space, there may be two or more PTZ cameras that can act as target cameras to collect video data from human targets in a manner that meets the requirements of the video image. However, in actual implementation, it is sufficient to confirm that one target camera and its target control parameters can output one channel of video data that meets the requirements of the video image. It is not necessary to identify all possible target cameras and target control parameters.
[0126] Step S140: Control the target camera to capture the target space under the target control parameters to obtain target video data, and use the target video data as the video acquisition result of multiple PTZ cameras.
[0127] With the target camera and target control parameters confirmed, the intelligent device controls the target camera to capture target video data according to the target control parameters. This provides video data that effectively reflects the speaker's perspective in the current multi-person communication scenario. If an upper-layer application (such as various video conferencing applications, live streaming applications, recording applications, etc.) requires live video and requests video data from the virtual camera driver, that is, when the virtual camera driver receives the target video request from the upper-layer application, the target camera sends the captured target video data to the virtual camera driver, which then sends the target video data to the upper-layer application. In this application scenario, the upper-layer application primarily sends the target video data to remote devices for remote participants to view. The virtual camera driver can directly provide the target video data to the upper-layer application without processing it. In this case, the upper-layer application's use of the target video data mainly involves sending it to a remote upper-layer application that communicates with it via the network for presentation, providing a good immersive experience for remote users participating in multi-person communication.
[0128] The following description, in conjunction with Figures 2, 4, and 5, illustrates the video acquisition process of an embodiment of this application in a specific application scenario, and thus describes the implementation process of this embodiment.
[0129] Please refer to Figure 2. In a target space 10, there are a first gimbal camera 11, a second gimbal camera 12, a third gimbal camera 13, and a smart device 15. The smart device 15 includes a first camera 151 and a microphone array 152. In a scenario involving remote multi-person communication, there are participants 16a, 16b, 16c, and 16d in the target space 10, who are seated around the conference table 14 as shown in Figure 2. During multi-person communication, participants 16d and 16b take turns as speakers, speaking on the communication focus information displayed by the smart device 15. When speaking, participants 16d and 16b usually face the smart device 15. At this time, the first camera 151 can capture participants 16d and 16b. However, because participants 16d and 16b are not directly facing the light-receiving surface of the first camera 151, the first video actually captures the side profiles of participants 16d and 16b. Moreover, since the first camera 151 can only zoom, participants 16d and 16b will be off-center in the video frame of the first video. If the first video is sent to a remote location to present the state of participants 16d and 16b when speaking, the remote participants will only perceive that there are participants off-center from the center of the video frame speaking in the global view of the target space 10, and the speakers are not transmitting information to them face-to-face.
[0130] Based on the specific implementation method described above, in remote multi-person communication scenarios based on video transmission, it is no longer necessary to directly use the first video stream to present the scene in the target space 10 to the remote end. Instead, based on the sound source location confirmed by the microphone array 152 and the content of the scene in the first video stream, the actual position of the speaker in the target space 10 is confirmed. Then, with the installation position information and equipment parameter information of multiple pan-tilt cameras in the target space 10 known, a target camera with a good shooting distance and shooting angle for the speaker is selected from the multiple pan-tilt cameras. In the target video data obtained by the target camera from the video capture of the area where the speaker is located, the human image target (the speaker's image) is located or close to the center area of the video screen, and the size of the human image target in the video screen is appropriate. When the target video data is sent to the remote participants for viewing, the remote participants can get a feeling of the speaker conveying information to them face to face, and the immersion of the remote participants during the video conference is better.
[0131] Taking participants 16d and 16b as speakers in Figure 2 as an example, during multi-person communication, the first camera 151 and microphone array 152 continuously collect data, obtaining the first video and sound source location respectively, and the first video and sound source location can be strictly corresponding based on the same time axis. During participant 16d's speech, the first video, as a video frame that can record the faces (front or side) of all participants, also records the image of participant 16d in a certain area of the video frame. At this time, with the positional relationship of the first camera 151 and microphone array 152 first calibrated in the same coordinate system, the approximate area of participant 16d's image in the first video frame can be determined based on the sound source location when the first camera 151 collects video data. Then, the approximate area is identified to confirm the image of the person whose lips are in an active state. The image of the person whose lips are in an active state at the sound source location can be confirmed as the image target corresponding to the speaker.
[0132] For a confirmed human target, the speaker's coordinates in both horizontal and vertical dimensions relative to the first camera 151 can be directly confirmed using the first video feed. Additionally, based on the general imaging principles of cameras, the length ratios of various segments in the imaging optical path can be determined to establish the speaker's depth parameters relative to the first camera 151. Thus, the speaker's spatial position is confirmed with the first camera 151 as a reference. Furthermore, the face orientation can be determined based on the proportions of the left and right sides of the human target. With the face orientation and corresponding spatial position confirmed, along with the installation location and equipment parameters of the PTZ camera, the speaker's position and orientation can be matched with the PTZ camera's shooting range. If the speaker is facing a relatively central position within the PTZ camera's shooting range, and a suitable shooting distance is maintained at this central position, resulting in the human target being centered and appropriately sized in the video feed, then this PTZ camera is the target camera.
[0133] For example, when participant 16d is the speaker, the data collected by the smart device 15 through the first camera 151 and the microphone array 152 can confirm that the spatial position of participant 16d is roughly at the upper left end of the conference table 14 (the bottom layer represents the spatial position as spatial coordinates), and the face is facing towards the smart device 15. At this time, the spatial position and face orientation are matched with the shooting range of each pan-tilt camera. The result is that participant 16d is facing away from the third pan-tilt camera 13, and the whole body is directly in front of the second pan-tilt camera 12, but the face orientation is more towards the first pan-tilt camera 11. Accordingly, the first pan-tilt camera 11 is used as the target camera. Furthermore, the first gimbal camera 11 can be adjusted in both the horizontal and vertical directions. In other words, the first gimbal camera 11 can adapt to the spatial position within the range of its own device parameter information, so that the shooting direction is as directly facing the participant's face as possible and centered. At the same time, it can zoom and adjust the magnification according to the installation position of the first gimbal camera 11 and the distance between the spatial position, so that the size ratio of the human image target in the corresponding video screen is appropriate. The various angle information and zoom and adjustment information of the first gimbal camera 11 confirmed in this state are the target control information.
[0134] Based on the target control information, the target camera (i.e., the first PTZ camera 11) is controlled to capture video of the spatial location (i.e., the location of participant 16d), resulting in the video image shown in Figure 4. Participant 16d's image is centered and appropriately sized in the video image. When this video image is displayed to a distant participant, it provides the distant participant with an immersive experience of the speaker directly addressing them. The PTZ camera is typically installed near the top of the target space 10, and accordingly, it is usually positioned diagonally downwards to capture the participant. In this case, other PTZ cameras located at the top of the target space 10 and at the same height as the participant may not be captured in the target video data. Figure 4 shows the layout relationship between the third PTZ camera 13 (represented by dashed lines) and the corresponding image capture range of the video image. During participant 16d's speech, the first video and sound source location are continuously processed. It can be observed that the human image target remains essentially unchanged. At this point, it is not necessary to confirm new target cameras and target control parameters; the video capture can be performed using the most recently confirmed target control parameters.
[0135] As the multi-person communication process progresses, after participant 16d finishes speaking, participant 16b continues to speak. While continuously processing the first video and sound source location, it can be found that the new human image target deviates significantly from the human image target when the target camera and target control parameters were confirmed last time, indicating that the speaker has changed. At this time, the target camera and target control parameters can be reconfirmed according to the processing procedure when participant 16d spoke. Accordingly, the third gimbal camera 13 is confirmed as the target camera. When the third gimbal camera 13 performs video acquisition based on the corresponding target control parameters, the second gimbal camera 12 is located outside the image acquisition range corresponding to the third gimbal camera 13. Figure 5 shows the layout relationship between the second gimbal camera 12 and the corresponding image acquisition range of the video screen, represented by the dashed line.
[0136] Based on the above processing, although multiple PTZ cameras are managed through a virtual camera driver, only one PTZ camera collects target video data at a time as the overall video acquisition result. In other words, the overall video acquisition result may be a video stream obtained by combining footage captured by multiple PTZ cameras based on changes in the speaker's posture. For example, in the target space 10 shown in Figure 2, participants 16d and 16b speak in turn while maintaining their respective postures. The final video acquisition result is a video with essentially the same composition as the video frame shown in Figure 4, and a video with essentially the same composition as the video frame shown in Figure 5, spliced together sequentially.
[0137] It should be understood that even if the speaker's identity does not change, the speaker's movement or change in facial orientation may cause the target camera or target control parameters to need to be adjusted accordingly. The specific adjustment logic is the same as the adjustment logic when the speaker's identity changes, and will not be explained in detail here.
[0138] Figure 6 is a schematic diagram of a video acquisition device provided in an embodiment of this application. This video acquisition device is applied to a smart device, which includes a first camera and a microphone array. The smart device is also connected to multiple pan-tilt cameras located in the same target space and stores the installation location information and device parameter information of the multiple pan-tilt cameras. Referring to Figure 6, the video acquisition device provided in this embodiment of the application includes a data acquisition unit 210, a data processing unit 220, a target confirmation unit 230, and an acquisition control unit 240.
[0139] The data acquisition unit 210 is used to acquire the first video stream captured by the first camera and the sound source location confirmed by the microphone array; the data processing unit 220 is used to confirm the face orientation and spatial position of the human target in the first video stream based on the sound source location and the first video stream; the target confirmation unit 230 is used to confirm the target camera that matches the face orientation and spatial position from multiple PTZ cameras, as well as the corresponding target control parameters, based on the installation location information and device parameter information; and the acquisition control unit 240 is used to control the target camera to capture the target space under the target control parameters to obtain target video data, and to use the target video data as the video acquisition result of multiple PTZ cameras.
[0140] Based on the above embodiments, the data processing unit 220 includes:
[0141] The area confirmation module is used to confirm the area to be identified in the video frame corresponding to the first video stream, pointing to the direction of the sound source.
[0142] The target confirmation module is used to detect the lip state of the face image of the region to be identified and confirm that the lips of the face target are in an active state.
[0143] The information confirmation module is used to confirm the face orientation and spatial position of the human target based on the human target and the preset parameter mapping relationship. The parameter mapping relationship is used to characterize the length ratio relationship in the imaging optical path of the first camera.
[0144] Based on the above embodiments, the data processing unit 220 includes:
[0145] The trigger processing module is used to confirm the face orientation and spatial position of the portrait target when the deviation between the portrait target and the historical portrait target when the most recently confirmed face orientation and spatial position reaches a preset trigger threshold, based on the portrait target and the preset parameter mapping relationship.
[0146] Based on the above embodiments, the video acquisition device further includes:
[0147] The data caching unit is used to cache the portrait target as a historical portrait target when the deviation between the portrait target and the historical portrait target when the most recently confirmed face orientation and spatial position reaches a preset trigger threshold, based on the portrait target and the preset parameter mapping relationship, after confirming the face orientation and spatial position of the portrait target.
[0148] Correspondingly, the trigger processing module also includes:
[0149] The portrait comparison submodule is used to confirm the deviation between the portrait target and the cached historical portrait targets. When the deviation reaches a preset trigger threshold, the facial orientation and spatial position of the portrait target are confirmed according to the portrait target and the preset parameter mapping relationship.
[0150] Based on the above embodiments, the information confirmation module includes:
[0151] The orientation recognition submodule is used to recognize the facial orientation of a human target and confirm the orientation of the human target's face.
[0152] The depth confirmation submodule is used to confirm the size and position parameters of the human figure target, and to confirm the depth parameters of the human figure target based on the size parameters and parameter mapping relationship;
[0153] The location confirmation module is used to confirm the spatial location of the human figure target based on the location parameters and depth parameters.
[0154] Based on the above embodiments, the target confirmation unit 230 includes:
[0155] The parameter matching module is used to identify, from multiple PTZ cameras, the target camera whose installation location and spatial location meet preset distance conditions, and whose image size and image position corresponding to the spatial location meet preset imaging conditions, based on the installation location information and device parameter information. The module then identifies the device parameters required for the target camera to meet the preset imaging conditions as the corresponding target control parameters.
[0156] Based on the above embodiments, the video acquisition device further includes:
[0157] The image preprocessing unit is used to acquire images through multiple pan-tilt cameras before acquiring the first video captured by the first camera and the location of the sound source confirmed based on the microphone array, so as to generate a spatial three-dimensional image of the target space.
[0158] Correspondingly, the parameter matching module includes:
[0159] The coordinate transformation submodule is used to transform the spatial position of the first camera in the first coordinate system to the second coordinate system to obtain the second spatial position, according to the transformation relationship between the first coordinate system and the second coordinate system of the three-dimensional spatial image.
[0160] The range confirmation submodule is used to confirm the image acquisition range of each PTZ camera in the second coordinate system based on the installation location information and equipment parameter information;
[0161] The range matching submodule is used to identify the pan-tilt camera corresponding to the alternative installation location information as the target camera and to identify a set of device parameters as the corresponding target control parameters when the alternative installation location information and the second spatial location meet the preset distance conditions, the second spatial location is within the image acquisition range corresponding to the alternative installation location information, the face is facing the pan-tilt camera corresponding to the alternative installation location information, and the image size and position of the second spatial location under a set of device parameters of the alternative installation location information meet the preset imaging conditions.
[0162] Based on the above embodiments, the video acquisition device further includes:
[0163] The data sending unit is used to send the target video data to the upper-layer application when it receives a target video request from the upper-layer application.
[0164] The video acquisition device provided in this application embodiment can be used to execute the video acquisition method provided in the above embodiment, and has corresponding functions and beneficial effects.
[0165] Figure 7 is a schematic diagram of the structure of a smart device provided in an embodiment of this application. As shown in Figure 7, the smart device includes a processor 310 and a memory 320, and may also include an input device 330, an output device 340, and a communication device 350. The number of processors 310 in the smart device can be one or more, and Figure 7 shows one processor 310 as an example. The processor 310, memory 320, input device 330, output device 340, and communication device 350 in the smart device can be connected through a bus or other means, and Figure 7 shows a connection via a bus as an example.
[0166] The memory 320, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as the program instructions / modules corresponding to the video acquisition method in this embodiment. The processor 310 executes various functional applications and data processing of the smart device by running the software programs, instructions, and modules stored in the memory 320, thereby realizing the aforementioned video acquisition method.
[0167] The memory 320 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on the use of the smart device. Furthermore, the memory 320 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 320 may further include memory remotely located relative to the processor 310, which can be connected to the smart device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0168] Input device 330 can be used to receive input digital or character information, and to generate signal inputs related to user settings and function control of the smart device. Output device 340 may include display devices such as a display screen.
[0169] The aforementioned intelligent devices can be used to execute any video capture method, possessing corresponding functions and beneficial effects.
[0170] This application also provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program performs relevant operations in the video acquisition method provided in any embodiment of this application and has corresponding functions and beneficial effects.
[0171] Those skilled in the art will understand that embodiments of this application may be provided as methods, systems, or computer program products.
[0172] Therefore, this application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0173] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0174] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0175] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0176] Note that the above description is merely a preferred embodiment and the technical principles employed in this application. Those skilled in the art will understand that this application is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this application. Therefore, although this application has been described in detail through the above embodiments, this application is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this application, and the scope of this application is determined by the scope of the appended claims.
Claims
1. Video capture methods, applied to smart devices, among which, The intelligent device includes a first camera and a microphone array. The intelligent device is also connected to multiple PTZ cameras located in the same target space and stores the installation location information and device parameter information of the multiple PTZ cameras. The video acquisition method includes: Acquire the first video stream captured by the first camera and the location of the sound source confirmed based on the microphone array; Based on the location of the sound source and the first video stream, the facial orientation and spatial position of the human figure in the first video stream are confirmed. Based on the installation location information and device parameter information, a target camera matching the face orientation and spatial position, as well as the corresponding target control parameters, are identified from the plurality of PTZ cameras; The target camera is controlled to capture video data of the target space under the target control parameters, and the target video data is used as the video acquisition result of the multiple PTZ cameras.
2. The video acquisition method according to claim 1, wherein, The step of determining the facial orientation and spatial location of the human target in the first video stream based on the sound source location and the first video stream includes: Confirm the area to be identified in the video frame corresponding to the first video stream, where the sound source is pointing. Perform lip state detection on the face image of the region to be identified to confirm that the lips of the portrait target are in an active state. Based on the portrait target and the preset parameter mapping relationship, the face orientation and spatial position of the portrait target are confirmed. The parameter mapping relationship is used to characterize the length ratio relationship in the imaging optical path of the first camera.
3. The video acquisition method according to claim 2, wherein, The step of determining the facial orientation and spatial position of the portrait target based on the portrait target and the preset parameter mapping relationship includes: If the deviation between the portrait target and the historical portrait target at the time of the most recent confirmation of face orientation and spatial position reaches a preset trigger threshold, the face orientation and spatial position of the portrait target are confirmed according to the portrait target and the preset parameter mapping relationship.
4. The video acquisition method according to claim 3, wherein, When the deviation between the portrait target and the historical portrait target at the time of the most recent confirmed face orientation and spatial position reaches a preset trigger threshold, after confirming the face orientation and spatial position of the portrait target according to the portrait target and the preset parameter mapping relationship, the method further includes: Cache the portrait target as the historical portrait target; Accordingly, when the deviation between the portrait target and the historical portrait target at the time of the most recent confirmation of face orientation and spatial location reaches a preset trigger threshold, the face orientation and spatial location of the portrait target are confirmed according to the portrait target and the preset parameter mapping relationship, including: The deviation between the portrait target and the cached historical portrait target is confirmed. If the deviation reaches a preset trigger threshold, the face orientation and spatial position of the portrait target are confirmed according to the portrait target and the preset parameter mapping relationship.
5. The video acquisition method according to claim 2, wherein, The step of determining the facial orientation and spatial position of the portrait target based on the portrait target and the preset parameter mapping relationship includes: Facial orientation recognition is performed on the portrait target to confirm the facial orientation of the portrait target; Confirm the size and position parameters of the human figure target, and confirm the depth parameters of the human figure target based on the size parameters and the parameter mapping relationship; The spatial location of the human figure target is determined based on the position and depth parameters.
6. The video acquisition method according to any one of claims 1-5, wherein, The step of identifying a target camera that matches the face orientation and spatial position from the plurality of pan-tilt cameras based on the installation location information and device parameter information, and the corresponding target control parameters, includes: Based on the installation location information and device parameter information, from the plurality of PTZ cameras, identify the target camera whose installation location information and spatial location meet the preset distance conditions, and whose image size and image position corresponding to the spatial location meet the preset imaging conditions, and confirm the device parameters required for the target camera to meet the preset imaging conditions as the corresponding target control parameters.
7. The video acquisition method according to claim 6, wherein, Before acquiring the first video stream captured by the first camera and the location of the sound source confirmed based on the microphone array, the method further includes: Images are acquired using the multiple PTZ cameras to generate a three-dimensional spatial image of the target space; Accordingly, based on the installation location information and device parameter information, the step of identifying a target camera from among the plurality of PTZ cameras whose installation location information and spatial location meet a preset distance condition, and whose image size and image position corresponding to the spatial location meet preset imaging conditions, and confirming the device parameters required for the target camera to meet the preset imaging conditions as the corresponding target control parameters, includes: The spatial position of the first camera in the first coordinate system is transformed to the second coordinate system according to the transformation relationship between the first coordinate system and the second coordinate system of the spatial three-dimensional image to obtain the second spatial position; Based on the installation location information and device parameter information, the image acquisition range of each PTZ camera in the second coordinate system is confirmed; When the candidate installation location information and the second spatial location meet a preset distance condition, the second spatial location is within the image acquisition range corresponding to the candidate installation location information, the face orientation is the pan-tilt camera corresponding to the candidate installation location information, and the image size and position of the second spatial location under a set of device parameters of the candidate installation location information meet preset imaging conditions, the pan-tilt camera corresponding to the candidate installation location information is identified as the target camera, and the set of device parameters is identified as the corresponding target control parameters.
8. The video acquisition method according to any one of claims 1-5, wherein, Also includes: Upon receiving a target video request from an upper-layer application, the target video data is sent to the upper-layer application.
9. Video capture device, used in smart devices, among which, The intelligent device includes a first camera and a microphone array. The intelligent device is also connected to multiple PTZ cameras located in the same target space and stores the installation location information and device parameter information of the multiple PTZ cameras. The video acquisition device includes: The data acquisition unit is used to acquire the first video captured by the first camera and the location of the sound source confirmed based on the microphone array; The data processing unit is used to determine the facial orientation and spatial position of the human figure in the first video based on the location of the sound source and the first video stream. The target confirmation unit is used to confirm, based on the installation location information and device parameter information, a target camera that matches the face orientation and spatial position from the plurality of PTZ cameras, as well as the corresponding target control parameters; The acquisition control unit is used to control the target camera to capture target video data in the target space under the target control parameters, and to use the target video data as the video acquisition result of the multiple PTZ cameras.
10. Smart devices, among which, include: One or more processors; Memory, used to store one or more computer programs; When the one or more computer programs are executed by the one or more processors, the smart device implements the video capture method as described in any one of claims 1-8.
11. A computer-readable storage medium having a computer program stored thereon, wherein, When executed by a processor, the computer program implements the video capture method as described in any one of claims 1-8.
Citation Information
Patent Citations
Intelligent video director method based on microphone array sound guidance
CN101567969A
Automatic capture type intelligent conference shooting system
CN108513063A
Image display method, device, system and equipment and readable storage medium
CN108900787A
Camera control method and display equipment
CN111669508A
Techniques and system for automatic video conference camera feed selection based on room events
US20120293606A1