Photographing control method and apparatus, and smart device, system and storage medium
By configuring a rotatable second camera on a smart device to work in conjunction with the built-in first camera, and automatically adjusting the shooting control parameters, the problem that a fixed camera cannot capture the central area of the speaker is solved, thus improving the clarity of the video and the quality of communication.
Patent Information
- Application Number
- PCT/CN2024/101668
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2026-01-02
AI Technical Summary
Because the camera on a smart device has a fixed shooting angle, it cannot capture video footage of the speaker in the center when the speaker is at the edge of the shooting space, resulting in a poor video communication experience.
The system employs a rotatable second camera in conjunction with the built-in first camera of the smart device. By recognizing the speaker's pixel coordinates and face size, it automatically adjusts the shooting control parameters of the second camera to achieve accurate shooting of the speaker in the center area.
It improves the clarity of video footage, ensuring a clear close-up of the speaker in the center, thus enhancing the quality of video communication. Furthermore, it eliminates the need for manual camera adjustments, simplifying the operation process.
Smart Images

Figure CN2024101668_02012026_PF_FP_ABST
Abstract
Description
Photographing control method and device, intelligent device, system, and storage medium TECHNICAL FIELD
[0001] Embodiments of the present application relate to the technical field of image processing, and in particular to a photographing control method and device, an intelligent device, a system, and a storage medium. BACKGROUND
[0002] Nowadays, intelligent technology has been widely applied in daily work and study scenes such as offices and classrooms. For example, intelligent devices such as conference tablets or smart blackboards are often used in business meetings or classroom teaching. One of the common functions of intelligent devices is to use the camera configured on the intelligent device to take pictures to realize video communication or video recording. For example, in a meeting or teaching scenario, the current speaker is photographed to enable remote participants to see the speaker's speaking picture.
[0003] In related technologies, the camera configured on the intelligent device is usually a non-rotatable camera. At this time, the photographing angle of the camera is fixed. When the speaker is located at the edge of the photographing space, the speaker photographed by the camera can only be in the edge area of the picture and cannot be located in the middle area of the picture, which greatly affects the photographing effect of the intelligent device when photographing the speaker and reduces the video communication experience.
[0004] SUMMARY
[0005] Embodiments of the present application provide a photographing control method and device, an intelligent device, a system, and a storage medium to solve the technical problem that the camera configured on the intelligent device in related technologies cannot photograph the video picture of the speaker located in the middle area when the speaker is located at the edge of the photographing space due to the fixed photographing angle.
[0006] In a first aspect, an embodiment of the present application provides a photographing control method applied to an intelligent device. The intelligent device includes a first camera, a display screen, a virtual camera driver, and a target application program. The intelligent device is connected with an external second camera. The second camera is a rotatable camera. The photographing control method includes the following steps.
[0007] Obtaining first video data photographed by the first camera, the first camera being configured to photograph a first photographing space;
[0008] Identifying a first pixel coordinate of a speaker and a first face size in a first video picture currently processed in the first video data, the speaker being located in the first photographing space;
[0009] Determining a first photographing control parameter of the second camera according to the first pixel coordinate and the first face size;
[0010] control the second camera to shoot using the first shooting control parameter, so that the second camera realizes the aligned shooting of the speaker under the first shooting control parameter;
[0011] obtain the second video data shot by the second camera;
[0012] obtain the second video data by the virtual camera driver program, and send the second video data to the target application program, so that the target application program can use the second video data on the display screen.
[0013] The above, obtaining the first video data shot by the first camera, and identifying the first pixel coordinates of the speaker in the first video data in the first video frame and the first face size, then determining the first shooting control parameter of the second camera according to the first pixel coordinates and the first face size, controlling the second camera to use the first shooting control parameter, obtaining the second video data shot by the second camera, then the virtual camera driver program can obtain the second video data and send the second video data to the target application program to use the second video data on the display screen. The technical means solves the technical problem that the camera configured on the intelligent device in the related art cannot shoot the video frame of the speaker in the middle area when the speaker is located at the edge of the shooting space because the shooting angle is fixed. By using the rotatable second camera and the intelligent device to cooperate with each other, the first camera configured by the intelligent device shoots the panoramic frame in the first shooting space, and the second camera shoots the close-up frame of the speaker in the middle area. Then, the video data (second video data) shot by the second camera is obtained by the virtual camera driver program and sent to the target application program, so that the target application program can obtain a clear speaker speaking frame, and the clarity of the video frame after network transmission is improved when the video frame needs to be transmitted. Moreover, the intelligent device can automatically determine the shooting control parameter of the second camera, without manually adjusting the second camera, so as to realize automatic intelligent adjustment, improve the alignment speed of the second camera, and ensure the shooting quality of the second camera. Moreover, without changing the hardware design of the intelligent device, only the software processing logic needs to be modified, so that the shooting control method is easy to implement. Moreover, by designing the virtual camera driver program, the target application program does not need to care about the number of underlying cameras and the control logic, and the number of underlying cameras can be easily expanded.
[0014] In an embodiment of the present application, the first shooting control parameter includes a first Pan parameter, a first Tilt parameter, and a first Zoom parameter;
[0015] The first shooting control parameter of the second camera is determined according to the first pixel coordinates and the first face size, specifically including:
[0016] determine a first Pan parameter corresponding to the first pixel coordinate according to a first correspondence relationship between pixel coordinates in the video data captured by the first camera and Pan parameters of the second camera;
[0017] determine a first Tilt parameter corresponding to the first pixel coordinate according to a second correspondence relationship between pixel coordinates in the video data captured by the first camera and Tilt parameters of the second camera;
[0018] determine a first Zoom parameter corresponding to the first face size according to a third correspondence relationship between face sizes in the video data captured by the first camera and Zoom parameters of the second camera.
[0019] In an embodiment of the present application, before the first video data captured by the first camera is obtained, the method further comprises:
[0020] establishing a first correspondence relationship between pixel coordinates in the video data captured by the first camera and Pan parameters of the second camera, and establishing a second correspondence relationship between pixel coordinates in the video data captured by the first camera and Tilt parameters of the second camera;
[0021] establishing a third correspondence relationship between face sizes in the video data captured by the first camera and Zoom parameters of the second camera.
[0022] In an embodiment of the present application, the first correspondence relationship between pixel coordinates in the video data captured by the first camera and Pan parameters of the second camera, and the second correspondence relationship between pixel coordinates in the video data captured by the first camera and Tilt parameters of the second camera are established, specifically comprising:
[0023] obtaining a second video frame based on the first camera, wherein the first camera captures a static second shooting space;
[0024] dividing the second video frame into a plurality of first video sub-frames, and determining a second pixel coordinate of a center point of each first video sub-frame in the second video frame;
[0025] obtaining a second Pan parameter and a second Tilt parameter corresponding to each of the first sub-video pictures, a third video picture captured by the second camera using the second Pan parameter, the second Tilt parameter and a second Zoom parameter has a similarity to a corresponding first sub-video picture reaching a similarity threshold, and the second Zoom parameter is a fixed parameter when the second camera captures the second capturing space;
[0026] determining a first correspondence relationship between pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera according to the second pixel coordinates and the second Pan parameter corresponding to each of the first sub-video pictures, and determining a second correspondence relationship between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera according to the second pixel coordinates and the second Tilt parameter corresponding to each of the first sub-video pictures.
[0027] In an embodiment of the present application, the establishment of the first correspondence relationship between the pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera and the establishment of the second correspondence relationship between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera specifically include:
[0028] obtaining an eighth video picture captured by the first camera and a ninth video picture captured by the second camera at the same capturing moment;
[0029] when a similarity between the ninth video picture and a fifth sub-video picture in the eighth video picture reaches a similarity threshold, obtaining a fifth pixel coordinate of a center point of the fifth sub-video picture in the eighth video picture, and obtaining a fourth Pan parameter and a fourth Tilt parameter used by the second camera when capturing the ninth video picture, the second camera using a fixed fourth Zoom parameter, and the fifth sub-video picture being one of the sub-video pictures obtained by dividing the eighth video picture;
[0030] determining the first correspondence relationship between the pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera according to a plurality of fifth pixel coordinates and a plurality of corresponding fourth Pan parameters, and determining the second correspondence relationship between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera according to a plurality of fifth pixel coordinates and a plurality of corresponding fourth Tilt parameters, the plurality of fifth pixel coordinates corresponding to a plurality of fifth sub-video pictures obtained at a plurality of capturing moments, and the pixel coordinate ranges of the fifth sub-video pictures being different.
[0031] In one embodiment of this application, establishing a third correspondence between the face size in the video data captured by the first camera and the Zoom parameter of the second camera specifically includes:
[0032] The system sequentially acquires multiple frames of fourth video images captured by the first camera and determines the second face size of the target face in each frame of the fourth video image. The first camera captures images in a third shooting space, which includes the target face. The second face size corresponding to each frame of the fourth video image is different.
[0033] Each time a frame of the fourth video image is acquired, the third Zoom parameter corresponding to the fourth video image is also acquired. The third face size of the target face in the fifth video image captured by the second camera using the third Zoom parameter, the third Pan parameter and the third Tilt parameter satisfies the preset face size.
[0034] Based on the second face size and third Zoom parameter corresponding to each frame of the fourth video image, a third correspondence between the face size in the video data captured by the first camera and the Zoom parameter of the second camera is determined.
[0035] As described above, by reasonably determining the first, second, and third correspondences, it is possible to determine the shooting control parameters through the first, second, and third correspondences, and when the second camera uses the shooting control parameters, it is possible to capture a magnified close-up image of the speaker located in the center area.
[0036] In one embodiment of this application, identifying the first pixel coordinates and first face size of the speaker within the currently processed first video frame in the first video data specifically includes:
[0037] Identify the speaker within the currently processed first video frame in the first video data;
[0038] Obtain a second sub-video frame containing the speaker's face within the first video frame;
[0039] The pixel coordinates of the center point of the second sub-video frame within the first video frame are determined and used as the first pixel coordinates. The size of the second sub-video frame is determined and used as the first face size.
[0040] In one embodiment of this application, identifying the speaker within the currently processed first video frame in the first video data specifically includes:
[0041] perform face recognition on the first video picture currently processed in the first video data and sixth video pictures that are continuous multiple frames before the first video picture in the first video data, and obtain a lip picture on each recognized face;
[0042] determine a speaker in the first video picture according to a motion amplitude of the lip in the lip picture, the motion amplitude of the lip of the speaker being the largest.
[0043] According to the above, the speaker in the video picture can be automatically determined based on face recognition and lip motion amplitude recognition.
[0044] In an embodiment of the present application, the intelligent device further comprises a microphone array.
[0045] The face recognition on the first video picture currently processed in the first video data and the sixth video pictures that are continuous multiple frames before the first video picture in the first video data, and the obtaining of the lip picture on each recognized face, specifically comprises:
[0046] obtaining DOA information of the speaker based on the microphone array;
[0047] obtaining a third sub-video picture in the first video picture currently processed in the first video data according to the DOA information and a fourth sub-video picture in the sixth video pictures that are continuous multiple frames before the first video picture in the first video data;
[0048] performing face recognition on the third sub-video picture and the fourth sub-video picture, and obtaining a lip picture on each recognized face.
[0049] According to the above, the direction of the speaker can be determined based on the DOA information of the microphone array, and then the sub-video picture corresponding to the direction where the speaker is located in the video picture can be obtained, and the speaker can be recognized in the sub-video picture, so that the data processing amount of the video picture can be reduced while ensuring the accuracy of the speaker recognition.
[0050] In an embodiment of the present application, after the obtaining of the second video data captured by the second camera, the method further comprises:
[0051] determining a third pixel coordinate and a fourth face size of the speaker in a seventh video picture currently processed in the second video data;
[0052] when the third pixel coordinate and the fourth face size satisfy a preset shooting condition, performing an operation of obtaining the second video data by the virtual camera driver program, the preset shooting condition being that the third pixel coordinate is located in a set pixel coordinate region and the fourth face size satisfies a preset face size.
[0053] The above, only when the pixel coordinates and the face size of the speaker in the video picture captured by the second camera satisfy the preset shooting condition, it is determined that the second camera captures the enlarged close-up picture of the speaker located in the middle area, at this time, the video data captured by the second camera is sent to the virtual camera driver program, so that the target application obtains the enlarged close-up picture of the speaker located in the middle area.
[0054] In an embodiment of the present application, the shooting control method further comprises:
[0055] When the third pixel coordinates and the fourth face size do not satisfy the preset shooting condition, the first shooting control parameter is adjusted according to the third pixel coordinates and the fourth face size to obtain a second shooting control parameter, and the second camera realizes the aligned shooting of the speaker under the second shooting control parameter and the video picture captured by the second camera satisfies the preset shooting condition;
[0056] Obtaining third video data captured by the second camera under the use of the second shooting control parameter;
[0057] The third video data is obtained by the virtual camera driver program and sent to the target application program, so that the target application program can use the third video data on the display screen.
[0058] The above, when the pixel coordinates and the face size of the speaker in the video picture captured by the second camera do not satisfy the preset shooting condition, the shooting control parameter used by the second camera is fine-tuned to ensure that the second camera can capture the enlarged close-up picture of the speaker located in the middle area.
[0059] In an embodiment of the present application, the obtaining of the second video data captured by the second camera further comprises:
[0060] Updating the first video picture currently processed;
[0061] Identifying the fourth pixel coordinates of the speaker in the updated first video picture;
[0062] When the coordinate distance between the fourth pixel coordinates and the first pixel coordinates is not greater than the distance threshold, it is determined that the speaker does not change and the operation of obtaining the second video data captured by the second camera is returned to be executed, so as to realize the continuous shooting of the speaker.
[0063] The fourth pixel coordinate of the speaker and the first pixel coordinate of the speaker are combined to determine whether the speaker has changed, and when the speaker has not changed, the second camera continues to use the current shooting control parameter, and the virtual camera driver uses the video data shot by the second camera to ensure that the enlarged close-up picture of the speaker is continuously shot.
[0064] In an embodiment of the present application, the shooting control method further comprises:
[0065] When the coordinate distance between the fourth pixel coordinate and the first pixel coordinate is greater than the distance threshold, the first video picture currently processed is continuously updated, the fourth pixel coordinate of the speaker in the updated first video picture is continuously identified, and coordinate counting is started, which is incremented by 1 each time the fourth pixel coordinate is determined;
[0066] If the current coordinate count reaches the count threshold, and the coordinate distance between the plurality of fourth pixel coordinates determined in the current coordinate count and the first pixel coordinate is greater than the distance threshold, it is determined that the speaker has changed, and the operation of identifying the first pixel coordinate of the speaker and the first face size in the first video picture currently processed in the first video data is returned to be executed, so that the second camera is aligned to shoot the changed speaker.
[0067] The fourth pixel coordinate of the speaker and the first pixel coordinate of the speaker are combined to determine whether the speaker has changed, and when the speaker has not changed, the second camera continues to use the current shooting control parameter, and the virtual camera driver uses the video data shot by the second camera to ensure that the enlarged close-up picture of the speaker is continuously shot.
[0068] In an embodiment of the present application, the shooting control method further comprises:
[0069] If the current coordinate count does not reach the count threshold and the coordinate distance between the newly determined fourth pixel coordinate and the first pixel coordinate is not greater than the distance threshold, it is determined that the speaker has not changed, and the operation of obtaining the second video data shot by the second camera is returned to be executed to realize continuous shooting of the speaker.
[0070] The fourth pixel coordinate of the speaker and the first pixel coordinate of the speaker are combined to determine whether the speaker has changed, and when the speaker has not changed, the second camera continues to use the current shooting control parameter, and the virtual camera driver uses the video data shot by the second camera to ensure that the enlarged close-up picture of the speaker is continuously shot.
[0071] In a second aspect, an embodiment of the present application further provides a photographing control apparatus applied to a smart device, the smart device comprising a first camera, a display screen, a virtual camera driver, and a target application program, the smart device being connected with an external second camera, the second camera being a rotatable camera, and the photographing control apparatus comprising:
[0072] a first video acquisition unit configured to acquire first video data captured by the first camera, the first camera being configured to capture a first photographing space;
[0073] a first data identification unit configured to identify a first pixel coordinate of a speaker in a first video frame currently processed in the first video data and a first face size, the speaker being located in the first photographing space;
[0074] a parameter determination unit configured to determine a first photographing control parameter of the second camera according to the first pixel coordinate and the first face size;
[0075] a parameter use unit configured to control the second camera to use the first photographing control parameter for photographing, so that the second camera realizes aligned photographing of the speaker under the first photographing control parameter;
[0076] a second video acquisition unit configured to acquire second video data captured by the second camera;
[0077] a first video sending unit configured to acquire the second video data by the virtual camera driver and send the second video data to the target application program, so that the target application program can use the second video data on the display screen.
[0078] In a third aspect, an embodiment of the present application further provides a smart device, comprising a first camera, a display screen, one or more processors, and a memory; the smart device being connected with an external second camera, the second camera being a rotatable camera;
[0079] the first camera being configured to perform photographing according to an indication of the processor;
[0080] the display screen being configured to perform display according to an indication of the processor;
[0081] the memory being configured to store one or more programs;
[0082] when the one or more programs are executed by the one or more processors, the one or more processors implement the photographing control method according to the first aspect;
[0083] The second camera is configured to capture an image using the shooting control parameter determined by the processor.
[0084] In a fourth aspect, an embodiment of the present application provides a shooting control system, which comprises the smart device and the second camera according to the third aspect.
[0085] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, the shooting control method according to the first aspect is implemented.
[0086] The shooting control device, the smart device, the system and the storage medium provided above have the advantages of the shooting control method. BRIEF DESCRIPTION OF DRAWINGS
[0087] FIG. 1 is a structural schematic diagram of a smart device according to an embodiment of the present application;
[0088] FIG. 2 is a front view of a smart device according to an embodiment of the present application;
[0089] FIG. 3 is a schematic diagram of a shooting control system according to an embodiment of the present application;
[0090] FIG. 4 is a flowchart of a shooting control method according to an embodiment of the present application;
[0091] FIG. 5 is a flowchart of a shooting control method according to another embodiment of the present application;
[0092] FIG. 6 is a horizontal field of view diagram of a microphone array according to an embodiment of the present application;
[0093] FIG. 7 is a schematic diagram of a third sub-video according to an embodiment of the present application;
[0094] FIG. 8 is a structural schematic diagram of a shooting control device according to an embodiment of the present application;
[0095] FIG. 9 is a structural schematic diagram of a shooting control system according to an embodiment of the present application. DETAILED DESCRIPTION
[0096] The present application will be further described below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are intended to explain the present application, but not to limit the present application. In addition, it should be noted that only the parts related to the present application are shown in the drawings for the convenience of description.
[0097] It should be noted that, due to the limitation of the length of the specification, all optional embodiments are not enumerated in the specification, and those skilled in the art should be able to think of any combination of technical features as long as the technical features do not contradict each other, which can constitute an optional embodiment.
[0098] The embodiments will be described in detail below.
[0099] In the related art, the number of cameras configured on the intelligent device is usually one, and the camera can realize shooting. Taking a classroom teaching scene as an example, the intelligent device is usually configured at the front end of the classroom, at this time, the camera configured on the intelligent device shoots from front to back in the classroom. Taking a meeting scene as an example, the intelligent device can shoot the local participants. In various application scenarios of the camera, shooting the speaker in the shooting space is a relatively important application scenario.
[0100] In the scenario of shooting the speaker, it is a relatively important requirement to make the speaker clearly appear in the middle area of the video picture. However, due to technical and cost considerations, the imaging quality of the camera configured on the intelligent device is generally not high, and the shooting angle of the camera is fixed and cannot realize rotating shooting. At this time, if the speaker is located at the edge of the shooting space, the speaker in the picture shot by the camera is also located at the edge, and the picture of the speaker in the middle area cannot be obtained. Moreover, when the speaker in the shot picture is zoomed in for close-up, the clarity of the zoomed-in picture is not high, so that the picture presents a blurred effect, especially when the picture is transmitted through the network, due to data compression and data loss in the transmission process, the clarity of the picture received by the receiving end will be further reduced, which seriously affects the video communication experience, especially in the remote meeting and remote teaching scenarios, which may also affect the meeting effect and teaching effect.
[0101] Based on this, the inventors of the present application design a new shooting control method, which designs two cameras, one is a first camera built in the intelligent device, and the other is a second camera rotatable independently of the intelligent device. By developing an algorithm for the two cameras to cooperate with each other, the position and size of the speaker in the video picture shot by the intelligent device are identified after the first camera shoots the shooting space (the position and size can reflect the position of the speaker in the shooting space and the distance between the speaker and the first camera), and the shooting control parameters used by the second camera are adjusted in combination with the position and size of the speaker, so that the second camera can shoot the speaker in alignment, and then obtain a video picture in which the speaker is located in the middle area and is clear, to solve the technical problems existing in the prior art.
[0102] The photographing control method provided in the embodiments of the present application can be executed by a smart device. The smart device can be implemented by software and / or hardware, and can be composed of two or more physical entities or one physical entity. At present, the smart device can be a smart blackboard used in a teaching scenario, a conference tablet or an interactive tablet used in a conference scenario, a notebook computer, or other electronic devices.
[0103] FIG. 1 is a structural schematic diagram of a smart device according to an embodiment of the present application. As shown in FIG. 1, the smart device at least includes a processor 11, a memory 12, a first camera 13, a microphone array 14, and a display screen 15. The processor 11, the memory 12, the first camera 13, the microphone array 14, and the display screen 15 can be connected by a bus or other means.
[0104] In an embodiment, the processor 11 can be one or more, and one processor 11 is taken as an example in FIG. 1. The processor 11 can include an application processor (AP), a graphics processing unit (GPU), a central processing unit (CPU), and other processing units.
[0105] The GPU and the memory chip, the interface circuit, and other required components (all of which are configured in the smart device) can constitute a graphics card. The graphics card can convert information required to be displayed by the smart device to drive the corresponding display screen to display. At present, the graphics card has data processing and logical analysis capabilities, that is, the graphics card can realize data processing and logical analysis functions by running a corresponding computer executable program (also referred to as a computer program).
[0106] The memory 12, as a computer readable storage medium, can be used to store software programs, computer executable programs and modules, such as program instructions / modules corresponding to the photographing control method in the embodiments of the present application. The memory 12 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required by at least one function; and the data storage area can store data created according to the use of the smart device and the like. In addition, the memory 12 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state memory device. In some examples, the memory 12 can further include a memory remotely arranged with respect to the processor 11, and the remote memory can be connected to the smart device through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. In the embodiments, the processor 11 runs one or more programs in the memory 12, and can implement the photographing control method provided by the embodiments of the present application.
[0107] The smart device is also configured with a camera, and the camera configured in the smart device is currently referred to as a first camera 13. Generally, the first camera 13 is a front camera, and the number of the first camera 13 is one. Optionally, the first camera 13 is fixedly installed at the back of the smart device and cannot be rotated, that is, the smart device is built-in with a non-rotatable first camera 13. The first camera 13 can be a wide-angle fixed-focus camera, and can also be other types of cameras, which are not limited in the embodiments. The first camera 13 is generally installed at the upper edge of the side of the smart device facing the user, for example, FIG. 2 is a front view of a smart device provided by an embodiment of the present application, and referring to FIG. 2, the first camera 13 is installed at the upper middle position of the surface of the smart device 10 facing the user. In actual application, the first camera 13 can also be installed at other positions, which are not limited in the embodiments.
[0108] The smart device is also configured with a microphone array 14, and the microphone array 14 is composed of a plurality of microphones. The installation position of the microphone array 14 is not limited at present, for example, referring to FIG. 2, the microphone array 14 is composed of 8 microphones 141, and every 4 microphones 141 in the 8 microphones 141 form a group and are distributed on both sides of the first camera 13.
[0109] Optionally, the smart device can further include a display screen 15 which displays based on the indication of the processor. The display screen 15 can be a liquid crystal display (LCD), an LED display, an organic light-emitting diode (OLED) display, a flex light-emitting diode (FLED) display, or the like. In an embodiment, the display screen 15 can further integrate a touch function, in which case the display screen includes a display panel and a touch panel. The display panel is used to complete visual output. The touch panel can be a touch component supporting infrared touch, electromagnetic touch, capacitive touch, and / or resistive touch, etc.
[0110] The smart device can further include one or more communication interfaces to realize communication with other devices through the communication interfaces. For example, the smart device realizes communication with the second camera through the communication interfaces.
[0111] In addition, the smart device can further include a power supply, a speaker, a physical button, and the like, which are not limited by embodiments.
[0112] On the basis of the hardware structure described above, the smart device supports at least one type of operating system, which can be an Android system, a Windows system, or a Linux system, etc. The smart device can install at least one application program under the operating system. The installed application program can be an application program provided by the operating system, an application program downloaded from a background server or a third-party device. The smart device can realize corresponding functions by running each application program. In an embodiment, the smart device at least installs an application program for realizing the photographing control method.
[0113] In an embodiment, the smart device is further connected with a camera independent of the smart device, which is referred to as a second camera at present. The second camera is a rotatable camera, and in an embodiment, a pan-tilt-zoom (PTZ) camera is described as the second camera. The PTZ camera can be understood as a camera with a pan-tilt head. The pan-tilt head can carry the camera to rotate in horizontal and vertical directions, so that the camera can take pictures from multiple angles. The PTZ camera used at present is a PTZ camera, which can realize the functions of omnidirectional movement (up and down, left and right), zooming (variable magnification and variable focus). When the PTZ camera is used for shooting, at least Pan parameter, Tilt parameter and Zoom parameter are required. The Pan parameter is a parameter for determining the position of the PTZ camera in the horizontal direction (which can be represented by an angle), that is, through the Pan parameter, the position to which the PTZ camera should be rotated in the horizontal direction can be determined, and different Pan parameters correspond to different horizontal rotation positions. The Tilt parameter is a parameter for determining the position of the PTZ camera in the vertical direction (which can be represented by an angle), that is, through the Tilt parameter, the position to which the PTZ camera should be rotated in the vertical direction can be determined, and different Tilt parameters correspond to different vertical rotation positions. The Zoom parameter is the magnification and focal length of the PTZ camera, that is, by adjusting the Zoom parameter, the PTZ camera can complete variable magnification and variable focus. By adjusting the Pan parameter and the Tilt parameter, the shooting angle of the PTZ camera can be changed, and at present, by setting appropriate Pan parameter and Tilt parameter for the PTZ camera, the object to be shot (such as the speaker) can be located in the middle area of the video picture shot by the PTZ camera. By adjusting the Zoom parameter of the PTZ camera, the size of the object to be shot in the video picture can be changed, and at present, by setting appropriate Zoom parameter for the PTZ camera, the object to be shot (such as the speaker) can occupy a suitable size in the video picture, and the clarity of the object to be shot in the video picture can also be ensured, that is, the video picture is a high-definition picture (because the PTZ camera is a variable focus camera, so after variable magnification, a high-definition picture can still be shot).
[0114] The second camera can be fixed on the smart device, such as being fixed on the upper edge, lower edge or side edge of the smart device by a fixing device. The second camera can also be fixed on the same support as the smart device, such as being installed on the support for installing the smart device. The second camera can also be installed on a wall, which can be a wall behind the smart device or a wall on which the smart device is fixed. Optionally, after the second camera is installed and fixed, the second camera has the same orientation as the first camera in the smart device, where the same orientation can be understood as the second camera and the first camera facing the same shooting area, such as when the first camera can shoot the front face of the speaker, the second camera can also shoot the front face of the speaker. FIG. 3 is a schematic diagram of a shooting control system provided by an embodiment of the present application, where the shooting control system includes a smart device 10 and a second camera 20. After the second camera 20 is installed, it is fixed at a middle position above the smart device 10, and the second camera 20 and the first camera of the smart device 10 both face the side of the user to ensure that the orientations of the two cameras are the same.
[0115] The second camera and the smart device can be connected in a wired (such as USB) or wireless (such as Wifi) manner to realize the transmission of data and instructions between the second camera and the smart device. Currently, the shooting control parameters used by the second camera to shoot the speaker are determined by the smart device, the second camera uses the shooting control parameters to shoot, and the video data shot by the second camera is transmitted back to the smart device for use.
[0116] It should be noted that in the related art, when using a PTZ camera, the horizontal rotation position, the vertical rotation position and the zoom of the PTZ camera need to be adjusted by manual operation. For example, when shooting a speaker in close-up by a PTZ camera, the user needs to first adjust the Pan parameter and the Tilt parameter of the PTZ camera to align the face of the speaker, and then adjust the Zoom parameter of the PTZ camera to make the size of the face of the speaker shot by the PTZ camera appropriate. However, manual adjustment requires a certain amount of time, which makes the speed of aligning the face of the PTZ camera very slow, and manual adjustment also has the problem of low control precision, which affects the user experience of using the PTZ camera and the shooting effect of the PTZ camera. In the embodiments, after the second camera and the smart device cooperate with each other, the shooting control parameters of the second camera are determined by the smart device, and the user does not need to manually adjust the second camera, which can improve the alignment speed of the second camera and ensure the shooting quality of the second camera.
[0117] After the smart device is connected with the second camera, the smart device can acquire the video data captured by the first camera and the video data captured by the second camera. On this basis, the virtual camera driver and the system application service are installed in the smart device, the virtual camera driver and the system application service are obtained by encapsulating computer programs that can run in the smart device, and the virtual camera driver and the system application service cooperate to enable the application program (located in the application layer) installed in the smart device to realize the non-sensing operation of the underlying camera logic, that is, the application program does not need to care about the number and control logic of the underlying camera, and only needs to acquire and use the video data sent by the virtual camera driver.
[0118] The virtual camera driver is used to expose to the application program in the upper layer, and the application program in the upper layer can use the video data sent by the virtual camera driver, that is, the virtual camera driver can be regarded as a bridge between the underlying physical camera and the application program in the upper layer. The system application service is used to read the video data (i.e., video stream) captured by the first camera and the second camera, determine the shooting control parameter to be used by the second camera in combination with the video data captured by the first camera, make the second camera use the determined shooting control parameter, and push the video data to the virtual camera driver. The system application service can be composed of a local AI algorithm module and a central logic control module, the local AI algorithm module can run in the CPU or the graphics card, and the central logic control module can run in the CPU.
[0119] In addition, the smart device also installs a target application program, which can be understood as an upper-layer application program that uses the video data captured by the virtual camera driver. That is, the target application program is an application program installed in the smart device and currently calling the virtual camera driver. For example, the target application program can be an application program for realizing video conference or an application program for realizing video teaching.
[0120] At present, the system application service and the virtual camera cooperate to realize the shooting control method.
[0121] FIG. 4 is a flowchart of a shooting control method provided by an embodiment of the present application, referring to FIG. 4, the shooting control method includes steps 310-360:
[0122] Step 310: acquiring first video data captured by a first camera, the first camera being used for capturing a first shooting space.
[0123] The first shooting space can be considered as a space region shot by the first camera in an application scenario of the smart device. For example, the smart device is applied in a conference scenario, and the first shooting space is a space where a conference room is located. The smart device is applied in a teaching scenario, and the first shooting space is a space where a classroom is located. Currently, one or more shot objects exist in the first shooting space. In the embodiments, the shot object is a person.
[0124] The first camera can obtain video data (i.e., a video stream) when shooting. Currently, the video data obtained by the first camera when shooting the first shooting space is referred to as first video data. The first video data can also be considered as panoramic data obtained after shooting the first shooting space. Optionally, the intrinsic parameter currently used by the first camera is a fixed intrinsic parameter, which can be pre-set and stored in the smart device.
[0125] Currently, the first video data is mainly obtained by the local AI algorithm module.
[0126] In step 320, the first pixel coordinates of a speaker in a first video picture currently processed in the first video data and the first face size of the speaker are identified, and the speaker is located in the first shooting space.
[0127] Illustratively, when the first video data is processed, the video picture in the first video data is mainly processed. One video picture can also be understood as one frame of image in the video data. It can be understood that the video data obtained by shooting is composed of multiple frames of images. After the multiple frames of images change continuously at a certain speed, according to the principle of visual persistence, the human eye can see a continuous and smooth visual effect, that is, the video is seen. Currently, when the smart device processes the video picture, the pixel coordinates of the speaker in the video picture and the face size of the speaker are mainly determined. In one embodiment, one frame of video picture currently processed by the smart device is referred to as a first video picture. Currently, the number of speakers is one, that is, the smart device can only determine the pixel coordinates of one speaker and the face size of the speaker at the same time. The speaker generally faces the smart device or is lateral to the smart device, so that the first camera can shoot the face of the speaker.
[0128] Illustratively, the smart device first determines the speaker in the first video picture. The way in which the smart device determines the speaker is not limited at present, for example, the smart device identifies each face in several consecutive frames of video pictures (including the first video picture) in the first video data, then determines the face with the largest lip movement amplitude among the faces, and determines the face as the face of the speaker. For another example, the smart device obtains the approximate direction of the speaker by using a microphone array, and then identifies the face of the speaker in combination with the direction of the speaker and each face in several consecutive frames of video pictures (including the first video picture) in the first video data.
[0129] After the speaker is identified, the pixel coordinates of the speaker in the first video frame and the face size of the speaker in the first video frame are obtained based on the face image of the speaker in the first video frame. The first camera has its own image pixel coordinate system during use. The image pixel coordinate system can be understood as a two-dimensional coordinate system established by a two-dimensional image plane. Each frame of image in the first video data uses the image pixel coordinate system. At this time, each pixel point in the image has a corresponding coordinate in the image pixel coordinate system. At present, the coordinate in the image pixel coordinate system is referred to as a pixel coordinate, and the pixel coordinates of the speaker in the first video frame are referred to as first pixel coordinates. The first pixel coordinates can be determined in the following manner: the face of the speaker in the first video frame is identified, and then the pixel coordinates of the center point (or other set point) of the face of the speaker in the first video frame in the image pixel coordinate system (i.e., the coordinates of the pixel point where the center point is located) are taken as the first pixel coordinates. The face size refers to the size of the face image in the video frame. At present, the face size of the speaker in the first video frame is referred to as a first face size. Optionally, the total number of pixel points occupied by the face image of the speaker in the first video frame is taken as the first face size. Further optionally, a minimum rectangular region containing the face image of the speaker in the first video frame is determined, and the width (number of pixel points), height (number of pixel points), or number of pixels occupied by the diagonal of the minimum rectangular region is taken as the first face size, or the width, height, or length of the diagonal of the face image is directly taken as the first face size. It should be noted that the face image of the speaker mentioned at present can be a frame containing only the face of the speaker, a frame containing the head where the face is located, or a frame containing the head and the neck.
[0130] It can be understood that the first face size and the first pixel coordinates are mainly determined by the local AI algorithm module at present.
[0131] Step 330: determining the first shooting control parameter of the second camera according to the first pixel coordinates and the first face size.
[0132] For example, after the first pixel coordinates and the first face size are determined, the shooting control parameter required by the second camera to shoot the speaker can be determined based on the first pixel coordinates and the first face size. The shooting control parameter required to shoot the speaker can include a Pan parameter and a Tilt parameter, or a Pan parameter, a Tilt parameter, and a Zoom parameter. In one embodiment, the shooting control parameter includes a Pan parameter, a Tilt parameter, and a Zoom parameter.
[0133] Currently, the shooting control parameter obtained based on the first pixel coordinate and the first face size is denoted as a first shooting control parameter. The Pan parameter, the Tilt parameter and the Zoom parameter included in the first shooting control parameter are denoted as a first Pan parameter, a first Tilt parameter and a first Zoom parameter respectively.
[0134] It can be understood that the rotation of the second camera is mainly related to the Pan parameter and the Tilt parameter, and the shooting angle of the second camera is different when the second camera rotates to different positions. When the second camera shoots the speaker, the speaker needs to be located as much as possible in the middle region of the video picture shot by the second camera, at this time, the shooting angle required by the second camera when shooting the speaker is mainly determined by the position of the speaker, that is, when the speaker is located at different positions, the second camera needs to adjust the Pan parameter and the Tilt parameter to make the second camera shoot the video picture in which the speaker is located in the middle region after adjustment. Generally speaking, since the shooting angle of the first camera is fixed, when the speaker is at different positions, the first pixel coordinate of the speaker in the first video picture will be different, therefore, in the embodiment, the first Pan parameter and the first Tilt parameter required for shooting the speaker can be determined in combination with the first pixel coordinate of the speaker. Optionally, the correspondence relationship between each pixel coordinate in the video data shot by the first camera and the Pan parameter and the correspondence relationship between each pixel coordinate and the Tilt parameter can be set in advance according to the actual situation, currently, the correspondence relationship between each pixel coordinate in the video data shot by the first camera and the Pan parameter is denoted as a first correspondence relationship, and the correspondence relationship between each pixel coordinate in the video data shot by the first camera and the Tilt parameter is denoted as a second correspondence relationship. It can be understood that after the Pan parameter and the Tilt parameter corresponding to the pixel coordinate are determined based on the first correspondence relationship and the second correspondence relationship, the middle region of the second camera when shooting under the parameters contains the physical entity corresponding to the pixel coordinate in the shooting space, that is, the second camera can shoot the physical entity corresponding to the pixel coordinate. Optionally, the pixel coordinate in the first correspondence relationship is taken as the independent variable and the Pan parameter is taken as the dependent variable, after the first pixel coordinate is brought into the first correspondence relationship, the obtained Pan parameter can be taken as the first Pan parameter. The pixel coordinate in the second correspondence relationship is taken as the independent variable and the Tilt parameter is taken as the dependent variable, after the first pixel coordinate is brought into the second correspondence relationship, the obtained Tilt parameter can be taken as the first Tilt parameter.
[0135] The zoom of the second camera is mainly related to the Zoom parameter. At present, when the second camera captures the speaker, the size of the face of the speaker in the video picture captured by the second camera needs to meet a certain size, so as to ensure that the speaking picture of the speaker is as clear as possible, that is, to realize the close-up of the speaker, therefore, the second camera needs to use appropriate Zoom parameter. It can be understood that when the distance between the speaker and the second camera is different, the Zoom parameter required for capturing the speaker will be different. Moreover, since the position of the first camera is fixed, when the distance between the speaker and the first camera is different, the size of the face of the speaker in the first video picture will be different, and the distance between the speaker and the first camera is inversely proportional to the size of the face of the speaker in the first video picture. That is, the first face size of the speaker can reflect the distance between the speaker and the first camera. Moreover, since the position of the second camera is fixed and the relative position relationship between the first camera and the second camera is also fixed, the distance between the speaker and the second camera can also be reflected based on the distance between the speaker and the first camera, based on which, in the embodiment, the first Zoom parameter required for capturing the speaker can be determined in combination with the first face size of the speaker. Optionally, the corresponding relationship between different face sizes in the video picture captured by the first camera and the Zoom parameter can be set in advance in combination with the actual situation, at present, the corresponding relationship is recorded as a third corresponding relationship. It can be understood that after the Zoom parameter corresponding to the face size is determined based on the third corresponding relationship, the face size obtained by the second camera when capturing under the Zoom parameter meets a certain size. Optionally, the face size in the third corresponding relationship is the independent variable, and the Zoom parameter is the dependent variable. After the first face size is brought into the third corresponding relationship, the obtained Zoom parameter can be used as the first Zoom parameter.
[0136] It can be understood that when the shooting control parameter includes the Pan parameter and the Tilt parameter, the first Pan parameter can be determined using the first corresponding relationship and the first Tilt parameter can be determined using the second corresponding relationship.
[0137] Optionally, the first shooting control parameter can be determined by the local AI algorithm module.
[0138] Step 340, control the second camera to capture using the first shooting control parameter, so that the second camera realizes the aligned shooting of the speaker under the first shooting control parameter.
[0139] For example, after determining the first shooting control parameter, the smart device instructs the second camera to use the first shooting control parameter. Alternatively, the smart device sends the first shooting control parameter to the second camera, and the second camera rotates left and right, up and down, and zooms based on the first shooting control parameter, so as to use the first shooting control parameter. Alternatively, the smart device determines the operation parameters of the motor for controlling the rotation left and right and the motor for controlling the rotation up and down in the holder of the second camera based on the first Pan parameter and the first Tilt parameter in the first shooting control parameter, and sends the operation parameters and the first Zoom parameter to the second camera. Then, the second camera drives the motor to operate according to the operation parameters and adjusts the zooming according to the first Zoom parameter, so as to use the first shooting control parameter.
[0140] It can be understood that after the second camera uses the first shooting control parameter to shoot, the speaker is basically located in the middle area in the video picture shot by the second camera, and the size of the face of the speaker can reach a certain size, so that the speaking picture of the speaker in the video picture is as clear as possible, that is, the aligned shooting of the speaker is realized.
[0141] Alternatively, the central logic control module can control the second camera to use the first shooting control parameter.
[0142] Step 350, obtaining second video data shot by the second camera.
[0143] Currently, the video data shot by the second camera is referred to as second video data. Alternatively, the smart device obtains the second video data shot by the second camera when it is determined that the second video data is needed. Currently, the smart device determines that the second video data shot by the second camera is needed after controlling the second camera to use the first shooting control parameter.
[0144] Step 360, obtaining the second video data by the virtual camera driver program, and sending the second video data to the target application program, so that the target application program can use the second video data on the display screen.
[0145] For example, after the second camera uses the first head shooting control parameter, the intelligent device obtains second video data shot by the second camera, and the virtual camera driver obtains the second video data. At this time, the virtual camera driver can receive one-way video data (currently the second video data) and send the received second video data to the target application. Wherein, after the target application obtains the second video data, the target application can use the second video data on the display screen. For example, when the target application is an application for realizing video conference (such as Meilipai conference or Tencent conference), after the target application obtains the second video data, the target application displays the second video data on the display screen and sends the second video data to other terminal devices participating in the conference for display.
[0146] It can be understood that in actual application, the intelligent device can select the video data shot by the first camera or the video data shot by the second camera for the virtual camera driver to obtain, so as to send to the target application. For example, when there is no speaker, the intelligent device sends the video data shot by the first camera to the virtual camera driver, and when there is a speaker, the intelligent device controls the second camera to shoot the speaker and then sends the video data shot by the second camera to the virtual camera driver. The intelligent device can also send the video data shot by the first camera and the video data shot by the second camera to the virtual camera driver, and the virtual camera driver processes the two-way video data, such as merging the two-way video data into one-way video data according to a set rule (such as embedding one-way video data into another way video data) and sending to the target application.
[0147] The first video data captured by the first camera is acquired, and the first pixel coordinates of the speaker in the first video frame in the first video data and the first face size are identified. Then, the first shooting control parameter of the second camera is determined according to the first pixel coordinates and the first face size. After the second camera uses the first shooting control parameter, the second video data captured by the second camera is acquired. Then, the second video data is acquired by the virtual camera driver program and sent to the target application program for the target application program to use the technical means of the second video data on the display screen. The technical problem that the camera configured on the intelligent device in the related art cannot capture the video frame of the speaker in the middle area when the speaker is located at the edge of the shooting space due to the fixed shooting angle of the camera is solved. By using the rotatable second camera and the intelligent device in cooperation, the panoramic frame in the first shooting space is captured by the first camera configured by the intelligent device, and the close-up frame of the speaker in the middle area is captured by the second camera. Then, the video data (second video data) captured by the second camera is acquired by the virtual camera driver program and sent to the target application program, so that the target application program can obtain a clear speaker speaking frame, and the clarity of the video frame after network transmission is improved when the video frame needs to be transmitted. In addition, the intelligent device can automatically determine the shooting control parameter of the second camera, without manually adjusting the second camera, so as to realize automatic intelligent adjustment, improve the alignment speed of the second camera, and ensure the shooting quality of the second camera. In addition, the hardware design of the intelligent device does not need to be changed, and only the software processing logic needs to be modified, so that the shooting control method is easy to implement. In addition, by designing the virtual camera driver program, the target application program does not need to care about the number of underlying cameras and control logic, and the number of underlying cameras can be easily expanded.
[0148] FIG. 5 is a flowchart of a shooting control method provided by another embodiment of the present application, which is a detailed shooting control method based on the shooting control method shown in FIG. 4. Referring to FIG. 5, the shooting control method includes steps 410-4120:
[0149] Step 410, establish a first correspondence relationship between the pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera, and establish a second correspondence relationship between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera.
[0150] Before the cooperation of the first camera and the second camera is used to realize the shooting of the speaker, the first correspondence relationship and the second correspondence relationship need to be established. Currently, after the relative position relationship between the smart device and the second camera is fixed (at this time, the first camera and the second camera are oriented in the same direction), the first correspondence relationship and the second correspondence relationship can be established. Optionally, after the smart device and the second camera are installed in a physical space (such as a conference room or a classroom) (that is, the relative position relationship no longer changes), the first correspondence relationship and the second correspondence relationship are established.
[0151] In one embodiment, when the first correspondence relationship and the second correspondence relationship are established, any one of the following two schemes can be used:
[0152] Scheme one, step 410 can specifically include steps 411-414:
[0153] Step 411, a second video frame based on the second video frame shot by the first camera is obtained, and the first camera shoots a static second shooting space.
[0154] Since the first correspondence relationship and the second correspondence relationship are both related to the pixel coordinates in the video data shot by the first camera, the first correspondence relationship and the second correspondence relationship can be established together. When the first correspondence relationship and the second correspondence relationship are established, the video data shot by the first camera in the second shooting space is first obtained. Currently, the second shooting space and the first shooting space can be the same space. In actual application, the second shooting space and the first shooting space can be different shooting spaces. For example, when the second camera and the smart device are both installed on a support, and the support can be moved, the user can push the support to different shooting spaces, but the relative position relationship between the smart device and the second camera does not change. Optionally, when the first camera shoots in the second shooting space, the second shooting space is a relatively static space, that is, all objects (including people and things) in the second shooting space are in a static state during the shooting of the first camera, so as to ensure the accuracy of the first correspondence relationship and the second correspondence relationship.
[0155] It can be understood that since the second shooting space is a relatively static space, the picture content of each video frame in the video data shot by the first camera in the second shooting space should be basically consistent. When the video data shot by the first camera is obtained, one video frame in the video data can be obtained. Currently, the obtained video frame is recorded as a second video frame. It can be understood that the second video frame can be any video frame in the video data shot by the first camera. Optionally, after the second video frame is obtained, the smart device can control the first camera to stop shooting. Alternatively, the smart device controls the first camera to shoot the second shooting space to obtain a picture, and the picture is used as the second video frame.
[0156] Step 412, the second video screen is divided into a plurality of first sub-video screens, and the center point of each first sub-video screen in the second video screen is determined as a second pixel coordinate.
[0157] For example, the second video screen is divided into a plurality of sub-video screens, and each sub-video screen is recorded as a first sub-video screen, wherein the number of first sub-video screens can be determined in combination with the resolution of the second video screen. It can be understood that the more the number of first sub-video screens, the more accurate the first corresponding relationship is, but the calculation amount is also larger. Therefore, the appropriate number of first sub-video screens can be set according to the actual situation, for example, when the resolution of the second video screen is 3840x2160, the second video screen is divided into 9x9 first sub-video screens. At this time, the height of each first sub-video screen is equal, but the width is not completely equal (because 3840 cannot be divided by 9, so the division method can be that the width of the first sub-video screen in the first eight columns is equal, and the width of the first sub-video screen in the last column is not equal to that of the first eight columns, so that the width of each first sub-video screen in the same column is equal to 3840). In practical application, the size of each first sub-video screen can also be completely equal.
[0158] After obtaining each first sub-video screen, the pixel coordinate of the center point of each first sub-video screen in the image pixel coordinate system of the first camera (i.e. the pixel coordinate of the center point in the second video screen) is determined. At present, the pixel coordinate of the center point is recorded as a second pixel coordinate, and each first sub-video screen has a corresponding second pixel coordinate.
[0159] Step 413, the second Pan parameter and the second Tilt parameter corresponding to each first sub-video screen are obtained, the similarity between the third video screen obtained by the second camera under the use of the second Pan parameter, the second Tilt parameter and the second Zoom parameter and the corresponding first sub-video screen reaches a similarity threshold, and the second Zoom parameter is a fixed parameter when the second camera shoots the second shooting space.
[0160] For example, after obtaining each first sub-video picture, the second camera is aligned to the shooting space corresponding to one of the first sub-video pictures (i.e., a sub-space in the second shooting space). The first sub-video picture can be the first sub-video picture in the upper left corner or other first sub-video picture. When the sizes of the first sub-video pictures are not completely consistent, the first sub-video picture with the largest size can be selected. The alignment process can be achieved by manually operating the second camera or by manually operating an intelligent device to control the Pan parameter and Tilt parameter of the second camera. Aligning the shooting space corresponding to the first sub-video picture can be understood as the center point of the first sub-video picture corresponding to the physical entity in the shooting space being located in a set pixel coordinate region (which can be set according to actual conditions, and is generally a region in the middle of the video picture) in the video picture captured by the second camera. At this time, it can be considered that the physical entity is located in the middle region of the video picture.
[0161] After the alignment of the second camera, the Zoom parameter of the second camera is adjusted so that the similarity between the video picture in the video data captured by the adjusted second camera (at present, the video picture captured by the second camera is recorded as a third video picture when the first correspondence is established) and the first sub-video picture used in the alignment reaches a similarity threshold value (which can be set according to actual conditions). Since the second shooting space is a static space, the intelligent device can also control the second camera to capture a picture, and the picture is used as the third video picture.
[0162] It can be understood that when the similarity between the first sub-video picture and the third video picture reaches the similarity threshold, it can be considered that the first sub-video picture and the third video picture are basically consistent, that is, the second camera can obtain a picture when the shooting space corresponding to the first sub-video picture is aligned, and the center points of the first sub-video picture and the third video picture are based on consistency. Optionally, the similarity can be calculated by the intelligent device. For example, the intelligent device controls the second camera to continuously adjust the Zoom parameter (the Zoom parameter can be automatically adjusted by the intelligent device according to a pre-set adjustment range, or the intelligent device can be manually controlled to adjust the Zoom parameter), obtains the third video picture shot by the second camera under each Zoom parameter, and calculates the similarity between the third video picture and the first sub-video picture each time a third video picture is obtained, and determines whether the similarity reaches the similarity threshold. If the similarity reaches the similarity threshold, the second camera is fixed to use the current Zoom parameter, that is, there is no need to adjust the Zoom parameter, and the recorded Pan parameter and Tilt parameter are recorded as the Pan parameter and Tilt parameter corresponding to the first sub-video picture. If the similarity does not reach the similarity threshold, the Zoom parameter is continuously adjusted, and the third video picture is obtained until the similarity between the third video picture and the first sub-video picture reaches the similarity threshold.
[0163] Then, the intelligent device controls the second camera to use the fixed Zoom parameter and continuously adjusts the Pan parameter and the Tilt parameter (the Pan parameter and the Tilt parameter can be automatically adjusted by the intelligent device according to a pre-set adjustment range, or the intelligent device can be manually controlled to adjust the Pan parameter and the Tilt parameter). After the second camera adjusts the Pan parameter and the Tilt parameter each time, the intelligent device obtains the third video picture based on the shooting of the second camera, and calculates the similarity between the third video picture and another first sub-video picture. If the similarity reaches the similarity threshold, it is considered that the other first sub-video picture and the third video picture are basically consistent, that is, the second camera can obtain a picture when the shooting space corresponding to the other first sub-video picture is aligned, and the center points of the other first sub-video picture and the third video picture are based on consistency. At this time, the recorded Pan parameter and Tilt parameter are recorded as the Pan parameter and Tilt parameter corresponding to the other first sub-video picture. Then, the intelligent device controls the second camera to adjust the Pan parameter and the Tilt parameter again until each first sub-video picture has corresponding Pan parameter and Tilt parameter.
[0164] In one embodiment, the Pan parameter and the Tilt parameter corresponding to each first sub-video picture are respectively denoted as a second Pan parameter and a second Tilt parameter, and the fixed Zoom parameter used by the second camera is denoted as a second Zoom parameter, that is, each first sub-video picture has a corresponding second Pan parameter, a second Tilt parameter, and a second Zoom parameter, and the second Zoom parameters corresponding to the first sub-video pictures are the same, that is, the second Zoom parameter is a fixed parameter. At present, the intelligent device can obtain the second Pan parameter and the second Tilt parameter corresponding to each first sub-video picture for use.
[0165] In step 414, a first correspondence relationship between pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera is determined according to the second pixel coordinates corresponding to each first sub-video picture and the second Pan parameter, and a second correspondence relationship between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera is determined according to the second pixel coordinates corresponding to each first sub-video picture and the second Tilt parameter.
[0166] Based on the second pixel coordinates (i.e., x, y) corresponding to each first sub-video picture and the second Pan parameter, a function fitting of a surface is performed, in which the second pixel coordinates are taken as independent variables and the corresponding second Pan parameters are taken as dependent variables, to obtain coefficients of the function, and then a function containing the coefficients is determined as the correspondence relationship between the pixel coordinates and the Pan parameter. The type of the surface used in the fitting process can be set in combination with actual conditions. At present, the correspondence relationship obtained by fitting is denoted as the first correspondence relationship.
[0167] Based on the second pixel coordinates (i.e., x, y) corresponding to each first sub-video picture and the second Tilt parameter, a function fitting of a surface is performed, in which the second pixel coordinates are taken as independent variables and the corresponding second Tilt parameters are taken as dependent variables, to obtain coefficients of the function, and then a function containing the coefficients is determined as the correspondence relationship between the pixel coordinates and the Tilt parameter. The type of the surface used in the fitting process can be set in combination with actual conditions. At present, the correspondence relationship obtained by fitting is denoted as the second correspondence relationship.
[0168] At this time, when the Pan parameter and the Tilt parameter obtained by substituting the pixel coordinates of a certain pixel point in the video picture captured by the first camera into the first correspondence relationship and the second correspondence relationship are used by the second camera, the central region of the video picture captured by the second camera includes the physical entity corresponding to the pixel coordinates.
[0169] Optionally, the first correspondence relationship and the second correspondence relationship can be determined in sequence or simultaneously.
[0170] The scheme two, the step 410 can specifically include steps 415-417:
[0171] The step 415, the eighth video picture based on the first camera shooting and the ninth video picture based on the second camera shooting at the same shooting moment are acquired.
[0172] In the process of determining the first correspondence and the second correspondence, the intelligent device instructs the first camera and the second camera to shoot at the same time. The shooting space when the first camera and the second camera shoot can be recorded as a fourth shooting space. Similar to the second shooting space, the fourth shooting space can be the same space as the first shooting space or a different space. The people and objects in the fourth shooting space can be stationary or moving, and the embodiments do not limit this.
[0173] Currently, the video picture shot by the first camera is recorded as the eighth video picture, and the video picture shot by the second camera is recorded as the ninth video picture. After the intelligent device instructs the first camera and the second camera to shoot at the same time, the first camera and the second camera will both feed back the current video data to the intelligent device. Then, the intelligent device can obtain the eighth video picture and the ninth video picture at the same shooting moment based on the two video data, and perform subsequent processing. It can be understood that when the intelligent device determines the first correspondence and the second correspondence, the eighth video picture and the ninth video picture at multiple shooting moments are needed.
[0174] The step 416, when the similarity between the ninth video picture and the fifth sub-video picture in the eighth video picture reaches a similarity threshold, the fifth pixel coordinate of the center point of the fifth sub-video picture in the eighth video picture is acquired, and the fourth Pan parameter and the fourth Tilt parameter used when the second camera shoots the ninth video picture are acquired. The second camera uses a fixed fourth Zoom parameter, and the fifth sub-video picture is one of the sub-video pictures obtained by dividing the eighth video picture.
[0175] For example, the intelligent device divides the video data shot by the first camera to obtain multiple regions. At this time, each video picture (i.e., each eighth video picture) in the video data will be divided into multiple sub-video pictures, that is, one sub-video picture corresponds to one region. The division process can refer to the division process in step 412. After the intelligent device starts to acquire the video data shot by the first camera, a region divided in the video data is selected. Optionally, when the sizes of the regions are not completely consistent, the region with the largest size can be selected.
[0176] After the region is selected, the intelligent device controls the second camera to aim at a corresponding shooting space (i.e., a sub-space in the fourth shooting space) of the region, and after aiming, the sub-video picture corresponding to the region in the video picture (i.e., the eighth video picture) shot by the first camera at the same shooting moment is located in the middle region of the ninth video picture shot by the second camera, wherein the aiming process can be realized by manually operating the second camera, or realized by manually operating the intelligent device to control the Pan parameter and the Tilt parameter of the second camera.
[0177] After the second camera is aimed, the Zoom parameter of the second camera is adjusted, and after the adjustment is completed, the similarity between the ninth video picture shot by the second camera and the sub-video picture (now recorded as the fifth sub-video picture) in the selected region in the eighth video picture shot by the first camera at the same shooting moment reaches a similarity threshold (which can be set in combination with actual conditions). That is, the ninth video picture is basically consistent with the fifth sub-video picture. In an optional manner, the intelligent device controls the second camera to continuously adjust the Zoom parameter (the Zoom parameter can be automatically adjusted by the intelligent device according to a pre-set adjustment range, or the Zoom parameter can be manually controlled by the intelligent device to be adjusted), and controls the first camera and the second camera to shoot, and after the eighth video picture and the ninth video picture at the same shooting moment are obtained each time, the similarity between the ninth video picture and the fifth sub-video picture in the eighth video picture is calculated, if the similarity reaches the similarity threshold, the second camera is fixed with the currently used Zoom parameter, and the Pan parameter and the Tilt parameter currently used by the second camera are recorded. The currently obtained Pan parameter and Tilt parameter can be considered as the fourth Pan parameter and the fourth Tilt parameter corresponding to the region where the fifth sub-video picture is located. The intelligent device further calculates the pixel coordinates (i.e., the pixel coordinates of the center point of the region where the fifth sub-video picture is located) of the center point of the fifth sub-video picture in the image pixel coordinate system of the first camera. At present, the pixel coordinates of the center point are recorded as the fifth pixel coordinates. At this time, the region where the fifth sub-video picture is located has the corresponding fourth Pan parameter, the fourth Tilt parameter and the fifth pixel coordinates, that is, the fifth pixel coordinates have the corresponding fourth Pan parameter and the fourth Tilt parameter. If the similarity does not reach the similarity threshold, the Zoom parameter is continuously adjusted until the calculated similarity reaches the similarity threshold.
[0178] Afterwards, the intelligent device controls the second camera to continue using the fixed Zoom parameter (the current one is recorded as the fourth Zoom parameter) and continuously adjust the Pan parameter and the Tilt parameter (the Pan parameter and the Tilt parameter can be automatically adjusted by the intelligent device according to the pre-set adjustment range, or the Pan parameter and the Tilt parameter can be manually controlled by the intelligent device to adjust). After each adjustment, the intelligent device continues to obtain the eighth video frame and the ninth video frame at the same shooting moment, and divides the eighth video frame to obtain one of the fifth video frames. The region corresponding to the current fifth video frame changes, that is, another region is currently used. Afterwards, the similarity between the fifth video frame and the ninth video frame is calculated, and if the similarity reaches the similarity threshold, the current Pan parameter and Tilt parameter used by the second camera are recorded. The current Pan parameter and Tilt parameter can be considered as the fourth Pan parameter and the fourth Tilt parameter corresponding to the region of the fifth video frame. The intelligent device also calculates the fifth pixel coordinate of the center point of the fifth video frame. If the similarity does not reach the similarity threshold, the Pan parameter and the Tilt parameter are continuously adjusted until the calculated similarity reaches the similarity threshold. Afterwards, the intelligent device repeats the foregoing process, that is, ensures that each region (that is, each region obtained after the video data is divided) has corresponding fourth Pan parameter, fourth Tilt parameter and fifth pixel coordinate.
[0179] In step 417, a first correspondence relationship between the pixel coordinates in the video data shot by the first camera and the Pan parameter of the second camera is determined according to the plurality of fifth pixel coordinates and the plurality of fourth Pan parameters, and a second correspondence relationship between the pixel coordinates in the video data shot by the first camera and the Tilt parameter of the second camera is determined according to the plurality of fifth pixel coordinates and the plurality of fourth Tilt parameters. The plurality of fifth pixel coordinates correspond to the fifth video frames obtained at the plurality of shooting moments, and the pixel coordinate ranges of the fifth video frames are all different.
[0180] The determination processes of the first correspondence relationship and the second correspondence relationship can refer to the determination processes of the first correspondence relationship and the second correspondence relationship in step 414.
[0181] It should be noted that the pixel coordinate ranges of the fifth sub-video pictures are all different, and it can be understood that the corresponding regions of the fifth sub-video pictures used for determining the fourth Tilt parameter and the fourth Pan parameter each time are all different in the video data. The size of the region after merging the corresponding regions of the fifth sub-video pictures is equal to the size of any eighth video picture. For example, the pixel coordinate range occupied by each frame of the eighth video picture in the video data captured by the first camera is 3840x2160, the video data is divided into 9x9 blocks, one block is used each time, and the similarity between the fifth sub-video picture and the ninth video picture in the block at the same shooting time is calculated, and then the fourth Pan parameter, the fourth Tilt parameter, and the fifth pixel coordinate corresponding to each region are obtained.
[0182] Step 420, establishing a third correspondence relationship between the face size in the video data captured by the first camera and the Zoom parameter of the second camera.
[0183] Before the first camera and the second camera are used to cooperate to realize shooting of the speaker, the third correspondence relationship needs to be established. At present, after the relative position relationship between the smart device and the second camera is fixed (at this time, the first camera and the second camera are directed to the same direction), the first correspondence relationship, the second correspondence relationship, and the third correspondence relationship can be established.
[0184] In one embodiment, step 420 includes steps 421-423:
[0185] Step 421, sequentially obtaining a plurality of fourth video pictures based on the video data captured by the first camera, and determining a second face size of the target face in each frame of the fourth video picture, the first camera shoots a third shooting space, the target face is included in the third shooting space, and the second face sizes corresponding to the fourth video pictures are all different.
[0186] When the third correspondence relationship is established, the video data captured by the first camera in the third shooting space is first obtained. At present, the third shooting space and the first shooting space can be the same space. In actual application, the third shooting space and the first shooting space can be different shooting spaces. For example, the second camera and the smart device are both installed on a support, and the support can be moved, so the user can push the support to different shooting spaces, but the relative position relationship between the smart device and the second camera does not change. The third shooting space and the second shooting space (or the fourth shooting space) can be the same space or different spaces. At present, there is a target object in the third shooting space, the target object is a person whose face can be completely shot by the first camera, and in the embodiment, the face of the target object is recorded as a target face, and the target face can be a front face.
[0187] In an embodiment, the target object continuously adjusts the distance between itself and the first camera in the third shooting space. After each adjustment of the distance (the magnitude of each adjustment is not limited at present), the target object first stops moving, and the smart device acquires the video data captured by the first camera (which can be acquired by manual indication), and then acquires a frame of video picture in the video data (which can be any frame of video picture in the video data captured after the target object temporarily stops moving), which is currently recorded as the fourth video picture, and the fourth video picture contains the target face. Alternatively, after the target object stops moving, the smart device controls the first camera to capture the third shooting space to obtain a picture, and takes the picture as the fourth video picture. It can be understood that the smart device can obtain the corresponding fourth video picture after each adjustment of the distance by the target object, and the distance between the target object and the first camera is different when each frame of the fourth video picture is captured.
[0188] It can be understood that the more the number of fourth video pictures (i.e., the number of adjustments of the distance by the target object) is, the more accurate the third correspondence determined is, but the larger the amount of calculation is. Therefore, the number of fourth video pictures can be set according to the actual situation.
[0189] For example, after each fourth video picture is obtained, the size of the target face in the fourth video picture is calculated, which is currently recorded as the second face size. The calculation method of the second face size is the same as that of the first face size. Each frame of the fourth video picture has a corresponding second face size. It can be understood that when the distance between the target object and the first camera is different, the second face size of the target face in the video picture captured by the first camera will also be different, i.e., the second face size corresponding to each frame of the fourth video picture will be different.
[0190] In actual application, the second face size of the target face in each frame of the fourth video picture can also be calculated after all the fourth video pictures are obtained.
[0191] Step 422: When each frame of the fourth video picture is acquired, the third Zoom parameter corresponding to the fourth video picture is also acquired, and the third face size of the target face in the fifth video picture captured by the second camera under the third Zoom parameter, the third Pan parameter and the third Tilt parameter satisfies the preset face size.
[0192] Optionally, after each stop of the target object, the intelligent device further controls the second camera to continuously adjust the Zoom parameter (the Zoom parameter can be automatically adjusted by the intelligent device according to a pre-set adjustment range, or the intelligent device can be manually controlled to adjust the Zoom parameter), and obtains a video picture (which can be a picture obtained by shooting, or a picture cut from video data obtained by shooting) of the second camera under each Zoom parameter. At present, the obtained video picture is recorded as a fifth video picture. When each fifth video picture is obtained, the face size of the target face in the fifth video picture is calculated. At present, the face size of the target face in the fifth video picture is recorded as a third face size. The third face size can be calculated by performing face recognition on the fifth video picture to obtain a face picture, and the third face size is determined in combination with the face picture. The process of determining the third face size in combination with the face picture can refer to the process of determining the first face size in combination with the face picture. After obtaining the third face size, it is determined whether the third face size meets a pre-set face size (which can be set according to actual conditions). If the third face size meets the pre-set face size, it indicates that the target face occupies a certain proportion in the fifth video picture (for example, the target face occupies 1 / 3 of the fifth video picture), and the user can clearly see the target face through the fifth video picture, that is, the zoom close-up of the target face is realized. At this time, the Zoom parameter currently used by the second camera is obtained and recorded as a third Zoom parameter. Otherwise, the second camera continues to adjust the Zoom parameter until the third face size meets the pre-set face size.
[0193] It can be understood that after the target object stops moving, the intelligent device can obtain the fourth video picture of the target object and the third Zoom parameter, and at this time, the corresponding relationship between the fourth video picture and the third Zoom parameter can be established. Optionally, since the target face of the target object does not change in size in the video picture captured by the first camera after the target object stops moving, the intelligent device can control the first camera to capture a fourth video picture and calculate the second face size each time the target object stops moving, and control the second camera to continuously adjust the Zoom parameter to obtain a suitable third Zoom parameter. It can also be understood that after the target object stops moving, the intelligent device can control the first camera to capture a video picture and control the second camera to capture a fifth video picture (i.e., the shooting time of the two cameras is the same) each time the Zoom parameter is adjusted, and when the third face size in the fifth video picture meets the preset face size, use the video picture obtained by the first camera at the same shooting time as the fourth video picture and calculate the second face size in the fourth face picture. Alternatively, after the target object stops moving, the intelligent device can control the second camera to capture a fifth video picture each time the Zoom parameter is adjusted, and when the third face size in the fifth video picture meets the preset face size, control the first camera to capture a fourth video picture and calculate the second face size in the fourth face picture.
[0194] After the intelligent device obtains the third Zoom parameter corresponding to the fourth video picture, the target object continues to move (which can be instructed by the intelligent device) to change the distance between the target object and the first camera, and when the target object stops moving, the intelligent device obtains the third face size in the fifth video picture and the Zoom parameter currently used by the second camera when the third face size meets the preset face size, and records it as the third Zoom parameter corresponding to the fourth video picture. Repeat this process until each fourth video picture has a corresponding third Zoom parameter. It can be understood that even if the distance between the target object and the first camera is different, by setting a suitable Zoom parameter (currently corresponding to the third Zoom parameter), the target object can occupy a certain proportion in the video data captured by the second camera, so as to realize the close-up of the target object at different distances.
[0195] Optionally, after each movement of the target object, the intelligent device first adjusts the Zoom parameter of the second camera to the minimum multiple (i.e., the video picture corresponds to the largest spatial range), and then continuously adjusts the Zoom parameter to continuously reduce the spatial range corresponding to the captured video picture.
[0196] Optionally, after each movement of the target object, the intelligent device first controls the second camera to aim at the target face, and then determines the third Zoom parameter. Aiming at the target face can be understood as that the target face is contained in the video data captured by the second camera, and the center point of the target face is in a set pixel coordinate region of the fifth video frame (which can be set according to actual conditions, and is generally a region in the middle of the video frame). At this time, it can be considered that the target face is located in the middle region of the video frame. The aiming process can be realized by manually operating the second camera, or realized by manually operating the intelligent device to control the Pan parameter and the Tilt parameter of the second camera. When the second camera is adjusted to the position corresponding to the appropriate Pan parameter and Tilt parameter, the target face can be aimed at. When it is determined that the target face has been aimed at (this process can be manually confirmed or automatically confirmed by the intelligent device), the intelligent device obtains the Pan parameter and the Tilt parameter currently used by the second camera, and records them as the third Pan parameter and the third Tilt parameter. At this time, after each movement of the target object, the third Pan parameter and the third Tilt parameter used by the second camera can be the same or different. Alternatively, after the intelligent device aims at the target face using the second camera for the first time, the third Pan parameter and the third Tilt parameter are fixedly used, and after each movement of the target object, the second camera uses the same third Pan parameter and the third Tilt parameter. It can be understood that generally, after each movement of the target object, the target object can be captured by the second camera. In actual application, the second camera can also not aim at the target object, as long as the target face appears in the fifth video frame captured by the second camera.
[0197] Step 423, determining a third correspondence relationship between the face size in the video data captured by the first camera and the Zoom parameter of the second camera according to the second face size corresponding to each fourth video frame and the third Zoom parameter.
[0198] Based on the second face size corresponding to each fourth video frame and the third Zoom parameter, a function fitting of a line (which can be a straight line or a curve) is performed, in which the second face size is taken as the independent variable, and the corresponding third Zoom parameter is taken as the dependent variable, to obtain the correspondence relationship between the face size and the Zoom parameter. The line type of the line used in the fitting process can be set according to actual conditions.
[0199] At present, the obtained correspondence relationship is recorded as the third correspondence relationship. When the Zoom parameter obtained by substituting the face size coordinate of the speaker in the video frame captured by the first camera into the third correspondence relationship is used by the second camera, the video frame captured by the second camera can be a close-up frame of the speaker.
[0200] Optionally, the third correspondence relationship can be determined first, and then the first correspondence relationship and the second correspondence relationship are determined.
[0201] After the first correspondence relationship, the second correspondence relationship and the third correspondence relationship are obtained, the first correspondence relationship, the second correspondence relationship and the third correspondence relationship can be applied. When the relative position relationship of the first camera and the second camera does not change, the first correspondence relationship, the second correspondence relationship and the third correspondence relationship can be continuously used. That is, when the second camera is installed, the first correspondence relationship, the second correspondence relationship and the third correspondence relationship are determined, and when the first correspondence relationship, the second correspondence relationship and the third correspondence relationship are needed in the subsequent use, the first correspondence relationship, the second correspondence relationship and the third correspondence relationship do not need to be determined again. At present, the first correspondence relationship, the second correspondence relationship and the third correspondence relationship are applied once as an example for description, at this time, the application of the first correspondence relationship, the second correspondence relationship and the third correspondence relationship can include steps 430-4120:
[0202] Step 430, obtaining first video data shot by the first camera, the first camera being used for shooting a first shooting space.
[0203] Step 440, identifying a speaker in a first video picture currently processed in the first video data.
[0204] Exemplarily, after obtaining the first video picture currently processed in the first video data, the speaker appearing in the first video picture is identified. Optionally, the speaker in the first video picture can be selected by manual operation, or the speaker in the first video picture can be automatically identified by the intelligent device. At present, the speaker is automatically identified by the intelligent device as an example for description. At this time, step 440 includes steps 441-442:
[0205] Step 441, performing face recognition on the first video picture currently processed in the first video data and a plurality of sixth video pictures continuously before the first video picture in the first video data, and obtaining a lip picture on each face identified.
[0206] Exemplarily, the face recognition on the first video picture is performed by using a deep learning method in an artificial intelligence algorithm (such as an Openvino or other deep learning tool package). The face recognition by using the deep learning method is a technology that has been realized, and will not be described herein. Through the face recognition, each face contained in the first video picture can be determined. At present, the first video picture can contain one or more faces. It can be understood that there is a case that no face is contained, at this time, the processing of the first video picture is abandoned, and the next video picture needing to be processed in the first video data is selected and the face recognition is continuously performed.
[0207] In the same way, the video frames (the sixth video frame is recorded for the moment) before the first video frame in the first video data are processed for face recognition. The number of frames can be set according to actual conditions, and is generally the number of frames required to determine the speaker.
[0208] After each face in the first video frame and the sixth video frame is recognized, a short video data of each face is obtained, which is composed of the face frames of the same face in the first video frame and the sixth video frame. Alternatively, the face frames of each face in the first video frame and the sixth video frame are obtained by cutting out or other methods, and the face frames of the same face are combined to form a short video data containing the face, which is recorded as face video data for the moment. Then, each lip frame in the face video data is identified, for example, by using artificial intelligence to identify the lip frame in each face frame in the face video data. The lip frame can reflect the opening and closing of the corresponding face's lips in the continuous time of the sixth video frame and the first video frame.
[0209] In one embodiment, the approximate direction of the speaker can also be determined in combination with the microphone array, and then the speaker is identified in the first video frame in combination with the direction of the speaker. At this time, the smart device also includes a microphone array, and step 441 specifically includes: obtaining the DOA information of the speaker based on the microphone array; according to the DOA information, the third sub-video frame is cut out in the first video frame currently processed in the first video data, and the fourth sub-video frame is cut out in the sixth video frame before the first video frame; face recognition is performed on the third sub-video frame and the fourth sub-video frame, and the lip frame on each face recognized is obtained.
[0210] For example, the microphone array can collect the speech content of the speaker (i.e., collect the voice of the speaker), and then based on the voice of the speaker collected by the microphone array, the angle of the speaker in the HFOV of the microphone array relative to the microphone array itself can be obtained by using the Direction Of Arrival (DOA) technology. The HFOV refers to the horizontal field of view, which can represent the collection range of the microphone array in the horizontal direction, i.e., the included angle between the two edges of the collection range of the microphone array in the horizontal direction can be taken as the HFOV of the microphone array. The angle of the speaker in the HFOV of the microphone array can be understood as the included angle of the line connecting the speaker with an edge (a pre-specified edge) in the collection range in the horizontal direction, for example, FIG. 6 is a schematic diagram of the horizontal field of view of a microphone array provided by an embodiment of the present application, referring to FIG. 6, the HFOV of the microphone array (i.e., the collection range of the microphone array in the horizontal direction) corresponds to a region 51, and a speaker 52 is located in the collection range. When the speaker speaks, the smart device can determine the direction of the speaker 52 in the horizontal field of view based on the sound signal collected by the microphone array and can express the direction by an angle, which can be the included angle a of the line connecting the speaker 52 with an edge (pre-specified) of the region 51. This angle can be considered as the angle of the speaker in the HFOV of the microphone array. In the embodiment, the angle of the speaker in the HFOV of the microphone array determined is recorded as DOA information. Based on the DOA information, the direction of the speaker relative to the microphone array can be known.
[0211] It can be understood that the first camera also has its own HFOV, which can represent the shooting range (or the angle range of the received image) of the first camera in the horizontal direction, i.e., the included angle between the two edges of the shooting range in the horizontal direction can be taken as the HFOV of the first camera. The range of the HFOV of the first camera and the range of the HFOV of the microphone array can be different, therefore, in the embodiment, a mapping relationship between the HFOV of the microphone array and the HFOV of the first camera is pre-established, based on which the angle in the HFOV range of the microphone array can be converted into the HFOV range of the first camera to obtain the angle in the HFOV of the first camera. For example, after determining the angle of the speaker in the HFOV of the microphone array (referring to FIG. 6), the angle of the speaker in the HFOV of the first camera (i.e., the included angle of the line connecting the speaker with a pre-specified edge in the HFOV of the first camera) can be obtained based on the mapping relationship, which can also represent the direction of the speaker relative to the first camera. It should be noted that the mapping relationship can be calculated based on the position parameters, intrinsic parameters, etc. of the microphone array and the first camera.
[0212] After the angle of the speaker in the HFOV of the first camera is obtained, an angle range is obtained according to the obtained angle as the center and a preset angle range. For example, the obtained angle is 10° and the preset angle range is 3°, then the angle range is 10°±3°, that is, the angle range of 7°-13°. Then, the region corresponding to the angle range in the first video picture is found, and the region is cropped from the first video picture to obtain a sub-video picture. At present, the cropped sub-video picture is recorded as a third sub-video picture. It can be understood that each angle in the HFOV of the first camera can find a corresponding perpendicular line in the video picture shot by the first camera, and the cropped sub-video picture is taken as the center line. For example, FIG. 7 is a schematic diagram of a third sub-video picture provided by an embodiment of the present application. Referring to FIG. 7, the angle of the speaker in the HFOV of the first camera is 10°, and the perpendicular line 62 in the first video picture 61 shot by the first camera corresponds to 10° in the HFOV. Therefore, the region 63 in the angle range of 10°±3° is taken as the third sub-video picture with the perpendicular line 62 as the center line. The third sub-video picture can be considered as a video picture containing the speaker, and the size of the third sub-video picture is smaller than the size of the first video picture, so that the workload of subsequent face recognition and lip recognition can be reduced. Each frame of the sixth video picture is cropped in the same way (the angle range used for cropping is consistent with the angle range used for cropping the first video picture), and the cropped sub-video picture is recorded as a fourth sub-video picture.
[0213] After the third sub-video picture and the fourth sub-video picture are obtained, the face recognition can be performed on the third sub-video picture and the fourth sub-video picture by using the deep learning method in the artificial intelligence algorithm, and then the lip picture of each face is obtained.
[0214] In step 442, the speaker in the first video picture is determined according to the motion amplitude of the lips in the lip picture. The motion amplitude of the lips of the speaker is the largest.
[0215] For example, the motion amplitude of the lips in the face can be determined based on the respective lip pictures corresponding to the same face (the respective lip pictures are arranged in the order of the corresponding video pictures in the video data). The motion amplitude can be calculated by the amplitude, frequency, and duration of the lip vibration (which can also be understood as the opening and closing of the lips). For example, different amplitudes correspond to different scores (the greater the amplitude, the higher the score), different frequencies also correspond to different scores (the higher the frequency, the higher the score), and different durations also correspond to different scores (the longer the duration, the higher the score). In addition, the amplitude, frequency, and duration each have a corresponding weight, and the weight can be flexibly adjusted in combination with the actual situation. Then, the amplitude score of the lip vibration (which can be the maximum amplitude or the average of the amplitude corresponding to each lip picture) is determined based on the respective lip pictures corresponding to the same face, the frequency (determined based on the duration and the number of opening and closing of the respective lip pictures) score of the lip vibration, and the duration score of the lip vibration. Then, the scores of the determined amplitude, frequency, and duration are weighted, and the obtained value can be used as the score of the lips corresponding to the face. Alternatively, after determining the amplitude, frequency, and duration scores of the lip vibration, it is first determined whether the amplitude, frequency, and duration scores are all valid. For example, the amplitude, frequency, and duration each have a corresponding threshold value. If the score of a certain dimension (amplitude, frequency, or duration) does not reach the corresponding threshold value, it is determined that the score of the dimension is invalid. If the scores of the three dimensions are all valid, the scores of the three dimensions are weighted, and the obtained value can be used as the score of the lips corresponding to the face. If the score of any one dimension is invalid, there is no need to weight the scores of the other two dimensions, i.e., the face corresponding to the lips is determined to be the face of the speaker.
[0216] After determining the score of the lips corresponding to each face, the face in which the lips with the highest score (i.e., the maximum motion amplitude) are selected as the face of the speaker, so as to identify the speaker in the first video picture.
[0217] It can be understood that since the motion amplitude of the lips needs to be calculated based on multiple video pictures, when the first video data is obtained, the first several video pictures (the number of which is less than the minimum number of video pictures required to determine the motion amplitude of the lips) cannot be identified as the speaker, and at this time, the second camera cannot be used to accurately capture the speaker.
[0218] It can be understood that identifying the same face in multiple video pictures is a technology that has been realized, and will not be described in detail. Even if the speaker is in a moving state, the position of the speaker in consecutive video pictures is still relatively close, so the speaker can be accurately identified based on the lip pictures of the speaker in consecutive video pictures.
[0219] It should be noted that in the current process of identifying the speaker, the facial images identified can be either frontal images or profile images including the lips.
[0220] Step 450: Obtain the speaker's face in the second sub-video frame within the first video frame.
[0221] Once the speaker is identified, their face can be extracted from the first video frame. Currently, the extracted face image is recorded as the second sub-video frame. Optionally, the extracted face image can be a frame containing only the speaker's face, a frame containing the head where the face is located, or a frame containing both the head and neck.
[0222] Step 460: Determine the pixel coordinates of the center point of the second sub-video frame within the first video frame and use them as the first pixel coordinates; determine the size of the second sub-video frame and use it as the first face size.
[0223] For example, after obtaining the second sub-video frame, the center point of the second sub-video frame is determined. For instance, a rectangular area containing the second sub-video frame is determined, and the center point of this rectangular area can be used as the center point of the second sub-video frame. Other methods can also be used to determine the center point. Then, the pixel coordinates of the center point in the image pixel coordinate system used by the first camera are determined, and these pixel coordinates are used as the first pixel coordinates of the center point within the first video frame.
[0224] After obtaining the second sub-video frame, its size within the image pixel coordinate system used by the first camera is determined. For example, a rectangular area containing the second sub-video frame can be identified, and its size can be obtained based on its width, height, or diagonal length (all represented in pixels) in the image pixel coordinate system. Other methods can also be used to determine the size of the second sub-video frame. This determined size can then be used as the size of the first face within the first video frame from the second sub-video frame.
[0225] Step 470: Determine the first Pan parameter corresponding to the first pixel coordinates according to the first correspondence relationship. The first correspondence relationship is the correspondence between the pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera.
[0226] For example, by substituting the coordinates of the first pixel as the independent variable into the first correspondence, the resulting Pan parameter can be used as the first Pan parameter.
[0227] Step 480: Determine the first Tilt parameter corresponding to the first pixel coordinates according to the second correspondence relationship. The second correspondence relationship is the correspondence between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera.
[0228] Exemplarily, the Tilt parameter obtained by bringing the first pixel coordinate as the independent variable into the second correspondence relationship can be taken as the first Tilt parameter. The speaker is located in the middle region in the video picture taken by the second camera using the first Pan parameter and the first Tilt parameter.
[0229] Step 490, determining the first Zoom parameter corresponding to the first face size according to a third correspondence relationship, the third correspondence relationship being a correspondence relationship between the face size in the video data taken by the first camera and the Zoom parameter of the second camera.
[0230] Exemplarily, the Zoom parameter obtained by bringing the first face size as the independent variable into the third correspondence relationship can be taken as the first Zoom parameter. The video picture taken by the second camera using the first Pan parameter, the first Tilt parameter and the first Zoom parameter can be considered as the enlarged close-up picture with the speaker located in the middle region.
[0231] Optionally, the steps 470-490 can be executed in sequence or simultaneously.
[0232] Step 4100, controlling the second camera to take pictures using the first shooting control parameter, so that the second camera realizes the aligned shooting of the speaker under the first shooting control parameter.
[0233] Step 4110, obtaining the second video data taken by the second camera.
[0234] Step 4120, obtaining the second video data by the virtual camera driver program and sending the second video data to the target application program, so that the target application program can use the second video data on the display screen.
[0235] Through the reasonable determination of the first correspondence relationship, the second correspondence relationship and the third correspondence relationship, the shooting control parameter can be determined through the first correspondence relationship, the second correspondence relationship and the third correspondence relationship, and the enlarged close-up picture with the speaker located in the middle region can be taken by the second camera using the shooting control parameter. Based on the face recognition and the lip movement amplitude recognition, the speaker in the video picture can be automatically determined. In combination with the DOA information of the microphone array, the direction of the speaker can be determined, and then the sub-video picture corresponding to the direction where the speaker is located in the video picture can be intercepted, and the speaker can be recognized in the sub-video picture, so that the data processing amount of the video picture can be reduced on the premise of ensuring the accuracy of the speaker recognition.
[0236] In an embodiment of the present application, after the second video data captured by the second camera is obtained, it is further determined whether the face picture of the speaker in the second video data is a high-quality close-up picture. If a high-quality close-up picture is not obtained, the first shooting control parameter can be fine-tuned to ensure that a high-quality close-up picture is captured. At this time, after step 4110, the method can further include: determining third pixel coordinates of the speaker in a seventh video picture currently processed in the second video data and a fourth face size; when the third pixel coordinates are located in a set pixel coordinate region and the fourth face size satisfies a preset face size, performing an operation of obtaining the second video data by the virtual camera driver, the preset shooting condition being that the third pixel coordinates are located in the set pixel coordinate region and the fourth face size satisfies the preset face size. When the third pixel coordinates and the fourth face size do not satisfy the preset shooting condition, adjusting the first shooting control parameter according to the third pixel coordinates and the fourth face size to obtain a second shooting control parameter, the second camera performing alignment shooting on the speaker under the second shooting control parameter and a video picture captured by the second camera satisfying the preset shooting condition; obtaining third video data captured by the second camera under the second shooting control parameter; obtaining the third video data by the virtual camera driver and sending the third video data to the target application program, so that the target application program can use the third video data on the display screen.
[0237] For example, after the second video data captured by the second camera is obtained, a video picture selected from the second video data is taken as the seventh video picture currently processed. The selected seventh video picture can be a first video picture obtained after the second camera uses the first shooting control parameter to perform shooting.
[0238] The pixel coordinates of the speaker in the seventh video picture (currently denoted as third pixel coordinates) and the face size (currently denoted as fourth face size) are determined. The face of the speaker can be recognized in the seventh video picture first, and the third pixel coordinates and the fourth face size are determined in combination with the face. The determination manner of the third pixel coordinates can refer to the determination manner of the first pixel coordinates, and the determination manner of the fourth face size can refer to the determination manner of the first face size.
[0239] After the third pixel coordinate and the fourth face size are obtained, it is determined whether the third pixel coordinate and the fourth face size satisfy a preset shooting condition. The preset shooting condition is that the third pixel coordinate is located in a set pixel coordinate region and the fourth face size satisfies a preset face size. The set pixel coordinate region refers to a region selected with the center point of the video picture as the center, which can be preset. When the third pixel coordinate is located in the set pixel coordinate region, it can be considered that the face of the speaker is located in the middle region of the seventh video picture. The preset face size is the same as the preset face size used when the third correspondence is determined. The third face size satisfying the preset face size indicates that the speaker occupies a certain proportion (such as 1 / 3) in the seventh video picture. When the preset shooting condition is satisfied, it indicates that the seventh video picture is an enlarged close-up picture with the speaker located in the middle region, that is, the second camera shoots a high-quality close-up picture. Therefore, the second camera can use the first shooting control parameter for shooting, and step 4120 is executed so that the target application uses the high-quality close-up picture of the speaker.
[0240] If the third pixel coordinate and the fourth face size do not satisfy the preset shooting condition, it indicates that the seventh video picture is not a high-quality close-up picture. Either the third pixel coordinate is not located in the set pixel coordinate region or the fourth face size does not satisfy the preset face size can be considered as not satisfying the preset shooting condition. At this time, the intelligent device fine tunes the first shooting control parameter, and the shooting control parameter obtained after fine tuning is recorded as the second shooting control parameter. It can be understood that the first shooting control parameter can be fine tuned multiple times at present until the video picture in the video data obtained by the second camera using the second shooting control parameter can satisfy the preset shooting condition.
[0241] The fine tuning process can be implemented by bisection method: when the third pixel coordinate is not located in the set pixel coordinate region, the relative positional relationship between the third pixel coordinate and the center point of the set pixel coordinate region can be determined, which can include that the third pixel coordinate is located on the upper side, the lower side, the left side, the right side, the upper left side, the upper right side, the lower left side, and the lower right side of the center point. The adjustment rules of the Pan parameter and the Tilt parameter (that is, the rules of adding or subtracting the Pan parameter and the Tilt parameter) under different relative positional relationships are pre-stored in the intelligent device. Then, the first Pan parameter and the first Tilt parameter are each halved to obtain a new Pan parameter (half of the first Pan parameter) and a new Tilt parameter (half of the first Tilt parameter), the new Pan parameter is calculated according to the calculation mode of the adjustment rule to obtain a fifth Pan parameter, and the new Tilt parameter is calculated according to the calculation mode of the adjustment rule to obtain a fifth Tilt parameter.
[0242] When the error between the fourth face size and the preset face size is not lower than the preset error, it can be considered that the fourth face size does not meet the preset face size. The calculation method of the error is not limited at present, for example, the calculated width difference and height difference are taken as the error. At present, the comparison result (whether greater than or less than the preset face size) of the fourth face size and the preset face size is determined, and the adjustment rule (i.e. the rule of adding or subtracting the Zoom parameter) corresponding to different comparison results is determined, then the first Zoom parameter is halved to obtain a new Zoom parameter (half of the first Zoom parameter), and the new Zoom parameter and the first Zoom parameter are calculated according to the calculation method of the adjustment rule to obtain a fifth Zoom parameter.
[0243] It can be understood that when the third pixel coordinate is located in the set pixel coordinate region but the fourth face size does not meet the preset face size, the first Pan parameter and the first Tilt parameter can be directly taken as the fifth Pan parameter and the fifth Tilt parameter, and the fifth Zoom parameter is determined. When the third pixel coordinate is not located in the set pixel coordinate region but the fourth face size meets the preset face size, the first Zoom parameter can be directly taken as the fifth Zoom parameter, and the fifth Pan parameter and the fifth Tilt parameter are determined.
[0244] After obtaining the fifth Pan parameter, the fifth Tilt parameter and the fifth Zoom parameter, the second camera is controlled to use the fifth Pan parameter, the fifth Tilt parameter and the fifth Zoom parameter (this process can refer to the process of controlling the second camera to use the first shooting control parameter). Then, the video data shot by the second camera is obtained, and the pixel coordinate and the face size of the speaker in the video picture of the video data are continuously determined, and whether it meets the preset shooting condition is judged.
[0245] If the preset shooting condition is met, the fifth Pan parameter, the fifth Tilt parameter and the fifth Zoom parameter are determined as the second shooting control parameter. Then, the video data (currently referred to as third video data) shot by the second camera using the second shooting control parameter is obtained, and the third video data is provided for the virtual camera driver program to obtain and send to the target application for use.
[0246] If the preset shooting condition is not met, the fifth Pan parameter, the fifth Tilt parameter and the fifth Zoom parameter are continuously adjusted. When the fifth Pan parameter is adjusted, the parameter used for addition and subtraction calculation with the fifth Pan parameter is half of half of the first Pan parameter (i.e., half of the first Pan parameter is divided into two). When the fifth Tilt parameter is adjusted, the parameter used for addition and subtraction calculation with the fifth Tilt parameter is half of half of the first Tilt parameter. When the fifth Zoom parameter is adjusted, the parameter used for addition and subtraction calculation with the fifth Zoom parameter is half of half of the first Zoom parameter. Then, the second camera is controlled to use the adjusted fifth Pan parameter, the fifth Tilt parameter and the fifth Zoom parameter. Video data captured by the second camera is obtained, and the pixel coordinates and the face size of the speaker in the video frame of the video data are continuously determined to determine whether the preset shooting condition is met. This process is repeated until the pixel coordinates and the face size of the speaker in the video frame of the video data captured by the second camera meet the preset shooting condition. At this time, the Pan parameter, the Tilt parameter and the Zoom parameter used by the second camera can be used as the second shooting control parameter. Video data captured by the second camera using the second shooting control parameter (currently referred to as third video data) is obtained, and the third video data is provided for the virtual camera driver program to obtain and send to the target application program.
[0247] The fine adjustment process can be implemented in other ways. For example, when the third pixel coordinates are not located in the set pixel coordinate region, the pixel distance (which can be determined by the absolute value of the difference between the coordinate values on the X axis) on the X axis and the pixel distance (which can be determined by the absolute value of the difference between the coordinate values on the Y axis) on the Y axis between the third pixel coordinates and the center point of the set pixel coordinate region in the image pixel coordinate system of the second camera, and the relative position relationship between the third pixel coordinates and the center point of the set pixel coordinate region can be determined. The relative position relationship can include that the third pixel coordinates are located on the upper side, the lower side, the left side, the right side, the upper left side, the upper right side, the lower left side and the lower right side of the center point. The adjustment amplitudes of the Pan parameter and the Tilt parameter under different X axis pixel distances, Y pixel distances and different relative position relationships are pre-stored in the intelligent device (which can be set according to actual conditions). Then, the adjustment amplitudes of the Pan parameter and the Tilt parameter are determined according to the determined X axis pixel distance, Y pixel distance and relative position relationship, and the first Pan parameter and the first Tilt parameter are adjusted based on the adjustment amplitudes. Currently, the adjusted parameters are used as the fifth Pan parameter and the fifth Tilt parameter, respectively.
[0248] When the error between the fourth face size and the preset face size is not lower than the preset error, it can be considered that the fourth face size does not meet the preset face size. The calculation method of the error is not limited at present, for example, the calculated width difference and height difference are taken as the error. The adjustment range of the Zoom parameter corresponding to different errors is pre-stored in the intelligent device. The adjustment range of the Zoom parameter can be determined according to the currently calculated error, and the first Zoom parameter is adjusted based on the adjustment range. At present, the parameter obtained after the adjustment is taken as the fifth Zoom parameter.
[0249] It can be understood that when the third pixel coordinate is located in the set pixel coordinate region but the fourth face size does not meet the preset face size, the first Pan parameter and the first Tilt parameter can be directly taken as the fifth Pan parameter and the fifth Tilt parameter, and the fifth Zoom parameter is determined. When the third pixel coordinate is not located in the set pixel coordinate region but the fourth face size meets the preset face size, the first Zoom parameter can be directly taken as the fifth Zoom parameter, and the fifth Pan parameter and the fifth Tilt parameter are determined.
[0250] After obtaining the fifth Pan parameter, the fifth Tilt parameter and the fifth Zoom parameter, the second camera is controlled to use the fifth Pan parameter, the fifth Tilt parameter and the fifth Zoom parameter (the process can refer to the process of controlling the second camera to use the first shooting control parameter). Then, the video data shot by the second camera is obtained, and the pixel coordinate and the face size of the speaker in the video picture of the video data are continuously determined, and it is judged whether the preset shooting condition is met. When the preset shooting condition is not met, the fifth Pan parameter, the fifth Tilt parameter and the fifth Zoom parameter are continuously adjusted in the manner described above until the pixel coordinate and the face size of the speaker in the video picture of the video data shot by the second camera meet the preset shooting condition. The Pan parameter, the Tilt parameter and the Zoom parameter used by the second camera when the preset shooting condition is met are taken as the second shooting control parameter. The video data (currently referred to as third video data) shot by the second camera using the second shooting control parameter is obtained, and the third video data is provided for the virtual camera driver program to obtain and send to the target application program for use.
[0251] Optionally, the process of determining the second shooting control parameter can be implemented by a local AI algorithm module, and the process of controlling the second camera to use the second shooting control parameter can be implemented by a central logic control module.
[0252] The above, only when the pixel coordinates and the face size of the speaker in the video picture captured by the second camera satisfy the preset shooting condition, it is determined that the second camera has captured the enlarged close-up picture of the speaker located in the middle region, at this time, the video data captured by the second camera is sent to the virtual camera driver program, so that the target application obtains the enlarged close-up picture of the speaker located in the middle region. And when the pixel coordinates and the face size of the speaker in the video picture captured by the second camera do not satisfy the preset shooting condition, the shooting control parameters used by the second camera are fine-tuned to ensure that the second camera can capture the enlarged close-up picture of the speaker located in the middle region.
[0253] In an embodiment of the present application, when the second video data containing the close-up picture of the speaker is provided to the target application program, it can also be detected whether the position of the speaker changes to realize the tracking shooting of the speaker. At this time, when step 4110 is executed, the following steps are also executed: updating the first video picture currently processed; identifying the fourth pixel coordinates of the speaker in the updated first video picture; when the coordinate distance between the fourth pixel coordinates and the first pixel coordinates is not greater than the preset distance, it is determined that the speaker has not changed and the operation of obtaining the second video data captured by the second camera is returned to be executed to realize the continuous shooting of the speaker.
[0254] For example, when the second video data captured by the second camera is obtained, in addition to sending the second video data to the virtual camera driver program, the first video data captured by the first camera is also continuously obtained, and the next frame of video picture to be processed is obtained after obtaining the first video picture in the first video data. The next frame of video picture to be processed can be separated from the first video picture by a preset number of frames (the preset number of frames is greater than or equal to 0). That is, the video picture separated from the first video picture by a preset number of frames is taken as the next video picture to be processed after obtaining the first video picture. The next video picture to be processed is updated as the first video picture currently processed, that is, the first video picture currently processed is updated, and the number of frames between two adjacent first video pictures is greater than or equal to 0.
[0255] Then, the pixel coordinates of the speaker in the first video picture are determined. At present, when the second video data captured by the second camera is obtained, the determined pixel coordinates of the speaker in the first video picture are marked as the fourth pixel coordinates. Optionally, the speaker needs to be re-identified in the first video picture (for details, please refer to the related description of step 440), and the fourth pixel coordinates of the speaker are determined.
[0256] Afterwards, a coordinate distance between the fourth pixel coordinate and the first pixel coordinate is calculated. It is determined whether the coordinate distance is greater than a distance threshold (a corresponding value can be set according to actual conditions). If the coordinate distance is not greater than the distance threshold, it is indicated that the speaker does not change. The change of the speaker can be that the speaker is replaced (i.e., the speaker at one position is replaced by the speaker at another position) or the position of the same speaker is changed (i.e., the speaker moves from one position to another position). When the speaker changes, the pixel coordinate of the speaker in the first video data also changes greatly. Therefore, when the coordinate distance is not greater than the distance threshold, it is indicated that the speaker does not change. At this time, the shooting control parameter used by the second camera does not need to change, and the smart device can continue to acquire the second video data and send the second video data to the virtual camera driver. In addition, the smart device continues to update the first video picture currently processed to continue to determine whether the coordinate distance is greater than the distance threshold, i.e., to continue to determine whether the speaker changes.
[0257] Optionally, when it is determined that the speaker does not change, the fourth pixel coordinate newly determined is updated as the first pixel coordinate. In this way, the latest first pixel coordinate is used to calculate the coordinate distance each time, so that the coordinate distance can reflect the latest position change of the speaker.
[0258] In one embodiment, when the coordinate distance between the fourth pixel coordinate and the first pixel coordinate is greater than the distance threshold, the first video picture currently processed is continuously updated, the fourth pixel coordinate of the speaker in the updated first video picture is continuously recognized, and the coordinate count is started. The coordinate count is increased by 1 each time the fourth pixel coordinate is determined. If the current coordinate count reaches a count threshold, and the coordinate distance between the multiple fourth pixel coordinates determined in the current coordinate count and the first pixel coordinate are all greater than the distance threshold, it is determined that the speaker changes, and the operation of recognizing the first pixel coordinate and the first face size of the speaker in the first video picture currently processed in the first video data is performed again, so that the second camera is aligned to shoot the changed speaker. If the current coordinate count does not reach the count threshold, and the coordinate distance between the fourth pixel coordinate newly determined and the first pixel coordinate is not greater than the distance threshold, it is determined that the speaker does not change, and the operation of acquiring the second video data shot by the second camera is performed again, so that the speaker is continuously shot.
[0259] For example, when the coordinate distance between the fourth pixel coordinate and the first pixel coordinate is greater than the distance threshold, it is determined that the speaker may have changed, and it is necessary to further determine whether the speaker has truly changed, i.e., whether the change of the speaker belongs to a real event. The process of determining whether the speaker has truly changed can be: continue to obtain the next processed video frame in the first video data and update it as the first video frame (i.e., update the first video frame), and again identify the fourth pixel coordinate of the speaker in the first video frame. And start the coordinate count, record the current coordinate count as 1, which indicates that the new fourth pixel coordinate is determined again after the first determination that the coordinate distance between the fourth pixel coordinate and the first pixel coordinate is greater than the distance threshold. Then, the coordinate distance between the newly determined fourth pixel coordinate and the first pixel coordinate is calculated again, and it is determined whether the coordinate distance is greater than the distance threshold. If the coordinate distance is greater than the distance threshold, continue to obtain the next processed video frame in the first video data and update it as the first video frame (i.e., update the first video frame), and again identify the fourth pixel coordinate of the speaker in the first video frame, and the coordinate count is incremented by 1. Then, the coordinate distance between the newly determined fourth pixel coordinate and the first pixel coordinate is calculated again. Repeat this process. If the current coordinate count reaches the count threshold (the specific value can be set according to the actual situation), and the coordinate distance between each fourth pixel coordinate and the first pixel coordinate calculated at the current time is greater than the distance threshold, it is determined that the position of the speaker has changed and the change has lasted for a period of time, and therefore it is determined that the speaker has truly changed. At this time, return to perform the operation of identifying the first pixel coordinate of the speaker and the first face size in the first video frame of the first video data currently being processed to re-determine the shooting control parameter suitable for the current speaker, and control the second camera to use the shooting control parameter to obtain the video data shot by the second camera and send it to the virtual camera driver. Optionally, after obtaining the video data shot by the second camera, it can be determined whether the pixel coordinate and face size of the speaker in the video frame of the video data meet the preset shooting condition, and the shooting control parameter can be fine-tuned when the preset shooting condition is not met.
[0260] It should be noted that during the process of continuously incrementing the coordinate count, the first pixel coordinate used to calculate the coordinate distance with each fourth pixel coordinate is the same pixel coordinate.
[0261] On this basis, if the coordinate count plus 1 does not reach the count threshold, and the distance between the newly determined fourth pixel coordinate (which can be considered as the fourth pixel coordinate corresponding to the current coordinate count) and the first pixel coordinate is not greater than the distance threshold, it can be considered that the position of the speaker may have a temporary change, which can be ignored. Therefore, it is determined that the speaker has not changed, and the shooting control parameter used by the second camera does not need to be changed. The intelligent device can continue to obtain second video data and send it to the virtual camera driver program. Optionally, the newly determined fourth pixel coordinate is updated to the first pixel coordinate. It can be understood that this process can also be considered as a de-bouncing filter, that is, the fourth pixel coordinate that is greater than the distance threshold when the count threshold is not reached is filtered out and not responded to.
[0262] It can be understood that the reason for the temporary change of the position of the speaker can be that the intelligent device has a temporary error when identifying the speaker, causing the position to change, or that the speaker has a short speech during the speech process.
[0263] It can be understood that the foregoing process of determining whether the speaker has changed can be implemented by a local AI algorithm module.
[0264] The foregoing process of detecting whether the speaker has changed is continuously performed to achieve tracking and shooting of the speaker.
[0265] As described above, the fourth pixel coordinate of the speaker and the first pixel coordinate can be used to determine whether the speaker has changed, and when the speaker has not changed, the second camera continues to use the current shooting control parameter, and the virtual camera driver program uses the video data shot by the second camera, to ensure that the enlarged close-up shot of the speaker is continuously shot. When the fourth pixel coordinate of the speaker and the first pixel coordinate are used to determine that the speaker has changed, the shooting control parameter suitable for the current speaker is re-determined, and the second camera uses the re-determined shooting control parameter to shoot, to ensure that the second camera can shoot the enlarged close-up shot of the changed speaker.
[0266] It can be understood that the foregoing shooting control method is not only suitable for shooting the speaker, but also suitable for shooting any object in the shooting space. In the implementation process, the face recognition can be modified to object recognition.
[0267] An embodiment of the present application further provides a photographing control device, which is applied to a smart device, and the smart device comprises a first camera, a display screen, a virtual camera driver, and a target application program. The smart device is connected with an external second camera, and the second camera is a rotatable camera. FIG. 8 is a structural schematic diagram of the photographing control device according to an embodiment of the present application. As shown in FIG. 8, the photographing control device comprises a first video acquisition unit 701, a first data identification unit 702, a parameter determination unit 703, a parameter use unit 704, a second video acquisition unit 705, and a first video sending unit 706.
[0268] The first video acquisition unit 701 is configured to acquire first video data photographed by the first camera, and the first camera is configured to photograph a first photographing space. The first data identification unit 702 is configured to identify a first pixel coordinate of a speaker in a first video picture currently processed in the first video data and a first face size of the speaker, and the speaker is located in the first photographing space. The parameter determination unit 703 is configured to determine a first photographing control parameter of the second camera according to the first pixel coordinate and the first face size. The parameter use unit 704 is configured to control the second camera to use the first photographing control parameter for photographing, so that the second camera realizes aligned photographing of the speaker under the first photographing control parameter. The second video acquisition unit 705 is configured to acquire second video data photographed by the second camera. The first video sending unit 706 is configured to obtain the second video data by the virtual camera driver, and send the second video data to the target application program, so that the target application program can use the second video data on the display screen.
[0269] In an embodiment of the present application, the first photographing control parameter comprises a first Pan parameter, a first Tilt parameter, and a first Zoom parameter. The parameter determination unit 703 comprises a first Pan parameter determination subunit, a first Tilt parameter determination subunit, and a first Zoom parameter determination subunit. The first Pan parameter determination subunit is configured to determine the first Pan parameter corresponding to the first pixel coordinate according to a first corresponding relationship, and the first corresponding relationship is a corresponding relationship between a pixel coordinate in the video data photographed by the first camera and a Pan parameter of the second camera. The first Tilt parameter determination subunit is configured to determine the first Tilt parameter corresponding to the first pixel coordinate according to a second corresponding relationship, and the second corresponding relationship is a corresponding relationship between a pixel coordinate in the video data photographed by the first camera and a Tilt parameter of the second camera. The first Zoom parameter determination subunit is configured to determine the first Zoom parameter corresponding to the first face size according to a third corresponding relationship, and the third corresponding relationship is a corresponding relationship between a face size in the video data photographed by the first camera and a Zoom parameter of the second camera.
[0270] In an embodiment of the present application, the photographing control device further comprises: a first correspondence determining unit, configured to, before obtaining the first video data captured by the first camera, establish a first correspondence between pixel coordinates in the video data captured by the first camera and a Pan parameter of the second camera, and establish a second correspondence between pixel coordinates in the video data captured by the first camera and a Tilt parameter of the second camera; and a second correspondence determining unit, configured to establish a third correspondence between a face size in the video data captured by the first camera and a Zoom parameter of the second camera.
[0271] In an embodiment of the present application, the first correspondence determining unit comprises: a third video obtaining sub-unit, configured to obtain a second video frame based on the second video data captured by the first camera, the first camera capturing a static second photographing space; a frame dividing sub-unit, configured to divide the second video frame into a plurality of first video sub-frames, and determine a second pixel coordinate of a center point of each first video sub-frame in the second video frame; a first parameter obtaining sub-unit, configured to obtain a second Pan parameter and a second Tilt parameter corresponding to each first video sub-frame, the similarity between a third video frame captured by the second camera using the second Pan parameter, the second Tilt parameter and a second Zoom parameter and the corresponding first video sub-frame reaching a similarity threshold, the second Zoom parameter being a fixed parameter when the second camera captures the second photographing space; and a first relationship determining sub-unit, configured to determine, according to the second pixel coordinate and the second Pan parameter corresponding to each first video sub-frame, the first correspondence between pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera, and determine, according to the second pixel coordinate and the second Tilt parameter corresponding to each first video sub-frame, the second correspondence between pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera.
[0272] In one embodiment of the present application, the first correspondence determining unit comprises: a seventh video obtaining subunit, configured to obtain an eighth video picture based on the first camera and a ninth video picture based on the second camera at the same shooting time; a third parameter obtaining subunit, configured to obtain a fifth pixel coordinate of a center point of a fifth sub-video picture in the eighth video picture when a similarity between the ninth video picture and the fifth sub-video picture reaches a similarity threshold, and obtain a fourth Pan parameter and a fourth Tilt parameter used by the second camera when shooting the ninth video picture, the second camera using a fixed fourth Zoom parameter, the fifth sub-video picture being one of the sub-video pictures obtained by dividing the eighth video picture; and a third relationship determining subunit, configured to determine a first correspondence between a pixel coordinate in the video data shot by the first camera and a Pan parameter of the second camera according to a plurality of fifth pixel coordinates and a plurality of corresponding fourth Pan parameters, and determine a second correspondence between a pixel coordinate in the video data shot by the first camera and a Tilt parameter of the second camera according to a plurality of fifth pixel coordinates and a plurality of corresponding fourth Tilt parameters, the plurality of fifth pixel coordinates corresponding to a plurality of fifth sub-video pictures obtained at a plurality of shooting times, and the pixel coordinate ranges of the fifth sub-video pictures being different.
[0273] In one embodiment of the present application, the second correspondence determining unit comprises: a fourth video obtaining subunit, configured to sequentially obtain a plurality of fourth video pictures based on the first camera, and determine a second face size of the target face in each of the fourth video pictures, the first camera shooting a third shooting space, the target face being in the third shooting space, and the second face sizes corresponding to the fourth video pictures being different; a second parameter obtaining subunit, configured to, when obtaining each of the fourth video pictures, obtain a third Zoom parameter corresponding to the fourth video picture, and the third face size of the target face in a fifth video picture shot by the second camera using the third Zoom parameter, a third Pan parameter and a third Tilt parameter satisfying a preset face size; and a second relationship determining subunit, configured to determine a third correspondence between a face size in the video data shot by the first camera and a Zoom parameter of the second camera according to the second face size corresponding to each of the fourth video pictures and the third Zoom parameter.
[0274] In an embodiment of the present application, the first data identifying unit 702 comprises: a speaker identifying subunit, configured to identify a speaker in a first video picture currently processed in the first video data; a speaker picture acquiring subunit, configured to acquire a second sub-video picture of a face of the speaker in the first video picture; and a first data determining subunit, configured to determine a pixel coordinate of a center point of the second sub-video picture in the first video picture as a first pixel coordinate, and determine a size of the second sub-video picture as a first face size.
[0275] In an embodiment of the present application, the speaker identifying subunit comprises: a lip picture identifying grandson unit, configured to perform face recognition on the first video picture currently processed in the first video data and a sixth video picture of a plurality of continuous frames before the first video picture in the first video data, and acquire a lip picture on each recognized face; and a speaker determining grandson unit, configured to determine the speaker in the first video picture according to a motion amplitude of the lip in the lip picture, wherein the motion amplitude of the lip of the speaker is the largest.
[0276] In an embodiment of the present application, the intelligent device further comprises a microphone array; and the lip picture identifying grandson unit is specifically configured to: acquire DOA information of the speaker based on the microphone array; acquire a third sub-video picture in the first video picture currently processed in the first video data and a fourth sub-video picture in the sixth video picture of a plurality of continuous frames before the first video picture according to the DOA information; and perform face recognition on the third sub-video picture and the fourth sub-video picture, and acquire a lip picture on each recognized face.
[0277] In an embodiment of the present application, the shooting control apparatus further comprises: a second data determining unit, configured to acquire a third pixel coordinate of the speaker in a seventh video picture currently processed in the second video data and a fourth face size after the second data determining unit acquires the second video data shot by the second camera; and a condition determining unit, configured to perform an operation of obtaining the second video data by the virtual camera driver when the third pixel coordinate and the fourth face size satisfy a preset shooting condition, wherein the preset shooting condition is that the third pixel coordinate is located in a set pixel coordinate region and the fourth face size satisfies a preset face size.
[0278] In an embodiment of the present application, the photographing control apparatus further comprises: a parameter adjustment unit, configured to adjust the first photographing control parameter according to the third pixel coordinate and the fourth face size when the third pixel coordinate and the fourth face size do not satisfy the preset photographing condition, to obtain a second photographing control parameter, the second camera implements the aligned photographing of the speaker under the second photographing control parameter, and the video picture obtained by the second camera satisfies the preset photographing condition; a sixth video acquisition unit, configured to acquire third video data photographed by the second camera under the second photographing control parameter; and a second video sending unit, configured to obtain the third video data by the virtual camera driver program, and send the third video data to the target application program, so that the target application program can use the third video data on the display screen.
[0279] In an embodiment of the present application, the photographing control apparatus further comprises: a picture updating unit, configured to update the first video picture currently processed when the second video data photographed by the second camera is acquired; a second data recognition unit, configured to recognize a fourth pixel coordinate of the speaker in the updated first video picture; and a first continuous photographing unit, configured to determine that the speaker does not change when the coordinate distance between the fourth pixel coordinate and the first pixel coordinate is not greater than the distance threshold, and return to perform the operation of acquiring the second video data photographed by the second camera, to implement the continuous photographing of the speaker.
[0280] In an embodiment of the present application, the photographing control apparatus further comprises: a counting unit, configured to continue to update the first video picture currently processed and continue to recognize the fourth pixel coordinate of the speaker in the updated first video picture when the coordinate distance between the fourth pixel coordinate and the first pixel coordinate is greater than the distance threshold, and start coordinate counting, the coordinate counting is increased by 1 each time the fourth pixel coordinate is determined; and a speaker change determination unit, configured to determine that the speaker changes when the current coordinate count reaches the count threshold, and the coordinate distance between the plurality of fourth pixel coordinates determined in the current coordinate count and the first pixel coordinate is greater than the distance threshold, and return to perform the operation of recognizing the first pixel coordinate and the first face size of the speaker in the first video picture currently processed in the first video data, to enable the second camera to implement the aligned photographing of the changed speaker.
[0281] In an embodiment of the present application, the photographing control apparatus further comprises: a second continuous photographing unit, configured to determine that the speaker does not change when the current coordinate count does not reach the count threshold and the coordinate distance between the newly determined fourth pixel coordinate and the first pixel coordinate is not greater than the distance threshold, and return to perform the operation of acquiring the second video data photographed by the second camera, to implement the continuous photographing of the speaker.
[0282] An embodiment of the present application further provides a smart device. Referring to FIG. 1, the smart device comprises a first camera 13, a display screen 15, one or more processors 11 (one is taken as an example in FIG. 1) and a memory 12. The smart device is connected with an external second camera (not shown in the figure), and the second camera is a rotatable camera. The first camera is configured to take a picture according to an indication of the processor; the display screen is configured to display according to an indication of the processor; the memory is configured to store one or more programs; and when the one or more programs are executed by the one or more processors, the one or more processors implement the picture taking control method according to any one of the foregoing embodiments; and the second camera is configured to take a picture using a picture taking control parameter determined by the processor.
[0283] The smart device further comprises a microphone array 14.
[0284] Details of the components can be referred to the foregoing description.
[0285] The smart device is configured to execute any picture taking control method, and has corresponding functions and advantages. Details not described herein can be referred to the foregoing description of the picture taking control method.
[0286] An embodiment of the present application further provides a picture taking control system. FIG. 9 is a structural schematic diagram of a picture taking control system according to an embodiment of the present application. Referring to FIG. 9, the picture taking control system comprises a smart device 10 and a second camera 20. Details of the smart device 10 and the second camera 20 can be referred to the foregoing description, and the smart device 10 and the second camera 20 have corresponding functions and advantages. Details not described herein can be referred to the foregoing description of the picture taking control method.
[0287] An embodiment of the present application further provides a storage medium comprising computer executable instructions. When the computer executable instructions are executed by a processor, the computer executable instructions are configured to perform the picture taking control method according to any embodiment of the present application, and have corresponding functions and advantages.
[0288] Those skilled in the art should understand that embodiments of the present application can be provided as a method, a system or a computer program product.
[0289] Accordingly, embodiments of the present application can be embodied in the form of complete hardware embodiments, complete software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present application can take the form of a computer program product on one or more computer-readable storage media (including, but not limited to, disk memory, CD-ROMs, optical storage devices, etc.) embodying computer readable program code. Embodiments of the present application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing system or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustration and / or block diagram block or blocks. These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart illustration and / or block diagram block or blocks. The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart illustration and / or block diagram block or blocks.
[0290] Computer readable media includes permanent and non-permanent, removable and non-removable media implemented in any method or technology for storage of information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disc read only memory (CD-ROM), digital versatile discs (DVDs) or other optical storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to computing devices. According to the definition herein, computer readable media does not include transitory media, such as modulated data signals and carrier waves.
[0291] It is also to be noted that the terms "comprising", "including", and any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises a... " does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0292] Note that the above merely describes preferred embodiments of the present application and the applied technical principles. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, reconfigurations, and substitutions can be made without departing from the scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the claims.
Claims
A shooting control method, applied to smart devices, wherein, The smart device includes a first camera, a display screen, a virtual camera driver, and a target application. The smart device is connected to an external second camera, which is a rotatable camera. The shooting control method includes: Acquire first video data captured by the first camera, wherein the first camera is used to capture a first shooting space; Identify the first pixel coordinates and first face size of the speaker within the currently processed first video frame in the first video data, wherein the speaker is located in the first shooting space; The first shooting control parameters of the second camera are determined based on the first pixel coordinates and the first face size; Control the second camera to take pictures using the first shooting control parameters, so that the second camera can achieve aligned shooting of the speaker under the first shooting control parameters; Acquire the second video data captured by the second camera; The second video data is obtained by the virtual camera driver and sent to the target application so that the target application can use the second video data on the display screen. According to the shooting control method of claim 1, wherein, The first shooting control parameters include the first Pan parameter, the first Tilt parameter, and the first Zoom parameter; The step of determining the first shooting control parameters of the second camera based on the first pixel coordinates and the first face size specifically includes: The first Pan parameter corresponding to the first pixel coordinate is determined according to the first correspondence relationship, wherein the first correspondence relationship is the correspondence between the pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera. The first Tilt parameter corresponding to the first pixel coordinate is determined according to the second correspondence relationship, wherein the second correspondence relationship is the correspondence between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera; The first Zoom parameter corresponding to the first face size is determined according to the third correspondence relationship, wherein the third correspondence relationship is the correspondence between the face size in the video data captured by the first camera and the Zoom parameter of the second camera. According to the shooting control method of claim 2, wherein, Before acquiring the first video data captured by the first camera, the method further includes: Establish a first correspondence between the pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera, and establish a second correspondence between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera; Establish a third correspondence between the face size in the video data captured by the first camera and the Zoom parameter of the second camera. According to claim 3, the shooting control method, wherein, The establishment of a first correspondence between the pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera, and the establishment of a second correspondence between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera, specifically include: Acquire a second video frame based on the first camera, wherein the first camera captures a static second shooting space; The second video frame is divided into multiple first sub-video frames, and the second pixel coordinates of the center point of each first sub-video frame in the second video frame are determined. The second Pan parameter and the second Tilt parameter corresponding to each first sub-video frame are obtained. The similarity between the third video frame captured by the second camera using the second Pan parameter, the second Tilt parameter and the second Zoom parameter and the corresponding first sub-video frame reaches the similarity threshold. When the second camera captures the second shooting space, the second Zoom parameter is a fixed parameter. A first correspondence between the pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera is determined based on the second pixel coordinates and the second Pan parameter corresponding to each first sub-video frame. A second correspondence between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera is determined based on the second pixel coordinates and the second Tilt parameter corresponding to each first sub-video frame. According to claim 3, the shooting control method, wherein, The establishment of a first correspondence between the pixel coordinates in the video data captured by the first camera and the Pan parameter of the second camera, and the establishment of a second correspondence between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera, specifically include: Acquire the eighth video frame captured by the first camera and the ninth video frame captured by the second camera at the same shooting time; When the similarity between the ninth video frame and the fifth sub-video frame in the eighth video frame reaches a similarity threshold, the fifth pixel coordinate of the center point of the fifth sub-video frame in the eighth video frame is obtained, as well as the fourth Pan parameter and the fourth Tilt parameter used by the second camera when capturing the ninth video frame. The second camera uses a fixed fourth Zoom parameter. The fifth sub-video frame is one of the sub-video frames after the eighth video frame is segmented. The pixel coordinates in the video data captured by the first camera and the second camera are determined based on multiple fifth pixel coordinates and corresponding multiple fourth Pan parameters. The first correspondence of the Pan parameter of the head is determined by the second correspondence between the pixel coordinates in the video data captured by the first camera and the Tilt parameter of the second camera, based on multiple fifth pixel coordinates and the corresponding multiple fourth Tilt parameters. The multiple fifth pixel coordinates correspond to the fifth sub-video frames obtained at multiple shooting times, and the pixel coordinate range of each fifth sub-video frame is different. According to claim 3, the shooting control method, wherein, The establishment of a third correspondence between the face size in the video data captured by the first camera and the Zoom parameter of the second camera specifically includes: The system sequentially acquires multiple frames of fourth video images captured by the first camera and determines the second face size of the target face in each frame of the fourth video image. The first camera captures images in a third shooting space, which includes the target face. The second face size corresponding to each frame of the fourth video image is different. Each time a frame of the fourth video image is acquired, the third Zoom parameter corresponding to the fourth video image is also acquired. The third face size of the target face in the fifth video image captured by the second camera using the third Zoom parameter, the third Pan parameter and the third Tilt parameter satisfies the preset face size. Based on the second face size and third Zoom parameter corresponding to each frame of the fourth video image, a third correspondence between the face size in the video data captured by the first camera and the Zoom parameter of the second camera is determined. According to the shooting control method of claim 1, wherein, The identification of the first pixel coordinates and first face size of the speaker within the currently processed first video frame in the first video data specifically includes: Identify the speaker within the currently processed first video frame in the first video data; Obtain a second sub-video frame containing the speaker's face within the first video frame; The pixel coordinates of the center point of the second sub-video frame within the first video frame are determined and used as the first pixel coordinates. The size of the second sub-video frame is determined and used as the first face size. According to the shooting control method of claim 7, wherein, The step of identifying the speaker within the currently processed first video frame in the first video data specifically includes: Face recognition is performed on the first video frame currently being processed in the first video data and the sixth video frame that is several consecutive frames before the first video frame in the first video data, and the lip images of each recognized face are obtained. Based on the range of lip movement in the lip image, the speaker in the first video frame is identified, and the speaker has the largest range of lip movement. According to the shooting control method of claim 8, wherein, The smart device also includes a microphone array; The step of performing face recognition on the currently processed first video frame in the first video data and the sixth video frame that appears multiple consecutive frames before the first video frame in the first video data, and obtaining the lip image of each recognized face, specifically includes: The speaker's DOA information is obtained based on the microphone array; Based on the DOA information, a third sub-video frame is extracted from the first video frame currently being processed in the first video data, and a fourth sub-video frame is extracted from the sixth video frame that is several consecutive frames before the first video frame. Face recognition is performed on the third and fourth sub-video frames, and the lips of each recognized face are obtained. According to the shooting control method of claim 1, wherein, After acquiring the second video data captured by the second camera, the process further includes: Determine the third pixel coordinates and fourth face size of the speaker within the currently processed seventh video frame in the second video data; When the third pixel coordinates and the fourth face size meet the preset shooting conditions, the operation of obtaining the second video data by the virtual camera driver is executed. The preset shooting conditions are that the third pixel coordinates are located within the set pixel coordinate area and the fourth face size meets the preset face size. According to the shooting control method of claim 10, wherein, Also includes: When the third pixel coordinates and the fourth face size do not meet the preset shooting conditions, the first shooting control parameters are adjusted according to the third pixel coordinates and the fourth face size to obtain the second shooting control parameters. The second camera achieves aligned shooting of the speaker under the second shooting control parameters and the video image captured by the second camera meets the preset shooting conditions. Obtain third video data captured by the second camera using the second shooting control parameters; The third video data is obtained by the virtual camera driver and sent to the target application so that the target application can use the third video data on the display screen. According to the shooting control method of claim 1, wherein, When acquiring the second video data captured by the second camera, the method further includes: Update the first video frame currently being processed; Identify the fourth pixel coordinate of the speaker within the updated first video frame; When the distance between the fourth pixel coordinate and the first pixel coordinate is not greater than a distance threshold, it is determined that the speaker has not changed and the operation of acquiring the second video data captured by the second camera is returned to achieve continuous shooting of the speaker. According to the shooting control method of claim 12, wherein, Also includes: When the coordinate distance between the fourth pixel coordinate and the first pixel coordinate is greater than a distance threshold, the currently processed first video frame continues to be updated and the fourth pixel coordinate of the speaker in the updated first video frame continues to be identified, and coordinate counting begins. The coordinate count is incremented by 1 each time a fourth pixel coordinate is determined. If the current coordinate count reaches the counting threshold, and the coordinate distance between the multiple fourth pixel coordinates determined within the current coordinate count and the first pixel coordinate is greater than the distance threshold, it is determined that the speaker has changed and the operation of recognizing the first pixel coordinate and first face size of the speaker in the first video frame currently being processed in the first video data is returned, so that the second camera is aimed at and captures the changed speaker. According to the shooting control method of claim 13, wherein, Also includes: If the current coordinate count has not reached the counting threshold and the coordinate distance between the latest determined fourth pixel coordinate and the first pixel coordinate is not greater than the distance threshold, it is determined that the speaker has not changed and the operation of obtaining the second video data captured by the second camera is returned to achieve continuous shooting of the speaker. A shooting control device, applied to smart devices, wherein... The smart device includes a first camera, a display screen, a virtual camera driver, and a target application. The smart device is connected to an external second camera, which is a rotatable camera. The shooting control device includes: The first video acquisition unit is used to acquire the first video data captured by the first camera, and the first camera is used to capture the first shooting space; The first data recognition unit is used to identify the first pixel coordinates and the first face size of the speaker in the first video frame currently being processed in the first video data, wherein the speaker is located in the first shooting space; The parameter determination unit is used to determine the first shooting control parameters of the second camera based on the first pixel coordinates and the first face size; The parameter usage unit is used to control the second camera to shoot using the first shooting control parameters, so that the second camera can achieve aligned shooting of the speaker under the first shooting control parameters; The second video acquisition unit is used to acquire the second video data captured by the second camera. The first video sending unit is configured to obtain the second video data from the virtual camera driver and send the second video data to the target application so that the target application can use the second video data on the display screen. A smart device, wherein, include: A first camera, a display screen, one or more processors, and memory; The smart device is connected to an external second camera, which is a rotatable camera; The first camera is used to take pictures according to the instructions of the processor; The display screen is used to display information according to the instructions of the processor; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the shooting control method as described in any one of claims 1-14; The second camera is used to take pictures using the shooting control parameters determined by the processor. A shooting control system, wherein, Includes the smart device and the second camera as described in claim 16. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by the processor, it implements the shooting control method as described in any one of claims 1-14.
Citation Information
Patent Citations
Person head shooting method, system and server
CN103167270A
Image processing method, apparatus and system
CN109492506A
Implementation method and device of face tracking camera shooting, computer equipment and storage medium
CN112839165A
Method and system for tracking and shooting spokesman by using binocular camera
CN117544855A
Automatic Switching Between Dynamic and Preset Camera Views in a Video Conference Endpoint
US20160134838A1