An automatic program director device
By designing an automatic guide device, using machine vision analysis and automatic lens guide module, automatic calculation and real-time adjustment of live view angles are realized, solving the problem that the guide angle cannot be adjusted in time in the existing technology, and improving the fluency and viewing of live broadcasts.
Patent Information
- Application Number
- CN202310002204.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-01-03
AI Technical Summary
In the prior art, the director angle cannot be adjusted in time and has low flexibility, which affects the fluency and viewing of live broadcasts.
An automatic guide device is designed, including the main camera, a gimbal close-up camera, a machine vision analysis module, an automatic lens guide module, a lens angle adjustment module, a video encoding module, a network output module and a video output interface. The device analyzes the anchor picture and item display screen through the machine vision analysis module, automatically calculates the live view angle, and realizes real-time angle adjustment through the lens angle adjustment module.
It achieves the goal of not losing in live broadcasts, provides better viewing and broader vision than single-view live broadcast devices, improves the flexibility of directors, and reduces the need for manual adjustments.
Smart Images

Figure CN116132794B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video directing, and in particular to an automatic directing device. Background Art
[0002] Currently, most normal personal live broadcasts rely on a single camera. Ordinary single cameras can only capture the scene that the current lens is aimed at, and the angle of the shooting needs to be adjusted manually in advance. A further smart tracking single camera can track the target within a limited range and adjust the angle, but the live broadcast of this device is prone to losing the target. When the target exceeds the limited angle, it is often impossible to continue tracking the target, and it is necessary to wait for manual adjustment of the lens angle.
[0003] In addition, single-camera live broadcast devices, such as mobile phones, are also unable to adjust the angle according to events outside the viewing angle. These defects greatly affect the smoothness and viewing quality of a live broadcast, making it impossible for online viewers to learn about the situation on the scene in a timely manner, and causing the live broadcaster to be distracted and unable to focus on the live broadcast content. Therefore, how to improve the flexibility of the director is a technical problem that technicians in this field urgently need to solve. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide an automatic directing device, which solves the technical problems in the prior art that the directing angle cannot be adjusted in time and has low flexibility.
[0005] In order to solve the above technical problems, the present invention provides an automatic broadcasting device, comprising:
[0006] Main camera 1, pan / tilt close-up camera 2, machine vision analysis module 3, automatic lens director module 4, lens angle adjustment module 5, video encoding module 6, network output module 7 and video output interface 8;
[0007] The main camera 1 and the pan-tilt close-up camera 2 are both connected to the machine vision analysis module 3; the main camera 1 transmits the captured anchor image to the machine vision analysis module 3; the pan-tilt close-up camera 2 transmits the captured item display image to the machine vision analysis module 3; the machine analysis module 3 analyzes the received anchor image and item display image to obtain a semantic analysis result;
[0008] The input end of the automatic shot directing module 4 is connected to the machine vision analysis module 3 to receive the semantic analysis result and calculate the live viewing angle;
[0009] The lens angle adjustment module 5 is connected to the output end of the automatic lens directing module 4 to obtain target video information according to the live viewing angle;
[0010] The video encoding module 6 is connected to the output end of the lens angle adjustment module 5, receives the target video information and encodes it to obtain an encoded video;
[0011] The network output module 7 is connected to the output end of the video encoding module 6, receives the encoded video, and transmits it to the video output interface 8. The input end of the video output interface 8 is connected to the network output module 7.
[0012] Optionally, the automatic director device further includes:
[0013] The output end of the external video access module 9 for receiving electronic video information is connected to the machine vision analysis module 3.
[0014] Optionally, the automatic director device further includes:
[0015] A left pan-tilt camera 10, the left pan-tilt camera 10 is connected to the machine vision analysis module 3, the left pan-tilt camera (10) is arranged on the left side of the host, and transmits the captured picture on the left side of the host to the machine vision analysis module 3.
[0016] Optionally, the automatic director device further includes:
[0017] A right pan-tilt camera 11, the right pan-tilt camera 11 is connected to the machine vision analysis module 3, the left pan-tilt camera is arranged on the right side of the host, and transmits the captured picture on the right side of the host to the machine vision analysis module.
[0018] Optionally, the automatic director device further includes:
[0019] A first lifting bracket 12; the first lifting bracket 12 is connected to the left pan-tilt camera 10.
[0020] Optionally, the automatic director device further includes:
[0021] A second lifting bracket 13; the second lifting bracket 13 is connected to the right pan-tilt camera 11.
[0022] The present invention also provides an automatic directing method, which is applied to the above automatic director device and includes:
[0023] Obtain the host picture and the item display picture;
[0024] Call the vision analysis model to analyze the host picture and the item display picture respectively to obtain the host semantic analysis result and the item display semantic analysis result;
[0025] Call the camera director calculation model to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result to determine the live broadcast perspective.
[0026] Optionally, before the visual analysis model is called to analyze the host screen and the item display screen respectively, it further includes:
[0027] Obtain the electronic video screen collected by the external video access module, the left side screen of the host collected by the left pan-tilt camera, and the right side screen of the host collected by the right pan-tilt camera;
[0028] Correspondingly, when the visual analysis model analyzes the host screen and the item display screen respectively to obtain the host semantic analysis result and the item display semantic analysis result, it includes:
[0029] Call the visual analysis model to analyze the host screen, the item display screen,
[0030] the left side screen of the host and the right side screen of the host respectively, to obtain the host semantic analysis result, the item display semantic analysis result, the electronic video semantic analysis result, the left side semantic analysis result of the host and the right side semantic analysis result of the host.
[0031] Optionally, when the visual analysis model analyzes the host screen and the item display screen respectively to obtain the host semantic analysis result and the item display semantic analysis result, it includes:
[0032] Call the neural network model in the visual analysis model to calculate the head pose of the host in the host screen, and use the head pose of the host as the host semantic analysis result;
[0033] Call the neural network gesture recognition model in the visual analysis model to calculate the item display gesture in the item display screen, and use the item display gesture as the item display semantic analysis result.
[0034] Optionally, when the camera director calculation model is called to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result, it includes:
[0035] Call the corresponding camera director calculation model for the live broadcast scenario to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result.
[0036] It can be seen that the automatic camera director device provided by the present invention includes a main camera, a pan-tilt close-up camera, a machine vision analysis module, an automatic camera director module, a camera angle adjustment module, a video encoding module, a network output module, and a video output interface. The main camera and the pan-tilt close-up camera are both connected to the machine vision analysis module. The main camera transmits the captured host live video to the machine vision analysis module. The pan-tilt close-up camera transmits the captured item display video to the machine vision analysis module. The machine analysis module analyzes the received host live video and the item display video to obtain a semantic analysis result. The input end of the automatic camera director module is connected to the machine vision analysis module, receives the semantic analysis result, and calculates the live broadcast perspective. The camera angle adjustment module is connected to the output end of the automatic camera director module, and obtains target video information according to the live broadcast perspective. The video encoding module is connected to the output end of the camera angle adjustment module, receives the target video information and encodes it to obtain an encoded video. The network output module is connected to the output end of the video encoding module, receives the encoded video, and transmits it to the video output interface. The input end of the video output interface is connected to the network output module. Compared with traditional single-camera live broadcast devices, since the present invention is provided with a main camera, a pan-tilt close-up camera, and a machine vision analysis module, an automatic camera director module, and a camera angle adjustment module for determining the camera perspective, the present invention will not lose the target during the live broadcast, and brings a better viewing experience and a wider field of view than single-perspective live broadcast devices. Compared with traditional intelligent tracking cameras, the semantic analysis of the machine vision analysis module of this device is rich, and it can perform live broadcasts according to changes in the scene and changes in the host's line of sight, without the need for manual switching and angle adjustment. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0038] Figure 1 It is a schematic structural diagram of an automatic camera director device provided by an embodiment of the present invention;
[0039] Figure 2 It is a schematic structural diagram of a camera installation method provided by an embodiment of the present invention;
[0040] Figure 3 It is a schematic structural diagram of an automatic camera director device with a lifting bracket provided by an embodiment of the present invention;
[0041] Figure 4 Schematic diagram of the structure of an automatic camera director device in the on state provided by an embodiment of the present invention;
[0042] Figure 5 Schematic diagram of the structure of an automatic camera director device in the off state provided by an embodiment of the present invention;
[0043] Figure 6 Flowchart of an automatic camera director method provided by an embodiment of the present invention;
[0044] Figure 7 Flow example diagram of an automatic camera director method provided by an embodiment of the present invention;
[0045] Appendix Figures 1-5 In the following, the reference numerals are explained as follows:
[0046] 1 - Main camera;
[0047] 2 - PTZ close - up camera;
[0048] 3 - Machine vision analysis module;
[0049] 4 - Automatic lens director module;
[0050] 5 - Lens angle adjustment module;
[0051] 6 - Video encoding module;
[0052] 7 - Network output module;
[0053] 8 - Video output interface;
[0054] 9 - External video access module;
[0055] 10 - Left PTZ camera;
[0056] 11 - Right PTZ camera;
[0057] 12 - First lifting bracket;
[0058] 13 - Second lifting bracket. Detailed implementation manners
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0060] Embodiment 1:
[0061] Please refer to Figure 1 , Figure 1 , which is a schematic structural diagram of an automatic device provided by an embodiment of the present invention. The device may include:
[0062] Main camera 1, pan-tilt close-up camera 2, machine vision analysis module 3, automatic lens director module 4, lens angle adjustment module 5, video encoding module 6, network output module 7, and video output interface 8;
[0063] Both the main camera 9 and the pan-tilt close-up camera 10 are connected to the machine vision analysis module 3; the main camera 1 transmits the captured host picture to the machine vision analysis module 3; the pan-tilt close-up camera 2 transmits the captured item display picture to the machine vision analysis module 3; the machine analysis module 3 analyzes the received host picture and item display picture to obtain a semantic analysis result;
[0064] The input end of the automatic lens director module 4 is connected to the machine vision analysis module 3, receives the semantic analysis result, and calculates the live broadcast perspective;
[0065] The lens angle adjustment module 5 is connected to the output end of the automatic lens director module 4, and obtains target video information according to the live broadcast perspective;
[0066] The video encoding module 6 is connected to the output end of the lens angle adjustment module 5, receives the target video information and encodes it to obtain an encoded video;
[0067] The network output module 7 is connected to the output end of the video encoding module 6, receives the encoded video, and transmits it to the video output interface 8. The input end of the video output interface 8 is connected to the network output module 7.
[0068] This embodiment does not limit the specific type of the main camera 1, as long as it can capture the picture of the host person. For example, the main camera 1 can be a pan-tilt camera, or the main camera can be a non-pan-tilt camera. The pan-tilt close-up camera 2 in this embodiment is mainly used to capture the items that the host person wants to display. The pan-tilt close-up camera 2 is a camera with a pan-tilt, which is equipped with a device for carrying the camera to rotate in two directions, horizontal and vertical. Installing the camera on the pan-tilt can enable the camera to take pictures from multiple angles.
[0069] In this embodiment, the machine vision analysis module 3 is equivalent to a processor, which is used to receive the video images collected by the main camera 1 and the pan-tilt close-up camera 2 and perform analysis. The automatic camera director module 4 in this embodiment is mainly used to receive the information analyzed by the machine vision analysis module 3, so as to calculate a suitable live broadcast perspective. Since the traditional rule-based camera director algorithm has weak flexibility and a single applicable scenario, and the present invention is a portable live broadcast device, it will be used in different scenarios and dynamically calculate the most suitable live broadcast perspective at each moment. This embodiment does not limit the specific method of calculating the live broadcast perspective in the automatic camera director module, as long as the live broadcast perspective can be calculated according to the semantic analysis result. It can be understood that different weight thresholds can be set for the semantic analysis results corresponding to the main camera 1 and the pan-tilt close-up camera 2 according to different live broadcast scenarios, so as to perform weighted calculation of the most suitable live broadcast perspective. The network output module 7 in this embodiment allows the encoded video to be played on a local display, or the encoded video stream to be uploaded to popular live broadcast platforms such as Douyu and Huya for playback.
[0070] Based on the above embodiment, the present invention uses the main camera 1, the pan-tilt close-up camera 2, the machine vision analysis module 3, the automatic camera director module 4, the lens angle adjustment module 5, the video encoding module 6, the network output module 7 and the video output interface 8 to collect the host picture and the item display picture for analysis and calculate a suitable live broadcast perspective. It can be seen that this embodiment is different from the traditional single-camera live broadcast device. Since the present invention uses the main camera 1 and the pan-tilt close-up camera 2 to separately capture the host picture and the item display picture, and uses the machine vision analysis module 3 and the automatic camera director module 4 in combination to calculate a suitable live broadcast perspective and timely adjust the live broadcast perspective, the target will not be lost during the live broadcast, and it brings a better viewing experience and a wider field of view than the single-perspective live broadcast device. Compared with the traditional intelligent tracking camera, the semantic analysis of this device is rich, and it can perform live broadcasts according to the changes of the scene and the changes of the host's line of sight, without the need for manual switching and angle adjustment.
[0071] Further, in order to improve the applicability of the automatic camera director device to different scenarios, the above automatic camera director device may further include:
[0072] The output end of the external video access module 9 for receiving electronic video information is connected to the machine vision analysis module 3.
[0073] The external video access module 9 in this embodiment can receive external video information. This embodiment does not limit the specific electronic video information. For example, the electronic video information can be a PPT presented by an external laptop desktop signal; or it can be a short video presented by an external laptop desktop signal, etc. In this embodiment, the machine vision analysis module 3 is used to collect the frame changes of the video stream corresponding to the external video access module 9. When the frame change of the video stream is greater than the frame changes of the adjacent several frames, it is regarded as having new content for display, and the corresponding semantic analysis result will be transmitted to the automatic camera director module 4 for final scheduling calculation.
[0074] Based on the above embodiments, since the external video access module 9 is used to obtain external electronic video information in the embodiments of the present invention, the applicable scenarios of the automatic camera director device are more extensive. Compared with the existing method of only obtaining video information through a camera, it can be applied to teaching live broadcasts, report live broadcasts, etc. that require PPT presentations.
[0075] Further, in order to obtain the picture on the left side of the host, the above automatic camera director device may further include:
[0076] A left pan-tilt camera 10, the left pan-tilt camera 10 is connected to the machine vision analysis module 3, the left pan-tilt camera 10 is arranged on the left side of the host, and transmits the collected picture on the left side of the host to the machine vision analysis module 3.
[0077] The left pan-tilt camera 10 in this embodiment is arranged on the left side of the host and transmits the collected picture on the left side of the host to the machine vision analysis module 3. In this embodiment, the machine vision analysis module 3 can calculate the optical flow change of the picture on the left side of the host collected by the left pan-tilt camera 10. When the optical flow change is too large and exceeds the threshold, it indicates that there may be a special event in this perspective. Similarly, the optical flow information will be further transmitted to the automatic camera director module 4 for comprehensive calculation.
[0078] Based on the above embodiments, since there is a dedicated left pan-tilt camera 10 to collect the picture on the left side of the host, the field of view of the automatic camera director device is more comprehensive.
[0079] Further, in order to obtain the picture on the right side of the host, the above automatic camera director device may further include:
[0080] A right pan-tilt camera 11, the right pan-tilt camera 11 is connected to the machine vision analysis module 3, the left pan-tilt camera is arranged on the right side of the host, and transmits the collected picture on the right side of the host to the machine vision analysis module.
[0081] In this embodiment, the right pan-tilt camera 11 is arranged on the left side of the host, and transmits the captured image of the right side of the host to the machine vision analysis module 3. In this embodiment, the machine vision analysis module 3 can calculate the optical flow change of the image of the right side of the host captured by the right pan-tilt camera 11. When the optical flow change is too large and exceeds the threshold, it indicates that there may be a special event in this perspective. Similarly, the optical flow information will be further transmitted to the automatic lens director module 4 for comprehensive calculation.
[0082] Based on the above embodiment, since there is a dedicated right pan-tilt camera 11 for capturing the image of the right side of the host, the field of view of this automatic director device is more comprehensive.
[0083] For ease of understanding, please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a camera installation method provided by an embodiment of the present invention. The cameras for this automatic director may specifically include:
[0084] A main camera 1, a left pan-tilt camera 10, and a right pan-tilt camera 11. The main camera 1, the left pan-tilt camera 10, and the right pan-tilt camera 11 are integrated on the same main board, and are respectively used to capture the image information of the host, the image of the left side of the host, and the image of the right side of the host.
[0085] Further, in order to facilitate the left pan-tilt camera 10 to obtain video images, the above automatic director device may further include:
[0086] A first lifting bracket 12; the first lifting bracket 12 is connected to the left pan-tilt camera 10.
[0087] This embodiment does not limit the specific form of the first lifting bracket 12. It is mainly possible to use the first lifting bracket 12 to adjust the height of the left pan-tilt camera 10. For example, the first lifting bracket 12 can be a screw-type lifting bracket, that is, a lifting bracket composed of a combination of cylinders of different sizes; or the first lifting bracket 12 can be a folding lifting bracket, that is, when adjusting the height of the left pan-tilt camera 10, the line is a broken line.
[0088] Based on the above embodiment, the embodiment of the present invention can adjust the height of the left pan-tilt camera 10 through the first lifting bracket 12, making the left pan-tilt camera 10 more flexible.
[0089] Further, in order to facilitate the left and right cameras 10 to obtain video images, the above automatic director device may further include:
[0090] A second lifting bracket 13; the second lifting bracket 13 is connected to the right pan-tilt camera 11.
[0091] This embodiment does not limit the specific form of the second lifting bracket 13. It is mainly possible to use the second lifting bracket 13 to adjust the height of the right pan-tilt camera 11.
[0092] This embodiment does not limit the specific form of the second lifting bracket 13. It is mainly possible to use the second lifting bracket 13 to adjust the height of the right pan-tilt camera 11. For example, the second lifting bracket 13 can be a screw-type lifting bracket, that is, a lifting bracket composed of a combination of cylinders of different sizes; or the second lifting bracket 13 can be a folding lifting bracket, that is, when adjusting the height of the right pan-tilt camera 11, the line is a broken line.
[0093] Based on the above embodiment, the embodiment of the present invention can adjust the height of the right pan-tilt camera 11 through the second lifting bracket 13, making the right pan-tilt camera 11 more flexible.
[0094] For ease of understanding, please refer to Figure 3 , Figure 3 which is a schematic structural diagram of an automatic camera control device with a lifting bracket provided by an embodiment of the present invention. The automatic camera control device with a lifting bracket may specifically include:
[0095] A left pan-tilt camera 10, a right pan-tilt camera 11, a first lifting bracket 12, and a second lifting bracket 13. Since the heights of the left pan-tilt camera 10 and the right pan-tilt camera 11 can be adjusted, when the automatic camera control device is not in use, all components can be retracted, and when needed, they can be extended. Therefore, an adjustable lifting bracket can also be installed for the pan-tilt close-up camera 2. For details, please refer to Figure 4 and Figure 5 , Figure 4 which is a schematic structural diagram of an automatic camera control device in the on state provided by an embodiment of the present invention. In the on state, the pan-tilt close-up camera 2, the left pan-tilt camera 10, and the right pan-tilt camera 11 can extend to capture video footage; Figure 5 which is a schematic structural diagram of an automatic camera control device in the off state provided by an embodiment of the present invention. In the off state, only the main camera 1 is shown, and the pan-tilt close-up camera 2, the left pan-tilt camera 10, and the right pan-tilt camera 11 are inside the automatic camera control device.
[0096] In summary, by applying the automatic camera director device provided in the embodiments of the present invention, through the main camera 1, the pan-tilt close-up camera 2, the machine vision analysis module 3, the automatic lens director module 4, the lens angle adjustment module 5, the video encoding module 6, the network output module 7, and the video output interface 8; both the main camera 9 and the pan-tilt close-up camera 10 are connected to the machine vision analysis module 3; the main camera 1 transmits the captured host live video to the machine vision analysis module 3; the pan-tilt close-up camera 2 transmits the captured item display video to the machine vision analysis module 3; the machine analysis module 3 analyzes the received host live video and item display video to obtain a semantic analysis result; the input end of the automatic lens director module 4 is connected to the machine vision analysis module 3, receives the semantic analysis result, and calculates the live broadcast perspective; the lens angle adjustment module 5 is connected to the output end of the automatic lens director module 4, and obtains target video information according to the live broadcast perspective; the video encoding module 6 is connected to the output end of the lens angle adjustment module 5, receives the target video information and encodes it to obtain an encoded video; the network output module 7 is connected to the output end of the video encoding module 6, receives the encoded video, and transmits it to the video output interface 8, and the input end of the video output interface 8 is connected to the network output module 7. It can be seen that, different from traditional single-camera live broadcast devices, the present invention uses the main camera 1 and the pan-tilt close-up camera 2 to bring a better viewing experience and a wider field of view than single-perspective live broadcast devices. Compared with traditional intelligent tracking cameras, the semantic analysis of the machine vision analysis module 3 of this automatic camera director device is rich, and there is no need for manual switching and angle adjustment. Moreover, external videos can be collected through the external video access module 9, making the video of the automatic camera director device more abundant; and, the left pan-tilt camera 10 and the right pan-tilt camera 11 are used to collect the left and right pictures of the host respectively, so that the target will not be lost during the live broadcast, and the live broadcast field of view is improved; and, the first lifting bracket 12 and the second lifting bracket 13 can be used to adjust the heights of the left pan-tilt camera 10 and the right pan-tilt camera 11, making the use of the left pan-tilt camera 10 and the right pan-tilt camera 11 more flexible.
[0097] The following introduces the automatic camera director method provided by the present invention. The automatic camera director method described below is applied to the automatic camera director device described above, and can be correspondingly referred to the automatic camera director device described above.
[0098] Specifically, please refer to Figure 6 , Figure 6 which is a flowchart of an automatic camera director method provided by an embodiment of the present invention, and specifically may include:
[0099] S100, obtain a host live video and an item display video.
[0100] In this embodiment, the host's video is captured by the main camera 1 in the automatic video switcher, and the item display video is captured by the pan-tilt close-up camera 2. This embodiment does not limit the specific host's video. For example, the host's video can be the video of the keynote speaker in a meeting; or the host's video can be the video of the teacher in an instructional video; or the host's video can be the video of the live seller in a live sales video. This embodiment does not limit the specific item display video. For example, the item display video can be the video of the product shown in the keynote speaker's hand in a meeting; or the item display video can be the video of the teaching tool in the teacher's hand in an instructional video; or the item display video can be the video of the item to be sold in a live sales video.
[0101] S101, call the visual analysis model to analyze the host's video and the item display video respectively, and obtain the host semantic analysis result and the item display semantic analysis result.
[0102] The visual analysis model in this embodiment can analyze the host's video and the item display video. Through the visual analysis model, the gaze change of the host and the details of the item display can be obtained. It can be understood that the host's video can be subjected to semantic analysis of the face pose orientation through the visual analysis model to estimate the head pose of the host; the item display video can capture the details of the item that the host wants to display through the visual analysis model. This embodiment does not limit the specific way of using the visual analysis model to analyze the host's video. For example, the head pose of the host can be calculated through the neural network model in the visual analysis model; or the change in the host's viewing angle can be calculated through the neural network model in the visual analysis model; or the upper body posture of the host can be calculated through the neural network model in the visual analysis model. This embodiment does not limit the specific way of using the viewing angle analysis model to analyze the item display video. For example, the change in the host's gesture can be calculated through the gesture recognition model; or the movement of the item can be calculated through the item recognition model.
[0103] S102, call the camera switch calculation model to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result, and determine the live viewing angle.
[0104] This embodiment does not limit the specific camera switch calculation model, as long as it can perform weighted calculation. It can be understood that generally, the video stream of the main camera 1 is defaulted as the output, that is, the weight of the host semantic analysis result is greater than the weight of the item display semantic analysis result. When it is detected that the head pose of the host in the host's video captured by the main camera 1 is significantly downward, the present invention will output the item display video captured by the pan-tilt close-up camera 2.
[0105] Based on the above embodiments, the automatic camera switching method provided by the present invention obtains the host live video and the item display video; calls the visual analysis model to analyze the host live video and the item display video respectively to obtain the host semantic analysis result and the item display semantic analysis result; calls the camera switching calculation model to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result to determine the live broadcast perspective. It can be seen that compared with the traditional single-camera camera switching method, the automatic camera switching method provided by the present invention can obtain the host live video and the item display video, and can calculate a suitable camera switching angle according to the camera switching calculation model. The whole process does not require manual switching and angle adjustment, improving the flexibility of camera switching and preventing the loss of the live broadcast target.
[0106] Further, in order to improve the diversity of the live video, before calling the visual analysis model to analyze the host live video and the item display video respectively, it may further include:
[0107] Obtain the electronic video image collected by the external video access module, the host left image collected by the left pan-tilt camera, and the host right image collected by the right pan-tilt camera;
[0108] Correspondingly, calling the visual analysis model to analyze the host live video and the item display video respectively to obtain the host semantic analysis result and the item display semantic analysis result may include:
[0109] Call the visual analysis model to analyze the host live video, the item display video, the host left image and the host right image respectively to obtain the host semantic analysis result, the item display semantic analysis result, the electronic video semantic analysis result, the host left semantic analysis result and the host right semantic analysis result.
[0110] In this embodiment, the electronic video image is collected by the external video access module. The host left image and the host right image are obtained by the left pan-tilt camera and the right pan-tilt camera. This embodiment does not limit the specific content of the electronic video image collected by the external video access module. For example, the electronic video image may be an electronic PPT image or an electronic PDF image. Since only the host live video is collected by the main camera, the external played electronic video, the host left image and the host right image cannot be collected. Therefore, the electronic video image, the host left image and the host right image are obtained by the external video access module, the left pan-tilt camera and the right pan-tilt camera to enhance the breadth of the camera switching perspective.
[0111] Based on the above embodiments, the visual analysis model of the embodiments of the present invention can also analyze the electronic video images collected by the external video access module, the left-side host images collected by the left pan-tilt camera, and the right-side host images collected by the right pan-tilt camera, making the semantic analysis content more abundant and increasing the diversity of the video stream, thereby improving the accuracy of the automatic video switcher.
[0112] Further, to improve the accuracy of the video live broadcast, the above-mentioned calling of the visual analysis model to analyze the host image and the item display image respectively to obtain the host semantic analysis result and the item display semantic analysis result may include:
[0113] Calling the neural network model in the visual analysis model to calculate the head pose of the host in the host image, and taking the head pose of the host as the host semantic analysis result;
[0114] Calling the neural network gesture recognition model in the visual analysis model to calculate the item display gesture in the item display image, and taking the item display gesture as the item display semantic analysis result.
[0115] This embodiment uses the neural network model to calculate the head pose of the host in the host image because generally, the head orientation of the host in the live broadcast indicates that there are events worthy of attention in the head orientation direction, and when presenting items, there are generally changes in gestures. Therefore, the gesture recognition model is selected to calculate the item display gesture in the item display image.
[0116] Further, to improve the accuracy of the camera video switcher, the above-mentioned calling of the camera video switcher calculation model to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result may include:
[0117] Calling the camera video switcher calculation model corresponding to the live broadcast scenario to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result.
[0118] This embodiment performs weighted calculation on the host semantic analysis result and the item display semantic analysis result by calling the camera video switcher calculation model corresponding to the live broadcast scenario. It can be understood that for different live broadcast scenarios, the weights of the host image and the item display image are different. Therefore, different camera video switcher calculation models can be designed according to different live broadcast scenarios, and the live broadcast image with a higher calculation result can be selected according to the preset weights. For example, for live selling goods, when both the item display image and the host image change simultaneously, the weight of the item display image is higher as the live broadcast image. For example, for a speech scenario, when both the item display image and the host image change simultaneously, the weight of the host image is higher as the live broadcast image.
[0119] In summary, by applying the automatic camera switching method provided by the embodiments of the present invention, the host screen and the item display screen are obtained; the visual analysis model is called to analyze the host screen and the item display screen respectively, and the host semantic analysis result and the item display semantic analysis result are obtained; the camera switching calculation model is called to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result to determine the live broadcast perspective. It can be seen that compared with the traditional single-camera camera switching method, the automatic camera switching method provided by the present invention can obtain the host screen and the item display screen, and can calculate a suitable camera switching angle according to the camera switching calculation model. The whole process does not require manual switching and angle adjustment, which improves the flexibility of camera switching and prevents the loss of the live broadcast target. Moreover, through the electronic video screen collected by the external video access module, the left pan-tilt camera and the right pan-tilt camera respectively collect the left side screen and the right side screen of the host, so that the target will not be lost during the live broadcast and the live broadcast vision is improved. In addition, different camera switching calculation models are called according to different scenarios, making the accuracy of camera switching higher.
[0120] For ease of understanding, please refer to Figure 7 , Figure 7 which is a flowchart example of an automatic camera switching method provided by an embodiment of the present invention. The automatic camera switching method may specifically include:
[0121] S200, obtaining five video streams of the host screen, the item display screen, the left side screen of the host, the right side screen of the host, and the computer PPT screen transmitted from five video input ends of the main camera, the pan-tilt close-up camera, the left pan-tilt camera, the right sub-pan-tilt camera, and the external video access module in a small meeting scenario;
[0122] S201, calling the neural network model, the gesture detection and recognition model, the first optical flow extraction model, the second optical flow extraction model, and the frame change model of the video stream in the machine analysis model to perform semantic analysis on the five video streams respectively, and obtaining the host semantic analysis result, the item display semantic analysis result, the first optical flow semantic analysis result, the second optical flow semantic analysis result, and the frame change semantic analysis result;
[0123] S202, obtaining the camera switching calculation model corresponding to the small meeting scenario to perform weighted calculation on the host semantic analysis result, the item display semantic analysis result, the first optical flow semantic analysis result, the second optical flow semantic analysis result, and the frame change semantic analysis result, and calculating the live broadcast perspective.
[0124] S203. Determine the camera angle based on the live broadcast perspective and determine the target video stream corresponding to the camera angle. In this embodiment, the weight of the external video access module is greater than the weight of the main camera, which is greater than the weights of the left pan-tilt camera and the right pan-tilt camera, which are greater than the weight of the pan-tilt close-up camera. When a PPT switch is detected in the video stream of the external video access module, the present invention will encode and send the video stream of the external video access module. When the video stream of the main camera is detected with the host's head posture significantly deviated to the left or right, the present invention will output the video stream of the left pan-tilt camera or the right pan-tilt camera. When the video stream of the main camera is detected with the host's head posture significantly deviated upward, the present invention will output the video stream of the external video access module. When the video stream of the main camera is detected with the speaker's head posture significantly deviated downward, the present invention will output the video stream of the pan-tilt close-up camera. When the video stream of the main camera is detected with the speaker's head posture in the center, the present invention will switch back to the video stream of the main camera. When a specific gesture is detected in the video stream with the weight of the pan-tilt close-up camera, the present invention will output the video stream of the pan-tilt close-up camera and adjust the angle according to the hand position or switch back to the video stream of the main camera. When the left pan-tilt camera or the right pan-tilt camera detects a participant's speech, the present invention will output the video stream of the left pan-tilt camera or the right pan-tilt camera and adjust the camera angle according to the participant's position. When the above situations occur simultaneously, the video stream with a higher calculation result will be selected according to the pre-set weights. For example, when a PPT switch and the speaker's head deviation to the left or right are detected simultaneously, the video stream of the external video access module will be displayed.
[0125] S204. Encode the target video stream; and send the encoded video.
[0126] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, refer to the description in the method section.
[0127] Finally, it should also be noted that in this article, relationships such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0128] The above has introduced in detail an automatic live director device provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An automatic live broadcast device, characterized in that, It includes a main camera (1), a pan-tilt close-up camera (2), a machine vision analysis module (3), an automatic lens director module (4), a lens angle adjustment module (5), a video encoding module (6), a network output module (7), and a video output interface (8); Both the main camera (1) and the pan-tilt close-up camera (2) are connected to the machine vision analysis module (3); the main camera (1) transmits the captured host live video to the machine vision analysis module (3); the pan-tilt close-up camera (2) transmits the captured item display video to the machine vision analysis module (3); the machine vision analysis module (3) analyzes the received host live video and item display video to obtain a host semantic analysis result and an item display semantic analysis result; The input end of the automatic lens director module (4) is connected to the machine vision analysis module (3), receives the semantic analysis result, and calculates the live broadcast viewing angle; The lens angle adjustment module (5) is connected to the output end of the automatic lens director module (4), and obtains target video information according to the live broadcast viewing angle; The video encoding module (6) is connected to the output end of the lens angle adjustment module (5), receives the target video information and encodes it to obtain an encoded video; The network output module (7) is connected to the output end of the video encoding module (6), receives the encoded video, and transmits it to the video output interface (8), and the input end of the video output interface (8) is connected to the network output module (7).
2. The automatic live broadcast device according to claim 1, characterized in that, It further includes: An external video access module (9) for receiving electronic video information, and the output end of the external video access module (9) is connected to the machine vision analysis module (3).
3. The automatic live broadcast device according to claim 1, characterized in that, It further includes: A left pan-tilt camera (10), the left pan-tilt camera (10) is connected to the machine vision analysis module (3), the left pan-tilt camera (10) is arranged on the left side of the host, and transmits the captured video on the left side of the host to the machine vision analysis module (3).
4. The automatic live broadcast device according to claim 3, characterized in that, It further includes: A right pan-tilt camera (11), the right pan-tilt camera (11) is connected to the machine vision analysis module (3), the right pan-tilt camera is arranged on the right side of the host, and transmits the captured video on the right side of the host to the machine vision analysis module (3).
5. The automatic live broadcast device according to claim 3, characterized in that, It further includes: A first lifting bracket (12); the first lifting bracket (12) is connected to the left pan-tilt camera (10).
6. The automatic live broadcast device according to claim 4, characterized in that, It further includes: A second lifting bracket (13); the second lifting bracket (13) is connected to the right pan-tilt camera (11).
7. An automatic live broadcast method, characterized in that, Applied to the automatic director device according to any one of claims 1 to 6, it includes: Obtain the host live video captured by the main camera (1) and the item display video captured by the pan-tilt close-up camera (2); The machine vision analysis module (3) calls a vision analysis model to analyze the host live video and the item display video respectively to obtain a host semantic analysis result and an item display semantic analysis result; The automatic camera director module (4) calls the camera director calculation model to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result to determine the live broadcast perspective.
8. The automatic live broadcast method according to claim 7, characterized in that, Before the machine vision analysis module (3) calls the vision analysis model to analyze the host screen and the item display screen respectively, it further includes: Obtaining the electronic video screen collected by the external video access module (9), the left side host screen collected by the left pan-tilt camera (10), and the right side host screen collected by the right pan-tilt camera (11); Correspondingly, the machine vision analysis module (3) calls the vision analysis model to analyze the host screen and the item display screen respectively, and obtains the host semantic analysis result and the item display semantic analysis result, including: The machine vision analysis module (3) calls the vision analysis model to analyze the host screen, the item display screen, the electronic video screen, the left side host screen and the right side host screen respectively, and obtains the host semantic analysis result, the item display semantic analysis result, the electronic video semantic analysis result, the left side host semantic analysis result and the right side host semantic analysis result.
9. The automatic live broadcast method according to claim 7, characterized in that, The machine vision analysis module (3) calls the vision analysis model to analyze the host screen and the item display screen respectively, and obtains the host semantic analysis result and the item display semantic analysis result, including: Calling the neural network model in the vision analysis model to calculate the head pose of the host in the host screen, and taking the head pose of the host as the host semantic analysis result; Calling the neural network gesture recognition model in the vision analysis model to calculate the item display gesture in the item display screen, and taking the item display gesture as the item display semantic analysis result.
10. The automatic live broadcast method according to claim 7, characterized in that, The automatic camera director module (4) calls the camera director calculation model to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result, including: Calling the camera director calculation model corresponding to the live broadcast scenario to perform weighted calculation on the host semantic analysis result and the item display semantic analysis result.
Citation Information
Patent Citations
Processing method and device for realizing intelligent product display in live broadcast process
CN112804585A
Lightweight multi-platform interactive video live broadcast cloud broadcast control system
CN214959711U