Video frame switching method during video playback, AI live production method and device, and system
By setting the display frame of video frames frame by frame and the machine learning generation display frame algorithm, the problem of unsmooth screen switching in video director is solved, smooth and smooth screen switching is achieved, and user experience is improved.
Patent Information
- Application Number
- CN202210752055.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-06-28
AI Technical Summary
The existing video director technology cannot achieve the smooth and smooth effect of artificial intelligence director when switching screens, causing dizziness and discomfort for viewers.
By setting the display frame of video frames frame by frame, we ensure that the distance between adjacent video frames does not exceed the maximum allowed movement distance, and combined with machine learning to generate display frame algorithm, smooth switching of video frames is achieved.
It avoids the large jump, stretching and deformation of the picture, improves the user's viewing experience, and achieves a smooth and smooth picture switching effect.
Smart Images

Figure CN115297277B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to, but is not limited to, video playback technology, and more particularly, to a method for switching video frames during video playback, an AI director method and apparatus, and a system. Background Art
[0002] During video directing, it is often necessary to switch video frames. For example, in the scenario of classroom directing, the playback video frame needs to be switched back and forth between a panoramic view, a close-up of students, etc. When switching video frames during artificial intelligence (AI) directing, it is necessary to achieve or even exceed the effect of manual switching while ensuring correctness, but the current effect still fails to meet the desired one. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.
[0004] An embodiment of the present disclosure provides a method for switching video frames during video playback, including:
[0005] Determine that it is necessary to transform the display frame of a video frame in a video frame sequence into a target display frame, where the display frame of the video frame is used to indicate the area where the picture to be played is located in the video frame;
[0006] Set the display frames of the video frames in the video frame sequence frame by frame to complete the transformation, where the distance between the display frames set for adjacent video frames is not greater than the maximum allowable moving distance of the display frame, and the distance between two display frames is determined according to the distance between the corresponding positioning points of the two display frames;
[0007] Determine the picture to be played in the video frame according to the display frames set for the video frames in the video frame sequence and perform playback processing.
[0008] An embodiment of the present disclosure further provides a video frame switching apparatus, including a processor and a memory storing a computer program, where when the processor executes the computer program, it can implement the video frame switching method according to any embodiment of the present disclosure.
[0009] The video frame switching method and apparatus of the above embodiments of the present disclosure can avoid the dizziness and discomfort of viewers caused by large jumps in the video frame, and achieve a smooth and fluent video frame switching effect.
[0010] An embodiment of the present disclosure further provides an AI directing method, including:
[0011] Process an input video frame sequence using a display frame generation algorithm based on machine learning to generate display frames of video frames;
[0012] When it is determined according to the change of the generated display box that the display box of the video frame in the video frame sequence needs to be transformed into a target display box, the screen switching during video playback is implemented according to the screen switching method described in any embodiment of the present disclosure.
[0013] An embodiment of the present disclosure further provides an AI live broadcast device, including a processor and a memory storing a computer program, wherein when the processor executes the computer program, it can implement the AI live broadcast method described in any embodiment of the present disclosure.
[0014] An embodiment of the present disclosure further provides an AI live broadcast system, including:
[0015] A video collector, configured to collect a video frame sequence and output it;
[0016] An AI live broadcast device, configured to receive the video frame sequence, perform AI live broadcast according to the AI live broadcast method described in any embodiment of the present disclosure, and output video data obtained by performing playback processing on the to-be-played screen;
[0017] A display, configured to display the to-be-played screen according to the video data.
[0018] The AI live broadcast method, device and system of the above embodiments of the present disclosure implement screen switching during video playback according to the screen switching method of the embodiments of the present disclosure, and can also avoid large jumps, stretching and deformation of the screen, making the screen switching smooth and improving the user's viewing experience.
[0019] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, which stores a computer program that can implement the screen switching method described in any embodiment of the present disclosure or the AI live broadcast method described in any embodiment of the present disclosure when executed by a processor.
[0020] An embodiment of the present disclosure further provides a computer program product, including a computer program that can implement the screen switching method described in any embodiment of the present disclosure or the AI live broadcast method described in any embodiment of the present disclosure when executed by a processor.
[0021] Other aspects can be understood after reading and understanding the drawings and the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The drawings are used to provide an understanding of the embodiments of the present disclosure, and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure, and do not constitute a limitation to the technical solutions of the present disclosure.
[0023] Figure 1Flow chart of a screen switching method according to an embodiment of the present disclosure;
[0024] Figure 2 Schematic diagram of an exemplary display frame transformation according to an embodiment of the present disclosure;
[0025] Figure 3 Flow chart of an AI director method according to an embodiment of the present disclosure;
[0026] Figure 4 Schematic diagram of the intersection over union of two rectangular frames;
[0027] Figure 5 Schematic diagram of an exemplary state transition according to an embodiment of the present disclosure;
[0028] Figure 6 Schematic diagram of a screen switching device according to an embodiment of the present disclosure;
[0029] Figure 7 Schematic diagram of an AI director system according to an embodiment of the present disclosure;
[0030] Figure 8 Flow chart of an AI director method according to another embodiment of the present disclosure. Detailed implementation manners
[0031] The present disclosure describes multiple embodiments, but the description is exemplary rather than restrictive, and it will be obvious to those of ordinary skill in the art that there can be more embodiments and implementation solutions within the scope of the embodiments described in the present disclosure.
[0032] In the description of the present disclosure, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment described as "exemplary" or "for example" in the present disclosure should not be construed as being more preferred or having more advantages than other embodiments. The "and / or" herein is a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. "Multiple" means two or more than two. In addition, in order to clearly describe the technical solutions of the embodiments of the present disclosure, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the words "first" and "second" do not limit the quantity and execution order, and the words "first" and "second" do not necessarily limit to being different.
[0033] In describing representative exemplary embodiments, the specification may have presented methods and / or processes as a particular sequence of steps. However, to the extent that the method or process does not depend on a particular order of the steps described herein, the method or process should not be limited to the particular order of steps described. As will be understood by those of ordinary skill in the art, other step orders are possible. Accordingly, the particular order of steps set forth in the specification should not be construed as a limitation on the claims. In addition, the claims directed to the method and / or process should not be limited to performing their steps in the order written, as those skilled in the art can readily understand that such orders may vary and still remain within the spirit and scope of the embodiments of the present disclosure.
[0034] The AI director can adaptively switch the playback screen according to changes in the shooting scene. For example, in the AI director for a classroom shooting scene, the playback screen needs to switch between the panoramic view of the classroom and the close-up view of the students. When the playback screen (referred to as the screen for short) switches from the panoramic view of the classroom to the close-up view of the students, it is called zooming in (or pushing in) of the screen. When the playback screen switches from the close-up view of the students to the panoramic view of the classroom, it is called zooming out (or pulling back) of the screen. In addition, the playback screen may also need to switch from the close-up view of one student to the close-up view of another student, which is called moving of the screen. In this article, the playback screen can also be referred to as a shot (not referring to a physical lens), and the switching of the playback screen can also be referred to as a shot transition.
[0035] During AI directing, algorithms such as deep learning can be used to determine the target screen to switch to. When actually performing the switch, one way is to control the movement of the pan-tilt head where the video collector such as a camera is located, or control the zoom of the physical lens of the video collector to achieve effects such as zooming in, zooming out, and moving of the above-mentioned screen. However, this method has additional requirements for the shooting device, which needs to support functions such as the movement of the pan-tilt head and the zoom of the physical lens, and also requires the shooting device to open control permissions to the AI directing device. Therefore, its application is greatly limited. Another way is to perform image processing on the video frames in the captured video frame sequence to achieve effects such as zooming in, zooming out, and moving of the above-mentioned screen. For example, after obtaining the video frame sequence of the classroom panorama by shooting, if it is necessary to switch the playback screen from the panoramic view of the classroom to the close-up view of the students, the area of the video frame containing the student screen can be enlarged and then played to achieve the effect of zooming in the screen. When the resolution of the video frame is sufficient, it can also meet the viewing needs of users.
[0036] When realizing the picture switching during video playback through image processing, when performing image processing on video frames, the area where the picture to be played of the video frame is located can be set. In this text, this area is called the display frame of the video frame. The display frame of the video frame can be represented by the position information of the vertices of the display frame in the video frame. Or rather, the area indicated by the display frame is represented by the position information of the vertices of the display frame in the video frame. For example, a rectangular display frame can be represented by the position coordinates of the four vertices of the rectangle, or by the position coordinates of two diagonal vertices, or by the position coordinates of a specified vertex and the length and width of the rectangle, and so on. When the playback picture is a panoramic picture of a classroom, the display frame of the video frame can be set as the area where the entire video frame is located, so that the playback picture is a panoramic picture of the classroom; and when it is necessary to switch the playback picture from the panoramic picture of the classroom to a picture of a student who needs a close-up, the display frame set for the video frame can be changed, and the display frame of the video frame can be set as the area where the student picture in the video frame is located, so that the playback picture is switched to a close-up picture of the student.
[0037] By changing the display frame set for the video frame, that is, the transformation of the display frame, before playing the video frame, and then determining the picture to be played in the video frame according to the display frame set for the video frame and performing playback processing (such as magnification) before playing, the picture switching can be realized. In this text, the display frame set based on the picture to be switched to is called the target display frame. That is to say, after the picture to be played in the video frame indicated by the target display frame is played, it is the picture to be switched to.
[0038] Based on the transformation of the display frame set for the video frame, the picture switching can be realized. When it is necessary to transform the display frame of the video frame into the target display frame, the transformation can be directly completed. For example, when it is necessary to switch the playback picture from the panoramic picture of the classroom to a picture of a student who needs a close-up, in the case where the display frame set for the previous video frame indicates the entire video frame area, the display frame of the current video frame is directly set as the area where the student picture in the video frame is located. This setting method is relatively simple, but it is easy to cause large jumps, stretches and deformations of the picture, resulting in dizziness and discomfort of the viewers. Therefore, it is necessary to provide a method for picture switching during video playback to avoid large jumps, stretches and deformations of the picture as much as possible, achieve a smooth and fluent picture switching effect, and improve the viewing experience of users.
[0039] For this reason, an embodiment of the present disclosure provides a method for picture switching during video playback, as Figure 1 shown, including:
[0040] Step 110, determining that it is necessary to transform the display frame of the video frame in the video frame sequence into a target display frame, where the display frame of the video frame is used to indicate the area where the picture to be played is located in the video frame;
[0041] The video frame sequence of this step can be captured by a camera or other video capture device, but the present disclosure is not limited thereto. In other embodiments, the method of this embodiment can also be used to implement screen switching when playing a video file.
[0042] The target display frame of this step can be determined according to user operations. For example, during manual live broadcast, the target display frame to be switched to is directly delimited on the video screen according to user input. The target display frame can also be jointly determined by combining algorithms and user operations. For example, according to user operations, the target image of the screen to be switched to, such as a certain student, is determined, and the target display frame surrounding the target image is generated through an image processing algorithm. During AI live broadcast, the target display frame can be generated using an AI algorithm, and the present disclosure is not limited thereto.
[0043] Step 120: Set the display frames of the video frames in the video frame sequence frame by frame to complete the transformation, where the distance between the display frames set for adjacent video frames is not greater than the maximum allowable moving distance of the display frame;
[0044] In this step, the distance between two display frames is determined according to the distance between the corresponding positioning points of the two display frames; in a video frame, the position coordinates of a point can be determined according to the number of pixels between this point and the origin (such as the upper left corner point or the center point of the video frame) in the X and Y directions, and the distance between points can be calculated according to the position coordinates of the two points.
[0045] Step 130: Determine the picture to be played in the video frame according to the display frames set for the video frames in the video frame sequence and perform playback processing.
[0046] In this step, the picture to be played in the video frame can be determined according to the display frame set for the video frame. If the area where the picture to be played is located is the entire video frame area, when performing playback processing on the picture to be played, corresponding display processing can be performed to generate corresponding video data to be displayed and output to the display for display. If the area where the picture to be played is located is a partial area in the video frame, when performing playback processing on the picture to be played, the picture to be played can be extracted first and magnified before performing display processing, so that the displayed picture can be adapted to the size of the display screen. If the aspect ratio of the picture to be played extracted according to the display frame does not match the aspect ratio of the display screen, the display frame can also be corrected first to extract a picture to be played with a suitable aspect ratio.
[0047] This embodiment realizes screen switching through video processing without the need to control the video capture device. Therefore, there are no requirements for the platform movement function, zoom ability, and access rights of the video capture device, as long as there is sufficient image resolution. It can also be used for screen switching during video file playback, so it can be applied in more scenarios and has better adaptability.
[0048] In this embodiment, the screen switching during video playback is achieved through the transformation of the display frame. When performing the transformation of the display frame, it is ensured that the distance between the display frames set for adjacent video frames is not greater than the maximum allowable moving distance of the display frame, so as to avoid the dizziness and discomfort of the viewers caused by large jumps in the screen, achieve a smooth and fluent screen switching effect, and improve the viewing experience of users.
[0049] In an exemplary embodiment of the present disclosure, setting the display frames of the video frames in the video frame sequence frame by frame to complete the transformation includes:
[0050] Set the display frame of the current video frame in the following manner: Determine the distance D between the display frame of the previous video frame and the target display frame m ; when D m is less than or equal to the maximum allowable moving distance D of the display frame ist , set the display frame of the current video frame as the target display frame to complete the transformation; when D m is greater than D ist , set the display frame of the current video frame to a display frame with a distance of D from the display frame of the previous video frame ist and a distance of D1 - D from the target display frame ist ;
[0051] If the transformation is not completed, take the next video frame as the current video frame and continue to set the display frame of the current video frame in the same manner until the transformation is completed.
[0052] The processing of the video frames in the video frame sequence is performed frame by frame. The video frame being processed is called the current video frame. After this video frame is processed, the next video frame will be processed in sequence, and at this time, the next video frame will become the current video frame.
[0053] After determining that the transformation of the display frame needs to be performed in this embodiment, the display frame of the current video frame being processed (assumed to be the i-th video frame in the video frame sequence) will be set in the above manner. If D m is less than or equal to D ist , it means that the distance between the display frame of the previous video frame (the (i - 1)-th video frame) and the target display frame is relatively small, and the display frame of the current video frame will be directly set as the target display frame. If D m is greater than D ist , it means that the distance between the display frame of the previous video frame and the target display frame is relatively large, and multiple video frames are required to complete the transformation. Then, the display frame of the current video frame will be set to a display frame with a distance of D from the display frame of the previous video frame ist, the distance to the target display box is D1 - D ist For the display box. When the transformation is not completed for the current video frame, that is, the display box set for the current video frame is not the target display box, the next video frame (the (i + 1)-th video frame) is taken as the current video frame, and the display box is set in the same way until the transformation is completed. Thus, it can be ensured that the distance between the display boxes set for adjacent video frames does not exceed the maximum allowable movement distance of the display box, achieving a smooth switching effect.
[0054] In an example of this embodiment, the display box is rectangular, the positioning points of the display box include the four vertices of the rectangle, and the display box is represented by the position information of the vertices of the display box in the video frame; the distance D m between the display box of the previous video frame and the target display box is the distance D A 、D B 、D C 、D D among them.
[0055] In Figure 2 the example shown, the display box of the previous video frame is a rectangle with vertices A o , B o , C o , D o . This rectangle can be the entire video frame area or a partial area in the video frame. The target display box is a rectangle with vertices A t , B t , C t , D t . In this example, D A is the distance from A o to A t (i.e., the length of line segment A o A t ), D B is the distance from B o to B t (i.e., the length of line segment B o B t ), D C is the distance from C o to C t (i.e., the length of line segment C o C t ), D D is the distance from D o to D t (i.e., the length of line segment D o D t . In this example, vertices A o , B o, C o , D o and vertex A t , B t , C t , D t are in one-to-one correspondence, where A o and A t are both the upper left vertices, B o and B t are both the upper right vertices, C o and C t are both the lower right vertices, D o and D t are both the lower left vertices, and the maximum value among D A , D B , D C , D D is D D .
[0056] In Figure 2 the example shown, the target display box is inside the display box of the previous video frame, which is the display box transformation for zooming in the picture. In another example, it is the display box transformation for zooming out the picture, and this example can also be represented by Figure 2 , and at this time the vertices are A t , B t , C t , D t The rectangle represents the display box of the previous video frame, and the vertices are A o , B o , C o , D o The rectangle represents the target display box, that is, the display box of the previous video frame is inside the target display box. In another example, it is the display box transformation for moving the picture, and at this time the display box of the previous video frame has no intersection or partial intersection with the target display box. It is easy to understand that in these examples, the distance D m between the display box of the previous video frame and the target display box can all be expressed as the distance D A , D B , D C , D D among the maximum values between the corresponding vertices of the display box of the previous video frame and the target display box, and this distance can objectively reflect the jump amplitude during picture switching.
[0057] In an example of this embodiment, the maximum allowable moving distance of the display box is equal to the time difference between the current video frame and the previous video frame multiplied by the maximum moving distance per unit time preset, and is expressed by the formula: D ist= (curr_time – last_time) * pixel_per_sec, where curr_time represents the playback time of the current video frame, last_time represents the playback time of the previous video frame, and pixel_per_sec is the maximum moving distance per unit time preset. In this example, the maximum allowable moving distance of the display box can vary with the time difference between adjacent video frames. The greater this time difference, the greater the maximum allowable moving distance D ist is. That is, D ist can vary with the change of the video frame rate. While maintaining the stable and smooth switching of the picture without affecting the visual experience, it can also avoid the impact on the visual experience caused by too long switching time and achieve a suitable balance. The maximum moving distance per unit time preset in this example can be a fixed value, but it can also be a variable value, such as varying with the change of the video scene. For example, when there are fast-moving objects in the scene, the maximum moving distance per unit time is set to be relatively large, and in other cases, the maximum moving distance per unit time is set to a relatively small value. Or, the value of the maximum moving distance per unit time is changed according to the atmosphere of the scene. The above maximum moving distance per unit time can be set according to empirical values and can be adjusted according to the actual effect during the design process.
[0058] In another example of this embodiment, the maximum allowable moving distance of the display box can also be set to a fixed value, which can also be set according to experience and can be adjusted during the experiment.
[0059] In an exemplary embodiment of the present disclosure, setting the display box of the current video frame to a distance D from the display box of the previous video frame ist and a distance D1 - D from the target display box ist for the display box includes:[[]]
[0060] Determine the distances between the four vertices of the display box of the current video frame and the corresponding vertices of the display box of the previous video frame: D A ’ = D ist * D A / D m , \ D B ’ = D ist * D B / D m , \ D C ’ = D ist * D C / D m , \ D D ’ = D ist * D D / Dm , where "*" represents the multiplication operation and " / " represents the division operation;
[0061] Determine the positions of the four vertices A c , B c , C c , D c of the display frame in the current video frame. Among them, A c is the point on the line segment A o A t whose distance to A o is D A '; B c is the point on the line segment B o B t whose distance to B o is D B '; C c is the point on the line segment C o C t whose distance to C o is D C '; D D is the point on the line segment D o D t whose distance to D o is D D '. In the example shown in Figure 2 , the display frame set for the current video frame is the rectangle with vertices A c , B c , C c , D c in the figure.
[0062] Although Figure 2 represents the transformation of the display frame during the process of zooming in on the picture, for the transformation of the display frame during the process of zooming out on the picture and the process of picture movement (no intersection or partial intersection between the pictures before and after switching), the positions of the four vertices of the display frame in the current video frame can also be determined in the above manner, so as to complete the setting of the display frame of the current video frame.
[0063] Through the above method of this embodiment, the vertices of the display frame of the current video frame can move from the corresponding vertices of the display frame of the previous video frame to the vertices of the target display frame in equal proportion, making the picture change during the picture switching process smoother and more stable, avoiding stretching and deformation of the picture, and thus improving the viewing effect.
[0064] During video playback, the screen may switch frequently, and the target display frame may be updated multiple times. In an exemplary embodiment of the present disclosure, after the transformation is completed (i.e., after the display frame of the video frame in the video frame sequence is transformed into the target display frame), the method further includes: when the target display frame is not updated, setting the display frame of the video frame in the video frame sequence to the target display frame until the target display frame is updated (i.e., the target display frame changes). After the display frame transformation is completed or during the transformation, if the target display frame is updated, the target display frame is transformed again according to the updated target display frame to transform the display frame set for the video frame into the updated target display frame frame by frame. In this embodiment, when the transformation of the display frame is completed once and the target display frame is not updated, no new transformation is performed, so that the playback screen after the switch remains unchanged, and the stability of the screen is enhanced, for example, it is fixed to a panoramic screen or a close-up screen.
[0065] An embodiment of the present disclosure also provides an AI directing method, such as Figure 3 Shown, including:
[0066] Step 210: Process the input video frame sequence using a display frame generation algorithm based on machine learning to generate a display frame of the video frame;
[0067] Step 220 : When it is determined that the display frame of the video frame in the video frame sequence needs to be transformed into the target display frame according to the change of the generated display frame, screen switching during video playback is implemented according to the screen switching method described in any embodiment of the present disclosure.
[0068] When AI directs, this embodiment automatically implements screen switching during video playback according to the screen switching method described in any embodiment of the present disclosure, thereby avoiding large screen jumps, stretching and deformation, making the screen switching smooth and improving the user's viewing experience.
[0069] In an exemplary embodiment of the present disclosure, the machine learning-based display frame generation algorithm is executed by a trained display frame generation model based on a deep learning network, wherein the deep learning network is trained using a video frame sequence as a sample as input data and display frames marked in the video frames based on focus events in the sample as target data.
[0070] Focus events vary with different scenarios. Taking the classroom scenario as an example, "a student stands up" and "a student sits down" can be regarded as focus events. After obtaining a sequence of video frames as samples, in one or more video frames of the samples where the focus event "a student stands up" occurs, a display box of a close-up view with the standing student as the target object is marked. In one or more video frames where the focus event "a student sits down" occurs, a display box of a panoramic view can be marked, such as taking the entire video frame area as a display box. Using the sequence of video frames as input data and the display boxes marked in the video frames based on the focus events in the samples as target data to train a deep learning network, the trained display box generation model based on the deep learning network can have the function of inferring according to the focus events in the input video frames and generating and outputting corresponding display boxes. Among them, the deep learning network can be a deep neural network such as a deep convolutional neural network (CNN), but it is not limited to this. Whether it is trained well can be verified using a sequence of video frames as a validation sample set. After inputting the validation sample into the deep neural network, by comparing the display box output by the deep neural network with the marked display box, it is judged whether the accuracy requirement is met. When the accuracy requirement is met, it is considered that the deep neural network has been trained well.
[0071] In an exemplary embodiment of the present disclosure, determining that the display box of a video frame in the video frame sequence needs to be transformed into a target display box according to the change of the generated display box includes:
[0072] Calculating the intersection over union (IoU) between the currently generated display box and the recorded target display box, and accumulating the number of times the calculated IoU is less than the set IoU threshold:
[0073] When the accumulated number of times reaches the set number threshold N (N≥2), updating the recorded target display box to the currently generated display box, determining that the display box of the video frame in the video frame sequence needs to be transformed into the updated target display box, and restarting the accumulation.
[0074] The display box generation model can generate display boxes at a fixed interval, such as generating one display box per frame or generating one display box every M frames (M≥2), but it is not limited to this. This interval can also change according to certain rules in different situations. In the initial state, the recorded target aggregation box can default to the entire video frame area.
[0075] In this embodiment, the update of the target display box is accompanied by the transformation of the display box. That is, after each update of the target display box, it is determined that the display box of the video frame in the video frame sequence needs to be transformed into the updated target display box. If the target display box is not updated and the transformation of the display box brought about by the previous update has been completed, the display box set for the video frame remains unchanged.
[0076] The intersection over union (IoU) between two bounding boxes is the ratio of the intersection between the two bounding boxes to the union of the two bounding boxes. As Figure 4 shown, the intersection between the two bounding boxes is represented by the box with diagonal lines, while the union of the two bounding boxes is the part enclosed by the outer contour lines of the two bounding boxes, which is equal to the sum of the areas of the two bounding boxes minus the intersection between the two bounding boxes. The IoU can reflect the degree of difference between the two bounding boxes to a certain extent. The smaller the IoU, the smaller the overlap between the two bounding boxes and the greater the difference. Therefore, in this embodiment, when the IoU between the display bounding box generated by the display bounding box generation model and the recorded target display bounding box is greater than or equal to the IoU threshold, it indicates that the generated display bounding box is very similar to the recorded target display bounding box and no transformation is required. When the IoU is less than the preset IoU threshold, to avoid jitter in the display bounding box output by the display bounding box generation model (i.e., the offset between the display bounding boxes output before and after when no transformation of the display bounding box is needed), both the IoU threshold and the number threshold can be set according to experience and adjusted. For example, the IoU threshold can be set to 0.3 and the number threshold can be set to 7 times, but this is only exemplary.
[0077] In this embodiment, it is determined that a transformation is needed only when the accumulated number reaches the set number threshold, which can achieve a smooth effect, make the display bounding box relatively stable, reduce transformations, thereby reducing screen jitter and unnecessary screen switching.
[0078] In an exemplary embodiment of the present disclosure, the AI director method further includes the following state update processing for some or all of the following:
[0079] When starting video playback, it is marked as the initial state;
[0080] When updating the recorded target display bounding box and the updated target display bounding box is inside the target display bounding box before the update, it is marked as the zoom-in state;
[0081] When updating the recorded target display bounding box and the target display bounding box before the update is inside the target display bounding box after the update, it is marked as the zoom-out state;
[0082] When updating the recorded target display bounding box and there is no intersection between the target display bounding box before the update and the target display bounding box after the update, it is marked as the move state;
[0083] When the transformation is completed and no further transformation of the display bounding box is required, it is marked as the fixed state.
[0084] The above-mentioned close state, far state, etc. can be used to represent the states of screen switching, and these states respectively correspond to specific transformations of the target display frame. In addition to the definitions of the above-mentioned states, for updating the recorded target display frame, when the target display frame before the update and the target display frame after the update partially intersect, it can be marked as the moving state, or marked as the close state or the far state according to the relative sizes of the target display frame before the update and the target display frame after the update, or ignored without marking.
[0085] In the state update process, it is not required that any conversion can be performed between all states. For example, in one example, only the initial state, the close state, the fixed state, and the far state are set. And the possible transitions between these states are defined. Taking Figure 5 the schematic diagram of the state transition shown as an example, in this example, the display frame corresponding to the panoramic view is set for the video frame in the initial state. It can be transitioned from the initial state to the close state. For example, the panoramic view is switched to the close-up view through the transformation of the display frame. After the transformation of the display frame is completed, it is transitioned from the close state to the fixed state, and the display frame set for the video frame is fixed as the currently recorded target display frame. At this time, the screen is fixed and no switching occurs. It can also be transitioned from the close state to the close state, which means that during a display frame transformation process, the recorded target display frame is updated again, and the target display frame after the update is located inside the target display frame before the update, showing the effect of continuous close-up of the screen on the screen. It can be transitioned from the fixed state to the close state. For example, the screen is switched from one close-up view to another specific view in the specific view through the transformation of the display frame. It can also be transitioned from the fixed state to the far state. For example, the screen is switched from the specific view to the global view through the transformation of the display frame. After the transformation of the display frame is completed, it enters the initial state. By restricting the minimum size of the display frame, the resolution of the close-up screen can be prevented from being too low.
[0086] The above-mentioned state update process is optional. Regarding the transformation of the display frame, in the case of the display frame of the previous video frame, that is, the current target display frame, no transformation of the display frame is required. In the case where the display frame of the previous video frame is different from the current target display frame, the display frame can be transformed according to the method of the embodiments of the present disclosure. The states marked by the above-mentioned state update process can provide useful state information for other controls or be provided to the outside through an interface.
[0087] In an exemplary embodiment of the present disclosure, the processing of the AI director is implemented through at least two threads. The playback processing is executed in the first thread, and the display frame generation algorithm (generating and outputting the display frame of the video frame according to the input video frame sequence) is executed in the second thread parallel to the first thread. The first thread and the second thread can be processed based on different copies of the same video frame sequence.
[0088] The display box generation algorithm takes relatively more time. If it is implemented in the same thread as the playback process, the real-time requirement for the display box generation algorithm is very high. If the real-time performance cannot meet the requirements, the frame rate of video playback will be reduced. In this embodiment, by implementing the playback process and the display box generation algorithm with different threads, the real-time requirement for the display box generation algorithm can be reduced, and the video playback frame rate can be increased. Thus, the delay caused by screen switching can be controlled to an extent that does not affect the viewing experience. For other processes in AI, such as the update of the target display box, the setting of the display box for video frames, and the status update, these processes can be implemented either in the first thread or in the second thread. The two threads run in parallel and can transfer data to each other. In one example, the second thread can also transfer the position information of the vertices of the display box currently output by the display box generation model to the first thread, and the first thread can update the target display box, set the display box for the video frame, and perform the playback process according to the display box of the video frame.
[0089] An embodiment of the present disclosure also provides a screen switching device, as Figure 6 shown, including a processor 60 and a memory 50 storing a computer program. Among them, when the processor 60 executes the computer program, it can implement the screen switching method as described in any embodiment of the present disclosure.
[0090] An embodiment of the present disclosure also provides an AI live director device, which can also be referred to Figure 6 to, including a processor and a memory storing a computer program. Among them, when the processor executes the computer program, it can implement the AI live director method as described in any embodiment of the present disclosure.
[0091] The processor in the above-mentioned screen switching device and AI live director device in the embodiments can be an integrated circuit chip with signal processing capabilities. For example, the processor can be a general-purpose processor, including a central processing unit (CPU for short) and a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0092] An embodiment of the present disclosure also provides an AI live director system, as Figure 7 shown, including:
[0093] A video collector 1, configured to collect a video frame sequence and output it;
[0094] The AI director device 2 is configured to receive the video frame sequence, perform AI direction according to the AI direction method described in any embodiment of the present disclosure, and output video data obtained after performing playback processing on the to-be-played picture.
[0095] The display 3 is configured to display the to-be-played picture according to the video data.
[0096] In this embodiment, the video collector 1, the AI director device 2, and the display 3 in the AI direction system may be physically independent of each other, or may be arbitrarily combined and integrated into one device. For example, the video collector 1, the AI director device 2, and the display 3 may be integrated into one device such as a mobile phone to implement the AI direction method of the embodiments of the present disclosure. Another example is that two of the video collector 1, the AI director device 2, and the display 3 may be integrated into one device, and the other is used as an independent device.
[0097] The AI direction system of this embodiment can execute the AI direction method of any embodiment of the present disclosure, automatically realize the picture switching during video playback, avoid large jumps, stretching, and deformation of the picture, make the picture switching smooth and fluent, and improve the user's viewing experience.
[0098] An embodiment of the present disclosure further provides a non-transitory computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the picture switching method described in any embodiment of the present disclosure, or implement the AI direction method described in any embodiment of the present disclosure.
[0099] An embodiment of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it can implement the picture switching method described in any embodiment of the present disclosure, or implement the AI direction method described in any embodiment of the present disclosure.
[0100] In an embodiment of the present disclosure, in order to enable smooth switching of the camera in AI direction, as well as the realization of the effects of zooming in, zooming out, and panning, an AI direction method is proposed. This AI direction method can be based on Figure 7 The AI direction system shown is implemented. As Figure 8 shown, the AI direction method of this embodiment includes:
[0101] Step 310, data splitting;
[0102] In this step, the image data of the video file or the video frame sequence collected by the video collector is read and divided into two parts (one of which can be obtained by copying), and are respectively sent to two parallel threads, called the first thread and the second thread. In this embodiment, the first thread implements Figure 8The processing from step 340 to step 360, including playback processing; implemented by the second thread Figure 8 The processing from step 320 to step 330, including algorithmic inference for generating a display box. The first thread can also be called the display thread, and the second thread can also be called the algorithmic inference thread. However, Figure 8 This thread division shown can be changed. For example, steps 330 to 350 among them can all be implemented in the first thread or the second thread, and so on.
[0103] In this embodiment, the playback processing and algorithmic inference are separated. The display thread reads the video frame sequence and can obtain the video data for final display with only a small amount of computation, which can improve the playback frame rate and reduce the real-time requirement for algorithmic inference. After the algorithmic inference thread reads the video frame sequence, it has sufficient time to perform inferences on various algorithms, and the accuracy and precision of the results can maintain a high standard, and the calculation results are fed back to the display thread. Of course, the algorithmic inference thread also needs to ensure a certain time efficiency.
[0104] Step 320, display box generation;
[0105] This step can perform inference on the input video frame sequence based on a network model to generate and output the display box of the video frame. Specifically, this step can include the following sub-steps: Step 3201, image preprocessing, such as cropping the video frames in the video frame sequence to meet the requirements of the network model for the input image; Step 3203, the network model performs inference based on the input video frame sequence to generate the display box of the video frame, and specific reference can be made to the above description. Step 3205, image postprocessing, such as extracting or calculating the position information of the vertices of the display box in the video frame from the data output by the network model.
[0106] Step 330, target display box update;
[0107] The display box output by the network model may be unstable and may experience a certain amount of jitter. This embodiment performs some smoothing processing for this phenomenon, and at the same time determines whether to update the current target display box according to the difference in the focus change between the previous and subsequent outputs.
[0108] In one example, when the network model (i.e., the display box generation model) outputs a new target box, the intersection over union (IOU) of the recorded target display box (which can also be referred to as the current target display box) Curr_Tgt_Box and the display box Next_Tgt_Box currently output by the network model can be calculated. The display box currently output by the network model is also the most recently output display box of the network model. If the calculated IOU is less than the set IOU threshold IOU_Threshold, i.e., IOU < IOU_Threshold, the count of IOU < IOU_Threshold will be incremented by 1, that is, this count will be accumulated. When the accumulated count reaches the set count threshold, the recorded target display box Curr_Tgt_Box will be updated to Next_Tgt_Box, and the accumulated count will be cleared, and the accumulation will start again.
[0109] In one example, the above update algorithm can be represented in the form of code as follows:
[0110] Calculate IOU(Curr_Tgt_Box, Next_Tgt_Box);
[0111] If IOU < IOU_Threshold:
[0112] Cnt++;
[0113] If Cnt > 6:
[0114] Cnt = 0;
[0115] At the same time, update Curr_Tgt_Box = Next_Tgt_Box;
[0116] In this example, Cnt++ means incrementing the accumulated count Cnt by 1. This example updates Curr_Tgt_Box when the accumulated count Cnt is greater than 6, indicating that the set count threshold is 7.
[0117] Step 340, state update;
[0118] When performing AI live production, the display box in the initial state can be set to the entire video frame area, so that the played picture starts from the panoramic picture captured by the camera. According to whether there is a focus event in the video (such as a student standing up to answer a question), it is decided whether to zoom in on the picture. If a focus event occurs, the panoramic picture can be gradually switched to the close-up picture of the target in the focus event through the transformation of the display box. If the focus event ends (such as the student sitting down), the picture can be slowly pulled back to the panoramic picture through the transformation of the set display box.
[0119] In one example, the schematic diagram of state transition is as Figure 5As shown, the state transition diagram includes the following states: an initial state (STATE_ORIG); a zoom-in state (STATE_ZOOM_IN), where the screen switching in this state is similar to the effect of a camera zooming in; a fixed state (STATE_ZOOMED), for example, the state when the playback screen has switched to a close-up screen and remains in that screen; a zoom-out state (STATE_ZOOM_OUT), for example, when combined with a focus event, the state when the screen is gradually pulled back to the original screen. The transitions between these states have been described above and will not be elaborated here.
[0120] It should be noted that in another embodiment, the above state update can also be simplified to two states, one is the updated state of the target display frame, and the other is the non-updated state of the target display frame. Each time the target display frame is updated, it is marked as the updated state, indicating that the screen is being switched; after the transformation to the current target display frame is completed, the updated state is set to the non-updated state, indicating that the playback screen is fixed. These two states can be represented by a flag bit.
[0121] Step 350, display frame setting;
[0122] In this step, based on the vertex positions (curr) of the display frame of the previous video frame and the vertex positions (target) of the target display frame to which it needs to be transformed, the vertex positions of the display frame set for the current video frame are calculated. The movement of the vertices of the display frame of the current video frame relative to the vertices of the display frame of the previous video frame corresponds to the distance of the screen being zoomed in or out.
[0123] In an example, the process of display frame setting includes the following steps:
[0124] Step 1, obtain the playback time curr_time of the current video frame and the playback time last_time of the previous video frame. According to the time difference between the current video frame and the previous video frame and the maximum allowable movement distance pixel_per_sec per unit time, calculate the distance Dist between the display frame set for the current video frame and the display frame of the previous video frame. Dist can also be referred to as the movement distance of the display frame set for the current video frame relative to the display frame of the previous video frame, and is expressed by the formula as follows:
[0125] Dist = (curr_time – last_time) * pixel_per_sec
[0126] Among them, pixel_per_sec can be an empirical value deduced according to the actual effect.
[0127] Step 2, according to the display frame last_box[A o ,Bo , C o , D o 's position and the target display box tgt_box[A t , B t , C t , D t 's position, and the moving distance Dist calculated in the previous step, calculate the display box curr_box[A of the current video frame c , B c , C c , D c 's position. The positions of these display boxes can be represented by the coordinates of the vertices. Please refer to Figure 2 .
[0128] An exemplary calculation process is as follows:
[0129] Step A: Calculate the display box last_box[A of the previous video frame o , B o , C o , D o 's four vertices to the corresponding vertices of the target display box tgt_box[A t , B t , C t , D t ':
[0130] D A = A o - A t ; D B = B o - B t ; D C = C o - C t ; D D = D o - D t ;
[0131] A o - A t in the above formula represents the distance from point A o to point A t . Similarly for others. The distance between two points can be calculated based on the coordinates of these two points.
[0132] Step B: Calculate Dm = MAX(D A , D B , D C , D D ), that is, Dm is the maximum value of D A , D B , D C , D D .
[0133] If Dmax < Dist, set the display box of the current video frame to the target display box, that is, let: A c = A t , B c = B t , C c = C t , D c = D t , and end this transformation;
[0134] If Dmax > Dist, each vertex of the display box of the current video frame needs to be moved towards the corresponding vertex of the target display box at an equal - ratio distance. The moving distance of each vertex is:
[0135] D A ’ = D ist * D A / D m
[0136] D B ’ = D ist * D B / D m
[0137] \ D C ’ = D ist * D C / D m
[0138] \ D D ’ = D ist * D D / D m
[0139] According to the moving distance towards the corresponding vertex of the target display box, the positions of each vertex of the display box of the current video frame can be determined, that is Figure 2 the four vertices A c , B c , C c , D c . A c , B c , C c , D c . Among them, A c is the point on the line segment A o A t that is at a distance of D o ’ from A A , B c is the point on the line segment B o B t that is at a distance of D oThe distance is D B ’s point, C c is the line segment C o C t on C to C o The distance is D C ’s point, D D is the line segment D o D t on D to D o The distance is D D ’s point. A c , B c , C c , D c The positions of A, B, C, and D can be specifically obtained according to the corresponding moving distance and the angle information of the vertex connection lines.
[0140] The AI director method of this embodiment can achieve smooth and stable switching to the target screen regardless of the current screen position. For example, when the AI director needs to switch the screen between the panoramic view of the classroom and the position of the student who needs a close-up, using the AI director method of this embodiment can ensure that the magnification and reduction changes of the screen can be achieved smoothly, so as to avoid stretching and deformation of the screen, causing dizziness and discomfort to the viewers. In addition, the camera switching can be minimized, and the playback frame rate can be effectively increased while ensuring a low delay.
[0141] The above embodiments of the present disclosure provide an AI director system and method that can achieve smooth screen switching, which can replace the traditional manual director method and has at least one of the following advantages:
[0142] Reduce the requirement for the real-time execution of the algorithm;
[0143] Improve the video playback frame rate;
[0144] The camera push and pull are natural and coherent, without screen deformation;
[0145] The screen can be moved during focusing while still maintaining the stability of the screen;
[0146] The display frame is stable, reducing screen jitter.
[0147] In any one or more of the above exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. The computer-readable medium may include a computer-readable storage medium corresponding to a tangible medium such as a data storage medium, or a communication medium including any medium that facilitates transfer of a computer program, such as according to a communication protocol, from one place to another. In this way, the computer-readable medium generally corresponds to a non-transitory tangible computer-readable storage medium or a communication medium such as a signal or a carrier wave. The data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. A computer program product may include a computer-readable medium.
[0148] By way of example, and not limitation, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection may be termed a computer-readable medium. By way of example, if instructions are transmitted using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that the computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather are directed to non-transitory tangible storage media. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, or Blu-ray disc, etc., where disks typically reproduce data magnetically, while discs use lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.
[0149] Instructions may be executed by one or more processors, such as, for example, one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Thus, the term "processor" as used herein may refer to any one of the foregoing structures or any other structure suitable for implementation of the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec. Also, the techniques may be implemented entirely in one or more circuits or logic elements.
[0150] The technical solutions of the embodiments of the present disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or a group of ICs (e.g., a chip set). The various components, modules, or units described in the embodiments of the present disclosure are emphasized to highlight the functional aspects of the devices configured to perform the described techniques, but do not necessarily need to be implemented by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or provided by a collection of interoperating hardware units, including one or more processors as described above, in conjunction with appropriate software and / or firmware.
Claims
1. A method for picture switching during video playback, comprising: Determine that it is necessary to transform the display frame of a video frame in a video frame sequence into a target display frame, wherein the display frame of the video frame is used to indicate the area where the picture to be played is located in the video frame; Set the display frames of the video frames in the video frame sequence frame by frame to complete the transformation, wherein the distance between the display frames set for adjacent video frames is not greater than the maximum allowable moving distance of the display frame, and the distance between two display frames is determined according to the distance between the corresponding positioning points of the two display frames; Determine the picture to be played in the video frame according to the display frames set for the video frames in the video frame sequence and perform playback processing.
2. The picture switching method according to claim 1, characterized in that: The step of setting the display frames of the video frames in the video frame sequence frame by frame to complete the transformation includes: Set the display box of the current video frame as follows: Determine the distance D between the display box of the previous video frame and the target display box m ; In the case where D m is less than or equal to the maximum allowable movement distance D ist of the display box, set the display box of the current video frame as the target display box to complete the transformation; In the case where D m is greater than D ist , set the display box of the current video frame to be at a distance D ist from the display box of the previous video frame and at a distance D1 - D ist from the target display box; In the case where the transformation is not completed, take the next video frame as the current video frame, and continue to set the display frame of the current video frame in the same manner until the transformation is completed.
3. The picture switching method according to claim 2, characterized in that: The maximum allowable moving distance of the display frame is set to a fixed value; or The maximum allowable moving distance of the display frame is equal to the playback time difference between the current video frame and the previous video frame multiplied by the preset maximum moving distance per unit time.
4. The picture switching method according to claim 2, characterized in that: The display frame is a rectangle, the positioning points of the display frame include the four vertices of the rectangle, and the display frame is represented by the position information of the vertices of the display frame in the video frame; The D m is the distance D between the corresponding vertices of the display box of the previous video frame and the target display box A , D B , D C , D D is the maximum value among them, where D A is the distance from A o to A t , D B is the distance from B o to B t , D C is the distance from C o to C t , D D is the distance from D o to D t ; A o , B o , C o , D o are the four vertices of the display box of the previous video frame, and A t , B t , C t , D t are the four vertices of the target display box corresponding to A o , B o , C o , D o respectively.
5. The picture switching method according to claim 4, characterized in that: Setting the display frame of the current video frame to a distance D from the display frame of the previous video frame ist and a distance D1 - D from the target display frame ist for the display frame, including: Determine the distances of the four vertices of the display box of the current video frame relative to the corresponding vertices of the display box of the previous video frame: D A ’ = D ist * D A / D m , \ D B ’ = D ist * D B / D m , \ D C ’ = D ist * D C / D m , \ D D ’ = D ist * D D / D m , where "*" represents the multiplication operation and " / " represents the division operation; Determine the four vertices A c , B c , C c , D c of the display box for the current video frame, where A c is the point on line segment A o A t that is at a distance of D o ' from A A , B c is the point on line segment B o B t that is at a distance of D o ' from B B , C c is the point on line segment C o C t that is at a distance of D o ' from C C , D D is the point on line segment D o D t that is at a distance of D o ' from D D .
6. The picture switching method according to claim 4, characterized in that: After completing the transformation, the method further includes: in the case where the target display frame is not updated, set the display frames of the video frames in the video frame sequence to the target display frame until the target display frame is updated.
7. An AI live production method, comprising: Process an input video frame sequence using a display frame generation algorithm based on machine learning to generate display frames of video frames; When it is determined according to the change of the generated display frames that it is necessary to transform the display frames of the video frames in the video frame sequence into a target display frame, implement picture switching during video playback according to the picture switching method described in any one of claims 1 to 6.
8. The AI live production method according to claim 7, characterized in that: The step of determining according to the change of the generated display frames that it is necessary to transform the display frames of the video frames in the video frame sequence into a target display frame includes: Calculate the intersection over union between the currently generated display frame and the recorded target display frame, and accumulate the number of times that the accumulated intersection over union is less than the set intersection over union threshold: When the accumulated number of times reaches the set number threshold N, update the recorded target display frame to the currently generated display frame, determine that it is necessary to transform the display frames of the video frames in the video frame sequence into the updated target display frame, and restart the accumulation, N≥2.
9. The AI live production method according to claim 7, characterized in that: The display box generation algorithm based on machine learning is executed by a trained display box generation model based on a deep learning network. The deep learning network is trained with a video frame sequence as sample input data and display boxes annotated in the video frames based on the focus events in the samples as target data.
10. The AI director method according to claim 8, characterized in that: The method further includes the following state update processes, either partially or in whole: When starting video playback, it is marked as the initial state; When the recorded target display box is updated and the updated target display box is inside the target display box before the update, it is marked as the zoom-in state; When the recorded target display box is updated and the target display box before the update is inside the updated target display box, it is marked as the zoom-out state; When the recorded target display box is updated and there is no intersection between the target display box before the update and the updated target display box, it is marked as the movement state; When the transformation is completed and there is no need to perform the transformation of the display box again, it is marked as the fixed state.
11. The AI director method according to claim 7, characterized in that: The playback process is executed in a first thread, and the display box generation algorithm is executed in a second thread parallel to the first thread. The first thread and the second thread process based on different copies of the same video frame sequence.
12. A screen switching device, characterized in that, It includes a processor and a memory storing a computer program. Among them, when the processor executes the computer program, it can implement the screen switching method according to any one of claims 1 to 6.
13. An AI director device, characterized in that, It includes a processor and a memory storing a computer program. Among them, when the processor executes the computer program, it can implement the AI director method according to any one of claims 7 to 11.
14. An AI director system, characterized in that, It includes: A video collector, configured to collect and output a video frame sequence; An AI director device, configured to receive the video frame sequence, perform AI direction according to the AI director method according to any one of claims 7 to 11, and output video data obtained after performing playback processing on the picture to be played; A display, configured to display the picture to be played according to the video data.
15. A non-transitory computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement the screen switching method according to any one of claims 1 to 6, or implement the AI director method according to any one of claims 7 to 11.
Citation Information
Patent Citations
Video playing method and device
CN106375772A
Ball game tracking camera shooting method and system
CN113052119A