Method, apparatus, device and storage medium for video switching
By receiving multiple video streams and scene set information, and combining content recognition results with scene information, the output video is determined, thus solving the problem of erroneous switching in the recording and broadcasting switching system and achieving more accurate and personalized video switching.
Patent Information
- Application Number
- CN202211209094.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2042-09-30
AI Technical Summary
Existing recording and switching systems are prone to errors when judging overall video switching, and their automatic switching logic cannot meet the actual needs of various users.
By receiving multiple video streams, the content to be triggered in the videos, and scene set information, the system obtains content recognition results, matches them with the content to be triggered, and combines the basic and related information in the scene information to determine the output video in the current scene. It supports priority order and modification instructions to personalize the switching logic.
It reduces the error rate of video switching, improves the accuracy of video switching, and meets the personalized needs of different users.
Smart Images

Figure CN115604412B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video image processing technology, and more specifically, to a method, apparatus, device, and storage medium for video switching. Background Technology
[0002] People watch various programs in their studies and daily lives, such as classroom lectures, live sports broadcasts, and variety shows. With the application of multimedia technology, people use recording systems to broadcast or record programs simultaneously. During live or recorded broadcasts, different camera positions are needed to adapt to different scenes. Current recording and switching systems are divided into two types: manual control and automatic switching.
[0003] Automatic switching in broadcast systems typically involves automatically recognizing video content and switching to that video immediately when a specific action is triggered. However, the entire recording environment is a unified whole, and relying solely on the trigger action of one video to determine the overall video switching often leads to errors. Furthermore, even within the same environment using identical camera positions, switching requirements can vary, making a static switching logic unsuitable for the diverse needs of different users. Summary of the Invention
[0004] To overcome the shortcomings of incomplete on-site environment assessment leading to erroneous video switching and the inability to change video switching logic as needed, this invention provides a video switching method, apparatus, device, and storage medium. The technical solution adopted by this invention is as follows.
[0005] In a first aspect, the present invention provides a method for video switching, comprising the steps of:
[0006] The system receives multiple video streams, video content to be triggered, and scene set information. The scene set information includes multiple scene information, which includes basic information and associated information. The basic information is formed based on the set of content to be triggered, and the associated information is used to record videos associated with the scene.
[0007] Obtain the content recognition results of the current video;
[0008] The content recognition results are matched with the content to be triggered to obtain the video trigger result;
[0009] The combination of video trigger results is matched with the basic information in the scene information to obtain the currently triggered scene;
[0010] Based on the association information in the scene information, the video associated with the currently triggered scene is determined as the output video.
[0011] In one implementation, when the number of output videos has an upper limit N, the scene set information also includes the priority order between the scenes;
[0012] The video switching method further includes the following steps:
[0013] When the number of output videos is greater than N, N output videos are selected from the determined output videos according to priority order.
[0014] In one implementation, the method further includes the step of: receiving a first modification instruction, and adding, deleting and / or modifying the basic information according to the first modification instruction.
[0015] In one implementation, the method further includes the step of: receiving a second modification instruction and adding and / or deleting content to be triggered in the video according to the second modification instruction.
[0016] In one implementation, the method further includes the step of: receiving a third modification instruction and modifying the priority order according to the third modification instruction.
[0017] In one implementation, at least one scene information includes: associated scene information, which is information about how a scene is associated with other scenes;
[0018] The video switching method further includes the following steps:
[0019] Based on the associated scene information and the associated information in the scene information, the video corresponding to the scene associated with the currently triggered scene is determined as the output video.
[0020] In one implementation, the method further includes the step of: receiving a fourth modification instruction and adding and / or deleting associated scene information in the scene information according to the fourth modification instruction.
[0021] Secondly, the present invention provides a video switching device, comprising:
[0022] The receiving module is used to receive multiple video streams, video content to be triggered, and scene set information. The scene set information includes multiple scene information, which includes basic information and associated information. The basic information is formed based on the set of content to be triggered, and the associated information is used to record videos associated with the scene.
[0023] The acquisition module is used to obtain the content recognition results of the current video;
[0024] The first matching module is used to match the content recognition result with the content to be triggered to obtain the video trigger result;
[0025] The second matching module is used to match the combination of video trigger results with the basic information in the scene information to obtain the currently triggered scene;
[0026] The output module is used to determine the video associated with the currently triggered scene as the output video based on the association information in the scene information.
[0027] Thirdly, the present invention provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method of any of the above embodiments.
[0028] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the program, when executed by a processor, implements the method of any of the above embodiments.
[0029] In this invention, pre-set triggerable content is used for different videos. This content is then combined into different scenes, and each scene is associated with a corresponding video. The recognized content from each video stream is received and matched with the triggerable content to obtain the video triggering result. This allows the system to determine the current scene and identify the video associated with that scene as the output video. This invention reconstructs a scene more closely resembling the current environment through the triggering actions of each video, and outputs the video based on the scene, thus reducing the error rate of video switching. Attached Figure Description
[0030] Figure 1 This is a flowchart of Embodiment 1 of the present invention.
[0031] Figure 2 This is a schematic diagram of the overall structure of Embodiment 2 of the present invention. Detailed Implementation
[0032] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0033] It should be noted that the terms "first, second, ..." used in the embodiments of the present invention are merely used to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, ..." can be interchanged in a specific order or sequence where permissible. It should be understood that the objects distinguished by "first, second, ..." can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein.
[0034] Example 1
[0035] Please see Figure 1 , Figure 1 This is a flowchart illustrating a video switching method according to Embodiment 1 of the present invention. The method includes steps S110, S120, S130, S140, and S150. It should be noted that steps S110-S150 are merely reference numerals used to clearly explain the embodiment and the accompanying drawings. Figure 1 The correspondence is not intended to limit the order of steps in this embodiment.
[0036] Step S110: Receive multiple video streams, video content to be triggered, and scene set information. The scene set information includes multiple scene information, which includes basic information and associated information. The basic information is formed based on the set of content to be triggered, and the associated information is used to record videos associated with the scene.
[0037] The system receives various video feeds. For example, in a classroom setting, received videos could include those tracking the teacher, those tracking students, panoramic videos of the teacher and students, videos of the teacher's computer, and videos of the teacher's handwriting. Each video has its own triggerable content. For instance, in a classroom setting, the triggerable content for the teacher tracking video could be the teacher operating the computer at the podium, the teacher walking at the podium, or the teacher standing at the podium. The triggerable content for the student tracking video could be a single student standing in the student area, or multiple students standing in the student area. The triggerable content for the teacher's computer could be the computer being used. Scene information is the collection of these triggerable content elements. For example, a scene could be created by combining "the teacher operating the computer at the podium" and "the computer is being used," or by combining "the teacher walking at the podium" and "a single student standing in the student area." Alternatively, a scene could be created using only "the teacher standing at the podium" as the triggerable content element. However, at least one scene must contain at least two triggerable content elements. Furthermore, each scene information element is associated with a video feed. For example, the scenes of "the teacher operating the computer at the podium" and "the computer is being used" can be associated with the teacher's computer video feed.
[0038] It should be noted that because different scenarios can freely choose which videos to associate with, a single video may be associated with multiple scenarios.
[0039] It should be noted that it is not necessary to arrange a scene for all the content to be triggered; some of the video's content to be triggered may not be included in any scene.
[0040] It should also be noted that a video can have multiple triggerable content items, and not all videos need to have triggerable content items. For example, in the example above, the whiteboard video does not need to have any triggerable content items set.
[0041] Step S120: Obtain the content recognition result of the current video.
[0042] Obtain the content recognition results from the videos. It's important to note that these results may not be for all videos; they can be specific to the videos containing the content to be triggered. For example, in the classroom teaching scenario, the received videos include: a panoramic view of the teacher, teacher tracking, a panoramic view of the students, student tracking, the teacher's computer screen, and the blackboard (6 videos in total). However, if the videos containing the content to be triggered are only teacher tracking, student tracking, and the teacher's computer screen (3 videos in total), then only the content recognition results from these 3 videos can be obtained.
[0043] For videos requiring specific results, recognition results can be obtained in the following ways. For example, several consecutive frames within a preset time period can be input into a pre-trained action recognition model to determine whether the images contain a preset action. This method only requires the obtained result and does not restrict how this result is obtained.
[0044] It should be noted that the content recognition results of the current video may not be unique; that is, a video may identify multiple pieces of content.
[0045] Step S130: Match the content recognition result with the content to be triggered to obtain the video trigger result.
[0046] Each video with pending trigger content has its own specific trigger content. For a series of videos awaiting triggering, the pending trigger content may include multiple elements. For example, in a teacher tracking video, the pending trigger content might include: the teacher operating a computer at the podium, the teacher walking at the podium, or the teacher standing at the podium. If the current content recognition result of the teacher tracking video is "the teacher is operating a computer at the podium," then the content recognition result matches the trigger content "the teacher is operating a computer at the podium," resulting in the video's trigger result being "the teacher is operating a computer at the podium." Similarly, for other videos, such as a video of a teacher using a computer, the trigger result is "the computer is being used."
[0047] It should be noted that, as mentioned earlier, the content recognition results of a video may not be unique, so it is possible for a video to have multiple matching recognition results, resulting in multiple triggering results.
[0048] Step S140: Match the combination of video trigger results with the basic information in the scene information to obtain the currently triggered scene.
[0049] Because the basic information in the scene information is formed based on the set of content to be triggered, combining the triggering results of the video can match it with the basic information in the scene information to obtain the currently triggered scene. Continuing the previous example, the scene information includes a scene composed of a teacher operating a computer at the podium and the computer performing an operation. If the combination of the triggering results in step S130 matches this scene, then this matched scene is the currently triggered scene.
[0050] It should be noted that since a video stream may include multiple triggering results, or combinations of different triggering results may match different scenarios, the current triggering scenario may also be multiple.
[0051] Step S150: Based on the association information in the scene information, determine the video associated with the currently triggered scene as the output video.
[0052] Each scene in the scene information is associated with a corresponding video. For example, the scene of the teacher operating the computer at the podium and the computer performing operations in the previous example is associated with the teacher's computer video. In this step, we confirm that the associated video is the output video.
[0053] It should be noted that there may be multiple triggering scenarios, so the determined output video may be multiple channels. However, in practice, there may be a limit to the number of output channels, and further confirmation of the final output video is required. Therefore, this method can also be understood as a part of the step of confirming the final output video. Those skilled in the art can then reasonably select the video output based on the determined output video.
[0054] In this method, pre-defined triggerable content is used for different videos. This content is then combined into different scenes, and each scene is associated with a corresponding video. The recognized content from each video stream is received and matched against the triggerable content to obtain the video triggering result. The method then determines the current scene and identifies the video associated with that scene as the output video. This invention reconstructs a scene more closely resembling the current environment through the triggering actions of each video, and outputs the video based on the scene, thus reducing the error rate of video switching.
[0055] In one implementation, when the number of output videos has an upper limit N, the scene set information also includes the priority order between the scenes;
[0056] When the number of output videos is greater than N, N output videos are selected from the determined output videos according to priority order.
[0057] As mentioned earlier, there may be multiple triggered scenarios, resulting in multiple output videos. When the number of determined output videos exceeds the upper limit N, it is necessary to select which of the determined output videos to output. In this implementation, each scenario has a priority order, and the output is selected based on this priority order. Of course, the upper limit is 1, meaning that if only one output is needed, the highest priority video is directly selected for output.
[0058] It should be noted that multiple scenes may be associated with the same video. Therefore, there may be 4 triggered scenes, but only 3 output videos are actually determined. If the upper limit is 3, then no further selection is needed. If the upper limit is 2, then selection is required. In the selection process, the same video associated with multiple scenes is sorted according to the scene with the highest priority to obtain the final output video.
[0059] It should be noted that it is possible that the number of output videos is less than the upper limit N. When the number is less than the upper limit, additional screens can be added or only certain videos can be played. These are all methods that can be formulated by those skilled in the art based on the actual situation. The specific formulation method is not the subject of this implementation.
[0060] In one implementation, the method further includes the step of: receiving a first modification instruction, and adding, deleting and / or modifying the basic information according to the first modification instruction.
[0061] This implementation method is designed to allow users to better personalize their switching logic according to their needs. Users can adjust scene information using the first modification command. If a scene is deemed inappropriate, it can be deleted using the first modification command. If a new scene needs to be added, it can be obtained by combining existing content to be triggered using the first modification command. If the content to be triggered within a scene is deemed inappropriate, the basic information in the scene information can be modified to add or delete content to be triggered within the scene.
[0062] In one implementation, the method further includes the step of: receiving a second modification instruction and adding and / or deleting content to be triggered in the video according to the second modification instruction.
[0063] Users can also modify the content to be triggered in each video stream using the second modification command. Modification generally involves adding or deleting content, while replacing content can be seen as deleting first and then adding. This implementation allows users to select more suitable content to be triggered.
[0064] In one implementation, the method further includes the step of: receiving a third modification instruction and modifying the priority order according to the third modification instruction.
[0065] This implementation method is also designed to allow users to better personalize their switching logic according to their needs. Users can adjust the priority order using a third modification command to adjust the output logic to meet their requirements.
[0066] In one implementation, at least one scene information includes: associated scene information, which is information about other scenes that the scene is associated with;
[0067] The video switching method further includes the following steps:
[0068] Based on the associated scene information and the associated information in the scene information, the video corresponding to the scene associated with the currently triggered scene is determined as the output video.
[0069] In reality, there may be several scenes that are related and need to be played together. However, due to reasons such as video content recognition errors, some scenes may not be triggered. In order to avoid this situation, this implementation method associates some scenes with other scenes, so that the associated scenes can be played together.
[0070] It should be noted that the association in this implementation is unidirectional. For example, A is associated with B, but B does not necessarily have to be associated with A.
[0071] In one implementation, the method further includes the step of: receiving a fourth modification instruction and adding and / or deleting associated scene information in the scene information according to the fourth modification instruction.
[0072] Users can also link different scenes together according to their needs, thereby enabling interaction between scenes.
[0073] Example 2
[0074] Corresponding to the method in Example 1, such as Figure 2 As shown, the present invention also provides a video switching device 5, including: a receiving module 501, an acquisition module 502, a first matching module 503, a second matching module 504, and an output module 505.
[0075] The acquisition module 501 is used to acquire the content recognition results of the current video;
[0076] The first matching module 502 is used to match the content recognition result with the content to be triggered to obtain the trigger result of the video;
[0077] The second matching module 503 is used to match the combination of video triggering results with the basic information in the scene information to obtain the currently triggered scene;
[0078] The output module 504 is used to determine the video associated with the currently triggered scene as the output video based on the association information in the scene information.
[0079] This device pre-sets triggerable content for different videos, combines this content into different scenes, and associates each scene with a corresponding video. It then receives the recognized content from each video stream and matches it with the triggerable content to obtain the video trigger result. From this, it determines the current scene and identifies the video associated with that scene as the output video. This invention reconstructs a scene more closely resembling the current environment through the triggering actions of each video, and outputs the video based on the scene, thus reducing the error rate of video switching.
[0080] In one implementation, when the number of output videos has an upper limit N, the scene set information also includes the priority order between the scenes;
[0081] The output module is also used to select N outputs from the determined output videos according to priority order when the number of output videos is greater than N.
[0082] In one embodiment, the video switching device further includes an instruction module.
[0083] The instruction module is used to receive a first modification instruction and add, delete, and / or modify the basic information according to the first modification instruction.
[0084] In one implementation, the instruction module is further configured to receive a second modification instruction and add and / or delete content to be triggered in the video according to the second modification instruction.
[0085] In one implementation, the instruction module is further configured to receive a third modification instruction and modify the priority order according to the third modification instruction.
[0086] In one implementation, at least one scene information includes: associated scene information, which is information about other scenes that the scene is associated with;
[0087] The output module is also used to determine the video corresponding to the scene associated with the currently triggered scene as the output video based on the associated scene information and the associated information in the scene information.
[0088] In one implementation, the instruction module is further configured to receive a fourth modification instruction and, according to the fourth modification instruction, add and / or delete associated scene information in the scene information.
[0089] Example 3
[0090] This invention also provides a storage medium storing computer instructions that, when executed by a processor, implement the video switching method of any of the above embodiments.
[0091] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, random access memory (RAM), read-only memory (ROM), magnetic disks, or optical disks.
[0092] Alternatively, if the integrated units of this invention are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this invention, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, terminal, or network device, etc.) to execute all or part of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, RAM, ROM, magnetic disks, or optical disks.
[0093] Corresponding to the computer storage medium described above, one embodiment also provides a computer device, which includes a memory, an encoder, and a computer program stored in the memory and executable on the encoder, wherein the encoder executes the program to implement any of the video switching methods described in the above embodiments.
[0094] The aforementioned computer device pre-sets triggerable content for different videos, combines this content into different scenes, associates different scenes with corresponding videos, receives the recognized content from each video stream, matches it with the triggerable content, obtains the video triggering result, determines the current scene, and identifies the video associated with that scene as the output video. This invention, through the triggering actions of each video, reconstructs a scene more closely resembling the current environment, outputting the video based on the scene, thus reducing the error rate of video switching.
[0095] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0096] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for video switching, characterized in that, Including the following steps: The system receives multiple video feeds, the content to be triggered in the videos, and scene set information. Each video with content to be triggered has its own content to be triggered. The scene set information includes multiple scene information, which includes basic information and related information. The basic information is formed based on the set of content to be triggered, and the related information is used to record videos associated with the scene. At least one scene's basic information contains at least two content to be triggered. The videos with content to be triggered include: videos tracking and filming teachers, videos tracking and filming students, panoramic videos of teachers and students, videos of teachers' computers, and videos of whiteboard writing. Obtain the current content recognition results of each video with content to be triggered; The content recognition result of the current video of each video with content to be triggered is matched with the content to be triggered, so as to obtain the trigger result of the video with each content to be triggered. The trigger results of each video with content to be triggered are combined to obtain one or more combined results; Each combination result is paired with the basic information in each scene information. The scene that matches the pairing result is the currently triggered scene. Based on the association information in the scene information, the video associated with each currently triggered scene is determined as the output video; When there is an upper limit N to the number of output videos, the scene set information also includes the priority order between each scene; The video switching method further includes the following steps: When the number of output videos is greater than N, N outputs are selected from the determined output videos according to the priority order among the scenes.
2. The video switching method according to claim 1, characterized in that, It also includes the steps of: receiving a first modification instruction, and adding, deleting and / or modifying the basic information according to the first modification instruction.
3. The video switching method according to claim 1, characterized in that, It also includes the steps of: receiving a second modification instruction and adding and / or deleting content to be triggered in the video according to the second modification instruction.
4. The video switching method according to claim 1, characterized in that, It also includes the step of: receiving a third modification instruction and modifying the priority order according to the third modification instruction.
5. The video switching method according to claim 1, characterized in that, At least one scene information includes: associated scene information, which is information about how a scene relates to other scenes; The video switching method further includes the following steps: Based on the associated scene information and the associated information in the scene information, the video corresponding to the scene associated with the currently triggered scene is determined as the output video.
6. The video switching method according to claim 5, characterized in that, It also includes the step of: receiving a fourth modification instruction, and adding and / or deleting associated scene information in the scene information according to the fourth modification instruction.
7. A video switching device, characterized in that, include: The receiving module is used to receive multiple video streams, the triggerable content of the videos to be triggered, and scene set information. Each video with triggerable content has its own triggerable content. The scene set information includes multiple scene information items, which include basic information and associated information. The basic information is formed based on the set of triggerable content, and the associated information records videos associated with the scene. At least one scene's basic information contains at least two triggerable contents. The videos with triggerable content include: videos tracking and filming teachers, videos tracking and filming students, panoramic videos of teachers and students, videos of teachers' computers, and videos of whiteboard writing. The acquisition module is used to acquire the current content recognition results of each video with content to be triggered; The first matching module is used to match the current video content recognition result of each video with content to be triggered with its respective content to be triggered, so as to obtain the trigger result of each video with content to be triggered. The second matching module is used to combine the triggering results of the videos with content to be triggered to obtain one or more combined results. Each combined result is matched with the basic information in each scene information. The scene that matches the matching result is the scene that is currently triggered. The output module is used to determine the video associated with each currently triggered scene as the output video based on the association information in the scene information; When there is an upper limit N to the number of output videos, the scene set information also includes the priority order between each scene; The output module is also used to select N outputs from the determined output videos according to the priority order among the scenes when the number of output videos is greater than N.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Video playing control method and device and computer readable storage medium
CN112312142A