A smart broadcasting system for dance scenarios

By designing an intelligent directing system for dance scenes and utilizing the collaborative work of multiple modules, the system can identify dance formations and lead dancers and switch camera angles, thus solving the problem of high difficulty in directing dance scenes and improving program production efficiency and viewing experience.

CN119520705BActive Publication Date: 2025-10-31COMMUNICATION UNIVERSITY OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411567344.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-05
Publication Date
2025-10-31
Estimated Expiration
2044-11-05

AI Technical Summary

Technical Problem

Existing technologies lack intelligent directing systems for dance scenes, and the semantic analysis of dance scenes is not sufficiently correlated with directing switching, resulting in high difficulty and resource consumption for manual directing.

Method used

An intelligent directing system for dance scenarios was designed, including a streaming decoding module, a blur detection module, a dance formation change judgment module, a lead dancer recognition module, and a shot switching module. By combining target tracking algorithms and deep learning-based human posture recognition, the system can recognize dance formations and lead dancers and switch shots.

Benefits of technology

It enables intelligent directing of dance scenes, improving program production efficiency and quality, reducing human resource consumption, and providing a better viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119520705B_ABST
    Figure CN119520705B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent directing system for dance scenes, comprising: a streaming decoding module for acquiring and decoding video streams from multiple cameras; a blur detection module for evaluating the clarity of each frame in the video stream; a dance formation change judgment module for analyzing the decoded video frames based on a target tracking algorithm to detect and determine whether the dancers' formations have changed; a lead dancer recognition module for identifying the lead dancer in the decoded video frames using a deep learning-based human pose recognition algorithm; a shot switching module for obtaining a shot transition probability matrix and a shot playback duration matrix based on statistical results of dance program video data, used to guide the shot dwell time and switching channel selection based on shot combination when there are no formation changes and the lead dancer has no special movements; and an intelligent directing logic switching module for automatically executing video switching instructions according to the needs of the dance scene, achieving intelligent directing effects. This invention enables intelligent directing of dance scenes, improving the efficiency and quality of program production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent broadcasting technology, and in particular relates to an intelligent broadcasting system for dance scenarios. Background Technology

[0002] As a crucial component of the gala, dance demands exceptionally high standards from the director, particularly in capturing and presenting visual content. Previously, this task relied heavily on human directors. However, human directing requires a high level of technical and artistic skill, necessitating the coordination of the entire team, execution of the program director's artistic vision, and the use of various auxiliary methods to showcase the performance. Therefore, human directing is challenging and consumes significant human resources.

[0003] With the rapid development of smart technology, people's lifestyles and work methods have become much more convenient. The combination of technology and the art of directing has given rise to intelligent directing systems. Intelligent directing is a video production system that utilizes artificial intelligence technology, aiming to achieve more intelligent, automated, and unmanned television directing switching. Intelligent directing provides a new solution for television program directing of all sizes.

[0004] Currently, the article "Intelligent Director Assists in the Innovation of New Media Programs for the 2021 Spring Festival Gala - A Brief Analysis of the Application of Artificial Intelligence Switching Technology" describes the application of intelligent director in the new media programs of the 2021 Spring Festival Gala. Artificial intelligence technology enabled real-time recognition and intelligent switching, accelerating program production and increasing innovation and appeal. The system is mainly based on facial recognition and key point detection, combined with sound information for comprehensive consideration, to accurately identify the key states of the host, guests, and other personnel guiding the director's switching. In addition, in 2019, Tanwi Mallick focused on Indian classical dance and studied methods for dance video analysis. This method combines musical and kinematic information to extract key postures in dance. To help analyze key elements and techniques in dance works, Stancliffe proposed a new method to address dance analysis capabilities by introducing video annotation technology. Zhang Y and Park HS proposed a semi-supervised learning framework that uses multi-view image streams to train a key point detector given limited labeled data (typically <4%). One algorithm performs cropping, panning, and scaling operations by optimizing the path of the cropping window in the original video, while simultaneously finding salient areas to retain, thus re-editing the entire film within a single shot. For shot selection, Gschwindt M, Camci E, and Bonatti R proposed a learning scheme that leverages the aesthetic characteristics of retrospective shooting to select the ideal shooting mode based on the aesthetic value of the video footage.

[0005] In summary, based on most existing research findings, studies related to intelligent broadcasting are inseparable from the understanding of video content and the extraction of semantic information from video using different methods. However, there are some shortcomings: there is currently a lack of intelligent broadcasting systems for dance scenes, and the semantic analysis of dance scenes usually focuses on the recognition of dance movements, which lacks a strong correlation with the factors that trigger broadcasting switching. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention proposes an intelligent directing system for dance scenarios. Different scenarios have different directing elements, and the switching of group dances focuses on dance formations and dancers' body language. It is necessary to extract semantic features that are strongly correlated with directing switching for dance scenarios and develop an intelligent directing system for dance scenarios to solve the problems existing in the prior art.

[0007] To achieve the above objectives, the present invention provides an intelligent directing system for dance scenarios, comprising:

[0008] The streaming decoding module is used to acquire video streams from multiple cameras and decode them to obtain video frames.

[0009] The blur detection module is used to evaluate whether each frame of the image in the video stream is sharp;

[0010] The dance formation change judgment module is used to analyze the decoded video frames based on the target tracking algorithm to detect and determine whether the dancers' formation has changed.

[0011] The lead dancer recognition module is used to identify lead dancers in the decoded video frames using a deep learning-based human pose recognition algorithm and to mark the appearance of lead dancers.

[0012] The shot switching module, based on the statistical results of the dance program video data, obtains a shot transition probability matrix and a shot playback duration matrix based on shot type. When there is no formation change and the lead dancer has no special movements, it guides the shot dwell time and switching channel selection according to the shot type combination.

[0013] The intelligent broadcasting logic switching module is used to automatically execute video switching instructions according to the needs of the dance scene, so as to achieve intelligent broadcasting effects.

[0014] Preferably, the fuzzy detection module employs the Laplace transform method and the Fourier transform method.

[0015] Preferably, the pull-stream decoding module includes:

[0016] The video stream address acquisition unit is used to acquire video stream address information from different cameras;

[0017] The video frame synchronization unit is used to ensure the temporal consistency of video frames acquired from different cameras.

[0018] Preferably, the dance formation change judgment module includes:

[0019] The target detection and tracking unit is used to detect dancers in video frames based on a target tracking algorithm.

[0020] The formation analysis unit is used to analyze the relative positions and movement trajectories of dancers, and to determine whether the formation has changed by setting dynamic thresholds.

[0021] Preferably, the lead dancer recognition module includes:

[0022] The human pose estimation unit is used to estimate the body key points of dancers using a deep learning-based human pose recognition algorithm to obtain human pose estimation results.

[0023] The lead dancer determination unit is used to identify the lead dancer based on the human posture estimation results and the contextual information of the dance scene.

[0024] Preferably, the shot switching module includes:

[0025] The shot classification unit is used to analyze the shot type of the current video frame and obtain the classification result of which category of "far, full, medium, near, and special" the shot type belongs to;

[0026] Hidden Markov Model Unit based on Shot State is used to model shot transition and dwell time with shot size as state variable, the transition relationship between shot sizes as state transition matrix, and the shot dwell time under each shot size as emission matrix.

[0027] The shot switching decision unit is used to determine whether to switch shots based on the classification results and the shot-based hidden Markov model.

[0028] Preferably, the intelligent broadcast control logic switching module includes:

[0029] The switching logic design unit is used to design switching logic based on the requirements of dance scenarios;

[0030] The switching instruction execution unit is used to execute video switching operations based on the output of the switching logic design unit.

[0031] Preferably, the intelligent broadcasting system adopts a combination of multi-process and multi-threaded approaches.

[0032] Compared with the prior art, the present invention has the following advantages and technical effects:

[0033] This invention proposes an intelligent directing system for dance scenarios. First, the streaming decoding module acquires and decodes multi-camera video streams, providing the raw video data needed for subsequent processing and ensuring the system can handle signals from different cameras. Second, the dance formation change judgment module uses a target tracking algorithm to track dancers in the video and determines whether the formation has changed by setting a dynamic threshold, thus switching accordingly when the formation changes. Next, the lead dancer recognition module uses a deep learning-based human posture recognition algorithm to identify the lead dancer, ensuring their dance posture is prominently displayed when they appear. Then, the shot switching module, based on statistical results of the dance program video data, obtains a shot transition probability matrix and a shot playback duration matrix based on shot size to guide shot dwell time and automatic switching. Finally, the intelligent directing logic switching module designs switching logic and generates switching instructions based on general directing principles for dance scenarios, achieving automatic intelligent directing switching and reducing the difficulty of directing.

[0034] This invention, based on the collaborative work of multiple modules, enables intelligent directing of dance scenes, improving the efficiency and quality of program production, reducing the consumption of human resources, and providing viewers with a better viewing experience. Attached Figure Description

[0035] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0036] Figure 1 This is a system schematic diagram according to an embodiment of the present invention;

[0037] Figure 2 This is a flowchart illustrating the system switching logic of an embodiment of the present invention. Detailed Implementation

[0038] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0039] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0040] Example 1

[0041] like Figure 1 As shown, this embodiment provides an intelligent directing system for dance scenarios, including:

[0042] The streaming decoding module is used to acquire video streams from multiple cameras and decode them to obtain video frames.

[0043] The blur detection module is used to evaluate whether each frame of the image in the video stream is sharp;

[0044] The dance formation change judgment module is used to analyze the decoded video frames based on the target tracking algorithm to detect and determine whether the dancers' formation has changed.

[0045] The lead dancer recognition module is used to identify lead dancers in the decoded video frames using a deep learning-based human pose recognition algorithm and to mark the appearance of lead dancers.

[0046] The shot switching module, based on the statistical results of the dance program video data, obtains a shot transition probability matrix and a shot playback duration matrix based on shot type. When there is no formation change and the lead dancer has no special movements, it guides the shot dwell time and switching channel selection according to the shot type combination.

[0047] The intelligent broadcasting logic switching module is used to automatically execute video switching instructions according to the needs of the dance scene, so as to achieve intelligent broadcasting effects.

[0048] In this embodiment, taking a three-camera setup as an example, feature extraction is performed on the decoded data, including blur detection, dance formation change judgment, group dance lead dancer identification, and scene classification. By analyzing the camera layout in the venue and combining it with test materials, it can be found that the first camera's view is mainly a panoramic view, the second camera's view is a distant view, and the third camera's view is mainly a panoramic view. Switching judgment is performed on a frame-by-frame basis. If the image is determined to be blurry, the switching qualification for that signal is cancelled; otherwise, subsequent switching decisions are executed. If a formation change occurs, a switching command is triggered, and the image is switched to a specific channel. If a lead dancer appears, the image is switched to a specific channel; otherwise, switching is performed according to the scene type. The system switching logic flowchart is as follows. Figure 2 As shown.

[0049] Furthermore, the fuzzy detection module employs the Laplace transform and Fourier transform methods.

[0050] Specifically, the blur detection module employs Laplace transform and Fourier transform methods to identify blur or distortion in images. For systems like intelligent broadcasting that require significant computational resources, algorithm selection must prioritize simplicity and speed while ensuring availability to guarantee the real-time performance of the entire system. Both methods are relatively simple and computationally efficient image processing techniques, especially the Fast Fourier Transform (FFT) algorithm, which efficiently processes the frequency domain information of images. This blur detection method offers advantages such as simplicity, speed, no training required, strong interpretability, good adaptability to specific blur types, visualization, and adjustable parameters.

[0051] Furthermore, the pull-stream decoding module includes:

[0052] The video stream address acquisition unit is used to acquire video stream address information from different cameras;

[0053] The video frame synchronization unit is used to ensure the temporal consistency of video frames acquired from different cameras.

[0054] Specifically, when building the simulation system, front-end data transmission can be achieved using virtual machine streaming, employing Ubuntu terminal FFmpeg commands to push multiple video streams to simulate real-time signals. In practical applications, signals from camera streams can be directly retrieved. During the decoding process, multiple threads are created to continuously read each video stream frame. After obtaining one frame, the image is converted into an RGB24 format array, and a data structure containing the image and metadata is constructed. The image data is then transmitted to various feature extraction modules for analysis.

[0055] Furthermore, the dance formation change judgment module includes:

[0056] The target detection and tracking unit is used to detect dancers in video frames based on a target tracking algorithm.

[0057] The formation analysis unit is used to analyze the relative positions and movement trajectories of dancers, and to determine whether the formation has changed by setting dynamic thresholds.

[0058] Specifically, the target tracking module in the system performs feature extraction to determine formation changes in group dance scenarios. It completes tasks including target detection, target tracking, average displacement calculation, dynamic threshold updating, and formation change judgment. The feature extraction algorithm is called as a function to analyze and process video frames, obtaining the data needed for switching decisions. In the function encapsulation section, the formation change judgment part is encapsulated as a function. The function input data is the video image data obtained by the decoding module. The video frames decoded by the system framework are fed into the target tracking algorithm. After image preprocessing, model inference and detection, and data analysis, the formation change judgment result is obtained. The system framework sets up interface functions to initialize the target tracking module, configure experimental models, extract image data from the input image, complete target detection, tracking, and formation change judgment tasks, store and transmit image and detection data, receive and process displacement data, and define a queue to store displacement data from adjacent frames. As the video plays, the video frame image data is updated, target position information is updated, detection results are updated, and judgment results are updated. Switching decisions are made and executed based on the detection results.

[0059] Furthermore, the lead dancer recognition module includes:

[0060] The human pose estimation unit is used to estimate the body key points of dancers using a deep learning-based human pose recognition algorithm to obtain human pose estimation results.

[0061] The lead dancer determination unit is used to identify the lead dancer based on the human posture estimation results and the contextual information of the dance scene.

[0062] Specifically, the human posture recognition module in the system extracts human limb movement information and identifies the lead dancer in a group dance scene based on the extracted posture information. It also implements system calls through function encapsulation, analyzing and processing video frames to obtain the data needed for switching decisions. To facilitate communication between the system framework and the human posture recognition module, the lead dancer recognition part is encapsulated as a function. This encapsulated function completes tasks including target detection, keypoint detection, motion amplitude judgment, counting the number of people with different motion amplitudes, and judging the lead dancer's status. The function input data is the video image data obtained from the decoding module. The system implementation first initializes the model in the system framework. As the video stream is decoded, video frame data is obtained. The images are processed and analyzed through function interfaces, completing image processing, model inference, keypoint information processing, motion amplitude judgment, iterative detection of multiple targets in a single frame, counting the number of people with different motion amplitudes for multiple targets, and judging the lead dancer's status. As the video playback updates the video frame image data, the judgment results are updated, and a switching decision is made and executed based on the detection results.

[0063] Furthermore, the shot switching module includes:

[0064] The shot classification unit is used to analyze the shot type of the current video frame and obtain the classification result of which category of "far, full, medium, near, and special" the shot type belongs to;

[0065] Hidden Markov Model Unit based on Shot State is used to model shot transition and dwell time with shot size as state variable, the transition relationship between shot sizes as state transition matrix, and the shot dwell time under each shot size as emission matrix.

[0066] The shot switching decision unit is used to determine whether to switch shots based on the classification results and the shot-based hidden Markov model.

[0067] Specifically, in dance program production, the appropriate combination of different shot types is crucial. A rich and reasonable combination of these shot types not only enhances the program's expressiveness but also better conveys the beauty and emotion of the dance. The alternation of different shot types creates a layered visual effect, avoiding visual fatigue caused by a single shot type and helping to meet the diverse needs of the audience. To achieve switching between different shot types, the shot type classification module in the feature extraction section first classifies the video footage into shot types as the basis for switching. The switching is performed according to a shot type switching model that includes a shot type transition probability matrix and a dwell time emission matrix. There are six possible shot types for the currently playing frame, and six possible shot types for the next playing frame. The next shot type is determined by the statistically obtained shot type transition probability matrix for the dance scene. The dwell time for each shot type is divided into 11 levels, and the playback time is determined by the dwell time emission probability matrix. In summary, the shot type of the next shot is selected according to the shot type transition probability, and the duration of each shot is selected according to the dwell time emission matrix.

[0068] Furthermore, the intelligent broadcast control logic switching module includes:

[0069] The switching logic design unit is used to design switching logic based on the requirements of dance scenarios;

[0070] The switching instruction execution unit is used to execute video switching operations based on the output of the switching logic design unit.

[0071] Specifically, the intelligent director switching module determines whether a switch is needed based on the data detected by the feature extraction part and the director switching logic. If the switching conditions are met, the switching operation is executed. The system framework receives information including video stream address, scene type, fuzzy identifier, formation change judgment result, and lead dancer recognition result. It determines whether a switch is needed and, based on the switching model, whether the switching conditions are met. This includes switching the broadcast to a long shot when a formation change is detected and switching to a wide shot when a lead dancer is detected. If a switch is detected, a switching command is sent. The switching is implemented by constructing a command object containing playback-related information, including the specified RTMP address, the timecode indicating the playback time point, the display timestamp (PTS) of the video frame, and the decoding timestamp (DTS) of the video frame. This command object is then sent to the server via WebSocket. Upon receiving the command, the server executes the corresponding playback operation.

[0072] Furthermore, the intelligent broadcasting system adopts a combination of multi-process and multi-threaded approaches.

[0073] In this embodiment, three processes are started to handle three tasks in parallel according to the system's functional requirements, respectively executing pull-stream decoding, feature detection, and switching decision. Since program recording typically involves multiple signal inputs, pull-stream decoding and feature detection employ multi-threading to process multiple video data streams in parallel. During the system testing phase, taking three-channel data as an example, three threads were started. Data interaction between processes is conducted through ZMQ (ZeroMQ), an open-source messaging library that uses ZMQ sockets for communication to complete data transmission tasks. The production tool first pulls and decodes the multi-camera signals to obtain image data, then extracts features from the image data, and finally makes switching decisions based on the extracted underlying features to achieve the switching.

[0074] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An intelligent broadcasting system for dance scenarios, characterized in that, Includes the following steps: The streaming decoding module is used to acquire video streams from multiple cameras and decode them to obtain video frames. The blur detection module is used to evaluate whether each frame of the video stream is clear; if it is determined to be blurry, the camera signal switching qualification is cancelled; if it is not blurry, the subsequent switching decision is continued. The dance formation change judgment module is used to analyze the decoded video frames based on the target tracking algorithm to detect and determine whether the dancers' formation has changed. The lead dancer recognition module is used to identify lead dancers in the decoded video frames using a deep learning-based human pose recognition algorithm and to mark the appearance of lead dancers. The shot switching module, based on the statistical results of the dance program video data, obtains a shot transition probability matrix and a shot playback duration matrix based on shot type. When there is no formation change and the lead dancer has no special movements, it guides the shot dwell time and switching channel selection according to the shot type combination. The shot switching module includes: The shot classification unit is used to analyze the shot type of the current video frame and obtain the classification result of which category of "far, full, medium, near, or special" the shot type belongs to; Hidden Markov Model Unit based on Shot State is used to model shot transition and dwell time with shot size as state variable, the transition relationship between shot sizes as state transition matrix, and the shot dwell time under each shot size as emission matrix. A shot switching decision unit is used to determine whether to switch shots based on the classification results and a shot-based hidden Markov model. The intelligent director logic switching module is used to determine whether a switch is needed based on the data detected by the feature extraction part and the director switching logic. If the switching conditions are met, the switching operation is executed. The conditions for determining whether the switch is met include switching the broadcast to a distant view when a change in formation is detected, and switching the view to a panoramic view when a lead dancer is detected.

2. The intelligent broadcasting system for dance scenarios according to claim 1, characterized in that, The fuzzy detection module employs the Laplace transform and Fourier transform methods.

3. The intelligent broadcasting system for dance scenarios according to claim 1, characterized in that, The pull-stream decoding module includes: The video stream address acquisition unit is used to acquire video stream address information from different cameras; The video frame synchronization unit is used to ensure the temporal consistency of video frames acquired from different cameras.

4. The intelligent broadcasting system for dance scenarios according to claim 1, characterized in that, The dance formation change judgment module includes: The target detection and tracking unit is used to detect dancers in video frames based on a target tracking algorithm. The formation analysis unit is used to analyze the relative positions and movement trajectories of dancers, and to determine whether the formation has changed by setting dynamic thresholds.

5. The intelligent broadcasting system for dance scenarios according to claim 1, characterized in that, The lead dancer recognition module includes: The human pose estimation unit is used to estimate the body key points of dancers using a deep learning-based human pose recognition algorithm to obtain human pose estimation results. The lead dancer determination unit is used to identify the lead dancer based on the human posture estimation results and the contextual information of the dance scene.

6. The intelligent broadcasting system for dance scenarios according to claim 1, characterized in that, The intelligent broadcast control logic switching module includes: The switching logic design unit is used to design switching logic based on the requirements of dance scenarios; The switching instruction execution unit is used to execute video switching operations based on the output of the switching logic design unit.

7. The intelligent broadcasting system for dance scenarios according to claim 1, characterized in that, The intelligent broadcasting system adopts a combination of multi-process and multi-threaded approaches.

Citation Information

Patent Citations

  • Intelligent director method for sports events

    CN110944123A

  • Intelligent director switching method based on multi-person posture estimation action amplitude detection

    CN116721468A