Soft hockey panoramic video automatic program directing method, equipment, medium and product

Through the improved YOLOv11 model and smooth transition mechanism, automatic director of soft hockey panoramic video is realized, solving the problems of high cost and complexity of traditional directors, improving the efficiency and quality of directors, and suitable for high-quality dissemination of soft hockey events.

CN120568083APending Publication Date: 2025-08-29INNER MONGOLIA UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510941670.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

Traditional soft hockey event directors rely on manual operations, which are costly and complex, limiting the large-scale promotion and popularization of events.

Method used

The improved YOLOv11 model is used to build an object detection model, combining the ConvNeXtV2 backbone network, the C3k2 module of the CA attention mechanism and the dynamic upsampling operator to automatically direct the panoramic video, and using the smooth transition mechanism of the historical frame queue and coordinate list to realize intelligent recognition and screen director of players and balls.

Benefits of technology

It realizes efficient director without manual intervention, ensures picture continuity and comfort, automatically adjusts focus, reduces the burden of manual director, reduces costs, and improves director efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120568083A_ABST
    Figure CN120568083A_ABST
Patent Text Reader

Abstract

The invention discloses a soft hockey panoramic video automatic director method and device, a medium and a product, and relates to the field of video director, and the method comprises the steps: building a target detection model based on an improved YOLOv11 model; the target detection model is used for determining a target detection result according to the soft hockey panoramic image; according to the improved YOLOv11 model, on the basis of the YOLOv11 model, a backbone network in which ConvNeXtV2 is introduced is adopted, a C3k2 module in which a CA attention mechanism is introduced is adopted, and an up-sampling mode of a dynamic up-sampling operator is adopted. Acquiring a soft hockey panoramic video stream; and according to each frame of picture in the soft hockey panoramic video stream, a smooth transition mechanism of a historical frame queue and a coordinate list is adopted, and the panoramic video is directed based on the target detection model. According to the invention, the video director efficiency can be improved on the basis of reducing the cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video directing, and in particular to a method, device, medium and product for automatic directing of a panoramic video of soft hockey. Background Art

[0002] With the booming development of softball hockey and the increasing number of competitions, the demand for high-quality broadcasts is becoming increasingly urgent. Video directing has become a core method for broadcasting events. Traditional event directing typically relies on the strategic placement of multiple cameras and the meticulous collaboration of a professional team of directors to capture the highlights of the game. However, traditional directing methods face high costs and complex operational processes, limiting the large-scale promotion and popularization of events.

[0003] In addition, how to significantly reduce the high cost required for traditional directing and greatly improve the efficiency of directing without human intervention is an urgent problem that needs to be solved. Summary of the Invention

[0004] The purpose of this application is to provide a method, device, medium and product for automatic directing of panoramic videos of soft hockey, which can improve the efficiency of video directing while reducing costs.

[0005] To achieve the above objectives, this application provides the following solutions:

[0006] In a first aspect, the present application provides a method for automatically directing a panoramic video of a soft hockey game, the method comprising:

[0007] Based on an improved YOLOv11 model, a target detection model is constructed; the target detection model is used to determine target detection results based on a panoramic image of a soft hockey game; the target detection results include: players, goalkeepers, balls, and referees; the improved YOLOv11 model is based on the YOLOv11 model and uses a backbone network that introduces ConvNeXtV2, a C3k2 module that introduces a CA attention mechanism, and an upsampling method using a dynamic upsampling operator;

[0008] Get the full video stream of soft hockey;

[0009] According to each frame in the soft hockey panoramic video stream, a smooth transition mechanism of historical frame queue and coordinate list is adopted, and the panoramic video is directed based on the target detection model.

[0010] Optionally, the target detection model is constructed based on the improved YOLOv11 model, specifically including:

[0011] Obtain panoramic videos of soft hockey at different time periods;

[0012] Based on the soft hockey panoramic videos of different time periods, the OpenCV library technology is used to extract the soft hockey panoramic images frame by frame;

[0013] Performing data preprocessing on the soft hockey panoramic image; the data preprocessing includes: image cropping and image screening;

[0014] Perform object annotation on the preprocessed soft hockey panoramic images and construct a dataset;

[0015] According to the dataset, an object detection model is built based on the improved YOLOv11 model.

[0016] Optionally, the data preprocessing of the soft hockey panoramic image specifically includes:

[0017] The panorama image of the soft hockey is cropped based on a sliding window method to obtain an image of a set size;

[0018] According to the image of the set size, the YOLOv8n pre-trained model is used to perform image screening; the YOLOv8n pre-trained model is used to identify whether there is an object in the image of the set size.

[0019] Optionally, the ConvNextV2 Block module in ConvNeXtV2 adopts a combination of large kernel convolution and inverse bottleneck structure.

[0020] Optionally, the method of directing the panoramic video based on each frame in the soft hockey panoramic video stream using a smooth transition mechanism of a historical frame queue and a coordinate list and based on a target detection model specifically includes:

[0021] When directing based on spherical coordinates, determine the initial center position pre_position (x0, y0) and the corresponding initial coordinates;

[0022] Get the coordinate list position and the history frame queue history_frame; and set the detection threshold n and the output video threshold t; the coordinate list position is initially empty; the detection threshold n is used to indicate that detection is performed once every n frames; the output video threshold t is used to indicate that the images in the history frame queue history_frame are output every t frames;

[0023] Each frame of the soft hockey panoramic image in the soft hockey panoramic video stream is added to the history frame queue history_frame; when the value count of a frame of the soft hockey panoramic image satisfies count%n=0, the object detection model is used to detect the ball and the coordinates of the detected ball are added to the coordinate list position;

[0024] When the length of the history frame queue history_frame reaches the output video threshold t, according to the tail coordinate (x t ,y t ) determining a displacement amount and a movement direction of the center coordinates of the i-th frame of the soft hockey panoramic image compared to the (i-1)-th frame of the soft hockey panoramic image, and determining the center coordinates of the i-th frame of the soft hockey panoramic image;

[0025] Loop through each frame of the soft hockey panorama image in the history frame queue history_frame, and output the image with the center of the obtained soft hockey panorama image as the center; after the output is completed, update the initial position coordinate pre_position to (x t ,y t ), clear the coordinate list position and the history frame queue history_frame; return to the step of adding each frame of the soft hockey panoramic image of the soft hockey panoramic video stream to the history frame queue history_frame until the output of the soft hockey panoramic video stream is completed;

[0026] According to the output images of the current soft hockey panoramic video stream, the broadcast is directed based on the ball coordinates in sequence.

[0027] Optionally, the method of directing the panoramic video based on each frame in the soft hockey panoramic video stream using a smooth transition mechanism of a historical frame queue and a coordinate list and based on a target detection model specifically includes:

[0028] When directing based on player coordinates, determine the initial center position pre_position(x0,y0) and the corresponding initial coordinates;

[0029] Get the coordinate list positions and the history frame queue history_frame; and set the detection threshold n and output video threshold t; the coordinate list positions is initially empty; the detection threshold n is used to indicate that detection is performed once every n frames; the output video threshold t is used to indicate that the images in the history frame queue history_frame are output every t frames;

[0030] Each frame of the soft hockey panoramic image in the soft hockey panoramic video stream is added to the history frame queue history_frame; when the value count of a frame of the soft hockey panoramic image satisfies count%n=0, the target detection model is used to detect the player and add the detected player's coordinates to the coordinate list person_positions; a K-Means clustering calculation is performed on the coordinate list person_positions to determine the center coordinates of the cluster where the players are relatively concentrated, and the center coordinates are added to the coordinate list positions; the coordinate list person_positions is initially empty;

[0031] When the length of the history frame queue history_frame reaches the output video threshold t, according to the tail coordinate (x t ,y t ) determining a displacement amount and a movement direction of the center coordinates of the i-th frame of the soft hockey panoramic image compared to the (i-1)-th frame of the soft hockey panoramic image, and determining the center coordinates of the i-th frame of the soft hockey panoramic image;

[0032] Loop through each frame of the soft hockey panorama image in the history frame queue history_frame, and output the image with the center of the obtained soft hockey panorama image as the center; after the output is completed, update the initial position coordinate pre_position to (x t ,y t ), clear the coordinate list position and the history frame queue history_frame; return to the step of adding each frame of the soft hockey panoramic image of the soft hockey panoramic video stream to the history frame queue history_frame until the output of the soft hockey panoramic video stream is completed;

[0033] According to the output images of the current soft hockey panoramic video stream, the broadcast is directed based on the player coordinates in sequence.

[0034] In a second aspect, the present application provides a soft hockey panoramic video automatic broadcasting device, the soft hockey panoramic video automatic broadcasting device comprising:

[0035] A model construction module is configured to construct a target detection model based on an improved YOLOv11 model; the target detection model is configured to determine target detection results based on a panoramic image of a soft hockey game; the target detection results include: players, goalkeepers, balls, and referees; the improved YOLOv11 model is based on the YOLOv11 model and uses a backbone network that incorporates ConvNeXtV2, a C3k2 module that incorporates a CA attention mechanism, and an upsampling method that uses a dynamic upsampling operator;

[0036] A video stream acquisition module is used to acquire a panoramic video stream of soft hockey;

[0037] The intelligent broadcast guidance module is used to guide the panoramic video based on each frame in the soft hockey panoramic video stream, using a smooth transition mechanism of historical frame queues and coordinate lists and based on the target detection model.

[0038] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the automatic directing method for panoramic video of soft hockey.

[0039] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for automatically directing a panoramic video of soft hockey.

[0040] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the method for automatically directing the panoramic video of soft hockey.

[0041] According to the specific embodiments provided in this application, this application has the following technical effects:

[0042] The present application provides a method, device, medium and product for automatic directing of panoramic videos of soft hockey. Based on the improved YOLOv11 model, a target detection model is constructed to accurately identify the target; according to each frame in the panoramic video stream of soft hockey, a smooth transition mechanism of historical frame queue and coordinate list is adopted, and the panoramic video is directed based on the target detection model; the continuity of the directed picture is ensured, and frequent jumps caused by detection errors or players moving too fast are avoided, thereby improving the viewing comfort. Through the smooth transition of historical frame queue and coordinate list, even if a frame fails to successfully detect the ball or player, it can still smoothly transition to the next frame according to the previous historical coordinates, ensuring that the picture does not lose focus or lose key information. The present application can automatically judge and adjust the focus of the picture without manual intervention; whether it is based on the movement of the ball or the distribution of the players, it can react in a short time, greatly improving the efficiency and quality of directing, and reducing the burden of manual directing. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0044] Figure 1 This is a flowchart of a method for automatically directing a panoramic video of soft hockey in one embodiment of the present application;

[0045] Figure 2 Schematic diagram of the YOLO11 model structure;

[0046] Figure 3 This is a schematic diagram of the C3k2_CA module;

[0047] Figure 4 This is a schematic diagram of the Dysample upsampling process;

[0048] Figure 5 Generate a flow chart for the Dysample sampling point;

[0049] Figure 6 Schematic diagram of the structure of the landmark detection model CCD-YOLO11. DETAILED DESCRIPTION

[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0051] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0052] In an exemplary embodiment, Figure 1 As shown, a method for automatically directing a panoramic video of soft hockey is provided, which includes the following steps S101 to S103.

[0053] S101, constructing a target detection model based on an improved YOLOv11 model; the target detection model is used to determine target detection results based on a panoramic image of a soft hockey game; the target detection results include: players, goalkeepers, balls, and referees; the improved YOLOv11 model is based on the YOLOv11 model and adopts a backbone network introduced in ConvNeXtV2, a C3k2 module introduced in a CA attention mechanism, and an upsampling method using a dynamic upsampling operator; the ConvNextV2 Block module in ConvNeXtV2 adopts a combination of large kernel convolution and inverse bottleneck structure.

[0054] S101 specifically includes:

[0055] S11, obtaining panoramic videos of soft hockey matches at different time periods; during an offline soft hockey league, a panoramic camera was set up at the competition venue to collect game data. The camera recorded the entire game process and captured panoramic video data at different time periods. To ensure video capture quality, a panoramic camera with a resolution of 8160*3616 and a frame rate of 24 was used for shooting;

[0056] S12, extracting panoramic images of the soft hockey puck by frame using the OpenCV library technology based on the soft hockey puck panoramic videos of different time periods; setting the sampling interval to 15 frames, and saving each extracted image;

[0057] S13, performing data preprocessing on the soft hockey panoramic image; the data preprocessing includes: image cropping and image screening;

[0058] S13 specifically includes:

[0059] S1, cropping the soft hockey panoramic image based on a sliding window method to obtain an image of a set size;

[0060] Since the size of the image extracted by frame extraction is 8160*3616, the high-resolution image contains a lot of pixels, and directly inputting it into the model will consume a lot of memory and computing resources. High-resolution images may contain redundant information, and the model may focus too much on details, resulting in overfitting, performing well on the training set but poorly on the test set. Therefore, the image is cropped, the image resolution is adjusted, the input size is reduced, the training efficiency is improved, overfitting is avoided, and computing resources are effectively utilized. In order to crop a large image of size 8160*3616 into multiple small images of size 1920*1080, this can be achieved by the sliding window method. The above method usually crops on the original image by sliding a fixed-size window, moving a certain step size each time until the entire large image is covered. The specific formula is as follows:

[0061]

[0062] Among them, W orig ,H orig Represents the width and height of the original image, W crop ,H crop Indicates the width and height of the cropped small image, N x , N y represents the number of horizontal and vertical sliding, S y , S x Represents the vertical and horizontal sliding steps respectively.

[0063] S2, using a YOLOv8n pre-trained model to perform image screening based on an image of a set size; the YOLOv8n pre-trained model is used to identify whether an object exists in the image of the set size.

[0064] S14, performing target annotation on the pre-processed soft hockey panoramic image and constructing a dataset;

[0065] There are four target categories: player, goalkeeper, ball, and referee. To ensure that subsequent directing is not disrupted by off-field objects, the LabelImg annotation tool focuses on on-field objects and avoids including off-field objects in the annotation process. For the player analogy, the stick held by the player is a key feature for distinguishing different types of objects (such as players and referees). In soft hockey, the player's movement, posture, and stick position are key to player identification. To ensure accurate target annotation, the stick held by the player must be specifically considered and included as part of the annotation frame. The player's movement often changes the stick's angle and position. Therefore, the annotation frame should not only encompass the player's body but also include the area where the player's hand grips the stick and the extended portion of the stick. The stick is a crucial component of the player, and its shape, length, and position are key factors in distinguishing players from other objects.

[0066] After labeling is completed, the training set, validation set, and test set are divided into 8:1:1 ratios.

[0067] S15, based on the dataset and the improved YOLOv11 model, build a target detection model.

[0068] YOLO11 is the latest target detection model released by the Ultralytics team. It aims to maintain high accuracy while further improving the model's computational efficiency and flexibility. YOLO11 not only inherits the excellent foundation of the previous YOLO version, but also integrates a variety of advanced technical concepts on this basis, and innovatively introduces new functions and improved modules to further optimize detection performance and adaptability. The structure of YOLO11 is as follows: Figure 2 shown.

[0069] YOLO11 mainly uses the C3k2 module for feature extraction and feature fusion at different stages, and provides a C3k option, allowing users to decide whether to enable stacking of multiple C3k modules. When the C3k option is turned on, the C3k2 module will use the C3k module stack to complete feature extraction. When it is turned off, the C3k2 module will be similar to the C2f module in YOLOv8, using Bottleneck to complete feature extraction. The C3k2 module is an improvement to the cross-stage bottleneck structure. After a convolution, it will split the feature map and use multiple 3*3 convolution kernels to complete feature extraction on independent feature maps, and complete feature fusion through convolution in subsequent stages. This method of splitting first and then convolution can effectively reduce computational costs while maintaining good feature expression capabilities. In addition, the C3k2 module also provides a stacking parameter N to control the number of stacking times, so as to flexibly adjust the network depth according to different tasks.

[0070] Compared to previous versions, YOLO11 introduces the C2PSA module in its backbone network. This module is a cross-stage partial spatial attention module composed of a stack of convolutional blocks and multiple partial spatial attention (PSA) modules. The convolutional blocks are responsible for preliminary processing of input features, while the multiple PSA modules continuously extract local spatial information, enhance the spatial correlation in the feature map, and improve the model's focus on key local areas. The design of this module can further strengthen the model's ability to focus on local features, to a certain extent helping the model to more accurately complete target positioning and classification.

[0071] The backbone network is the core component of the object detection model, responsible for extracting features from the input image. The backbone network's feature extraction capabilities directly determine the model's ultimate detection performance. The backbone network structure used by YOLO11 is similar to CSPDarknet, consisting of Conv and C3k2 modules, primarily constructed through stacking 3x3 convolutional kernels. However, in soft hockey scenarios, when a player is obstructed, the 3x3 convolution is limited in its ability to capture player features due to its smaller receptive field. Key features may not be extracted, leading to missed detections and thus compromising the model's detection performance. To address this issue, ConvNeXtV2 was introduced to improve YOLO11n's backbone network, enhancing the model's feature extraction capabilities and enabling the model to maintain high recognition performance despite dynamic changes, occlusions, and overlaps.

[0072] ConvNeXtV2 is a new convolutional neural network architecture, a further improvement of ConvNeXt. It inherits and expands the original design concept, incorporates more advanced technologies, and designs a new ConvNeXtV2 block to enhance the network's feature extraction and global perception capabilities. The structure of the ConvNextV2 network is shown in Table 1. The entire network is mainly composed of Conv and multiple ConvNextV2 blocks stacked together.

[0073] Table 1

[0074] Input Operator Depth OutChannel Stride 640*640*3 Conv 1 40 4 160*160*40 ConvNeXtV2Block 2 40 1 160*160*40 Conv 1 80 2 80*80*80 ConvNeXtV2Block 2 80 1 80*80*80 Conv 1 160 2 40*40*160 ConvNeXtV2Block 6 160 1 40*40*160 Conv 1 320 2 20*20*320 ConvNeXtV2Block 2 320 1

[0075] The ConvNextV2 Block is the core module of ConvNexV2. During feature extraction, this module utilizes a unique approach, combining large-kernel convolution with an inverse bottleneck architecture. Large-kernel convolution effectively compensates for feature information loss due to insufficient receptive field and exhibits improved adaptability to convolution operations at varying scales. The inverse bottleneck architecture rapidly transforms spatial dimensions and semantically compresses features extracted by large-kernel convolution, further improving computational speed.

[0076] The C3k2 module is a key feature extraction and fusion component in YOLO11. It is used extensively in both the backbone and neck networks. This module combines dynamic stacking with channel separation strategies to efficiently extract and fusion features. In actual matches, players often engage in rapid movements in competition for the ball, and the relationship between background and target constantly changes, leading to a closer correlation between global and local features and placing higher demands on the model's feature fusion capabilities. However, the C3k2 module's relatively fixed feature extraction and fusion methods make it difficult to fully capture deep spatiotemporal dependencies, limiting the flexible fusion of global and local features and thus impacting detection performance. To address this issue, the C3k2 module has been improved based on the CA attention mechanism to enhance its feature fusion capabilities.

[0077] The CA attention mechanism uses global pooling to encode input features and splits the encoded features into two directions: one for capturing long-range dependencies between features and the other for preserving positional information. This design not only improves the network's ability to perceive global information but also maintains the accuracy of local positional information, thereby enhancing the network's ability to express features in complex tasks.

[0078] The long-range dependency information and position information provided by the CA attention mechanism can complement the global and local features provided by the large-core convolutional backbone network. Long-range dependencies can help the model perceive long-range features, avoid focusing only on local features, and thus improve the overall understanding of the scene. Position information can help the model accurately perceive the target position and prevent the loss of local position information due to excessive focus on global features. This paper integrates the CA attention mechanism into the Bottleneck structure and designs the C3k2_CA module. The structure is as follows: Figure 3 This improved module can more effectively utilize the global perception and location information provided by the CA attention mechanism, increase the focus on key areas, more fully integrate global and local features, and improve the model's target positioning ability in motion scenes.

[0079] In the neck network of YOLO11n, the nearest neighbor interpolation method is used for upsampling, converting the low-resolution feature maps extracted by the backbone network into high-resolution feature maps, and fusing feature information at different scales to improve the model's detection ability for targets of different sizes. However, players on the court often move at high speeds, which is prone to motion blur. Nearest neighbor interpolation has difficulty effectively processing these blurred areas, resulting in blurred edges or loss of key details in the generated high-resolution feature maps, which in turn limits the model's ability to capture detailed features. Therefore, the upsampling method of YOLO11n has been improved, using the dynamic upsampling operator DySample for dynamic upsampling.

[0080] The core idea of ​​dynamic upsampling by DySample is to use the sampling point generator to dynamically generate a sampling set for the input feature map X, and then use the dynamic sampling set to upsample through the grid sampling function to obtain the feature map X′. The specific upsampling process is as follows: Figure 4 shown.

[0081] The workflow of the sampling point generator is as follows Figure 5 As shown in the figure, before generating sampling points, the sampling point generator performs a linear projection on the input feature map X. One of the linear projections uses a constraint factor of 0.5 to constrain the range of movement of the sampling position. The dot product of the two projections will produce an offset set that matches the size of the upsampling scale factor s. After pixel shuffling, the sampling offset is obtained with a spatial size that matches the original grid size. The sampling offset is added to the original grid to obtain the final set of sampling points.

[0082] The dynamic sampling feature of DySample helps the model quickly extract representative features during upsampling, allowing it to more accurately capture key player characteristics when motion blur occurs. Dynamic sampling also helps the model more rationally allocate sampling points during upsampling, focusing them more closely on important areas related to the player. This generates clearer and more semantically rich high-resolution feature maps, making the upsampling process more natural and further enhancing the model's multi-scale feature recognition capabilities.

[0083] All the improvements are integrated to form the target detection model CCD-YOLO11. Its overall structure is as follows: Figure 6 Compared to YOLO11, CCD-YOLO11 optimizes YOLO11 in multiple areas, further improving the model's detection performance in soft hockey scenarios. In the backbone network, the large-kernel convolutional backbone network ConvNeXtV2 is used to give the model a larger receptive field and enhance the model's ability to extract global and local features. In the neck network, the C3k2_CA module is used to incorporate long-range dependencies and position information, strengthening the global relevance of features and improving the model's feature fusion capabilities. In the upsampling layer, DySample is used for dynamic upsampling to generate high-quality, high-resolution feature maps, optimize upsampling quality, and improve the model's multi-scale feature detection capabilities.

[0084] Use the above CDD-YOLOv11 model to train the dataset produced in step 3. The training set is trained with a fixed input image size of 640*640, 300 epochs of training, 16 batch sizes, and a stochastic gradient descent (SGD) optimizer. The momentum factor is set to 0.937, the initial learning rate is 0.02, and the weight decay coefficient is 0.0005.

[0085] S102, obtaining a panoramic video stream of soft hockey;

[0086] S103 , according to each frame in the soft hockey panoramic video stream, a smooth transition mechanism of a historical frame queue and a coordinate list is adopted, and the panoramic video is directed based on the target detection model.

[0087] S103 includes directing of ball coordinates and directing of player coordinates;

[0088] When directing based on spherical coordinates, the following steps are specifically included:

[0089] (1) Determine the initial center position pre_position (x0, y0) and the corresponding initial coordinates (8160*0.5, 3616*0.7);

[0090] (2) Obtain the coordinate list position and the history frame queue history_frame; and set the detection threshold n and the output video threshold t; the coordinate list position is initially empty; in order to meet the real-time requirements and avoid a large amount of time consumption for image reasoning and a long history frame queue resulting in excessive delay in output video, the detection threshold n is used to indicate that detection is performed once every n frames; the output video threshold t is used to indicate that the images in the history frame queue history_frame are output every t frames;

[0091] (3) Each frame of the soft hockey panoramic image in the soft hockey panoramic video stream is added to the history frame queue history_frame; when the value count of a frame of the soft hockey panoramic image satisfies count%n=0 (the remainder of count divided by n is 0), the target detection model is used to detect the ball, and the coordinates of the detected ball (bounding box) are added to the coordinate list position;

[0092] (4) When the length of the history frame queue history_frame reaches the output video threshold t, the displacement (xt, yt) of the center coordinates of the i-th frame of the soft hockey panoramic image compared to the i-1-th frame of the soft hockey panoramic image is determined according to the tail coordinates (xt, yt) in the coordinate list position. len and y len ) and the direction of movement (flag x , flag y ), and determining the center coordinates of the panorama image of the soft hockey puck of the i-th frame;

[0093] in, Calculate the center of the frame to be output. The center coordinates of the "first" frame in the historical frame queue are (x0, y0). The center coordinates of the i-th frame (x i ,y i )for:

[0094] x i =x0+flag x *(i-1)*x len ;

[0095] y i =y0+flag y *(i-1)*y len ;

[0096] (5) Loop through each frame of the soft hockey panoramic image in the history frame queue history_frame, and output the image with a size of H*W, with the center of the obtained soft hockey panoramic image as the center; after the output is completed, update the initial position coordinate pre_position to (x t ,y t ), clear the coordinate list position and the history frame queue history_frame; return to the step of adding each frame of the soft hockey panoramic image of the soft hockey panoramic video stream to the history frame queue history_frame until the output of the soft hockey panoramic video stream is completed;

[0097] (6) According to the output images of the current soft hockey panoramic video stream, the broadcast is directed based on the ball coordinates.

[0098] When directing based on player coordinates, the following steps are included:

[0099] 1) Determine the initial center position pre_position (x0, y0) and the corresponding initial coordinates (8160*0.5, 3616*0.7);

[0100] 2) Obtain the coordinate list positions and the history frame queue history_frame; and set the detection threshold n and the output video threshold t; the coordinate list positions is initially empty; the detection threshold n is used to indicate that detection is performed once every n frames; the output video threshold t is used to indicate that the images in the history frame queue history_frame are output every t frames;

[0101] 3) Each frame of the field hockey panoramic image in the field hockey panoramic video stream is added to the history frame queue history_frame; when the value count of a field hockey panoramic image frame satisfies count%n=0, the target detection model is used to detect the player and add the detected player's coordinates to the coordinate list person_positions; K-Means clustering is performed on the coordinate list person_positions to determine the center coordinates of the cluster where the players are relatively concentrated, and the center coordinates are added to the coordinate list positions; the coordinate list person_positions is initially empty;

[0102] The specific clustering steps include:

[0103] (1) Randomly select k coordinate points from the given coordinate list person_positions as the initial cluster centers.

[0104] (2) For each coordinate point in the coordinate list person_positions, calculate the Euclidean distance d(P,Q) between it and each cluster center, and then assign the coordinate point to the cluster with the nearest cluster center.

[0105]

[0106] (3) For each cluster, calculate the average value of all coordinate points in the cluster and use the average value as the new cluster center. j The data point set it contains is {(x i1 ,y i1 ), (x i2 ,y i2 )...(x inj ,y inj )}, where n j is cluster c j The number of data points in the cluster. The cluster center coordinates (x′ j , y′ j ) is calculated as:

[0107]

[0108] (4) Repeat steps (2) and (3) until the cluster center coordinates no longer change significantly (i.e., the maximum number of iterations is reached or the change in the cluster center is less than a certain threshold).

[0109] (5) Count the number of samples in each cluster, find the cluster with the largest number of samples, and obtain the center coordinates of the cluster.

[0110] 4) When the length of the history frame queue history_frame reaches the output video threshold t, according to the tail coordinate (x t ,y t ) determines the displacement (x) of the center coordinates of the i-th frame of the soft hockey panoramic image compared to the i-1-th frame of the soft hockey panoramic image len and y len ) and the direction of movement (flag x and flg y ), and determining the center coordinates of the panorama image of the soft hockey puck of the i-th frame;

[0111] in,

[0112] Calculate the center of the frame to be output. The center coordinates of the "first" frame in the historical frame queue are (x0, y0). The center coordinates of the i-th frame (x i ,y i )for:

[0113] x i =x0+flag x *(i-1)*x len ;

[0114] y i =y0+flag y *(i-1)*y len ;

[0115] 5) Loop through each frame of the soft hockey panoramic image in the history frame queue history_frame, and output the image with a size of H*W, with the center of the obtained soft hockey panoramic image as the center. After the output is completed, update the initial position coordinate pre_position to (x t ,y t ), clear the coordinate list position and the history frame queue history_frame; return to the step of adding each frame of the soft hockey panoramic image of the soft hockey panoramic video stream to the history frame queue history_frame until the output of the soft hockey panoramic video stream is completed;

[0116] 6) Direct the broadcast based on the player coordinates in sequence according to the output images of the current soft hockey panoramic video stream.

[0117] The CDD-YOLOv11 model-based real-time detection system for soft hockey in panoramic videos has been trained and achieved remarkable results. After constructing and annotating the dataset, the model was trained for 300 rounds to continuously optimize its ability to recognize various objects. Through this process, the model demonstrated excellent detection accuracy across multiple categories. Specifically, the model demonstrated very high accuracy when detecting the four main objects: players, balls, referees, and goalkeepers. The CDD-YOLOv11 performance using the mAP50 and mAP50.95 evaluation metrics is shown in Table 2:

[0118] Table 2

[0119] Model mAP@0.5(%) mAP@0.5:0.95(%) Params(M) GFLOPS(G) YOLOv5s 85.5 54.3 9.1 23.8 YOLOv6s 77.2 48.8 16.3 44.0 YOLOv8s 85.6 54.6 11.1 28.4 YOLOv9s 82.3 53.0 7.1 26.7 YOLOv10s 85.8 55.0 7.2 21.4 YOLOv11s 85.5 55.6 9.4 21.3 YOLOv12s 84.3 54.6 9.1 19.3 CCD-YOLO11n(ours) 86.6 56.6 8.1 22.2

[0120] The intelligent broadcasting method based on the ball and players, combined with target detection and cluster analysis technology, realizes dynamic and precise picture adjustment in real-time video broadcasting. The technical effects of these two methods include:

[0121] (1) Smooth transition to avoid frequent screen jumps: Through the smooth transition mechanism of the historical frame queue and coordinate list, whether it is a ball-based or player-based broadcast, the screen adjustment will not appear abrupt. The transition every t frames ensures the continuity of the broadcast screen and avoids frequent jumps caused by detection errors or excessive player movement speed, thereby improving viewing comfort.

[0122] (2) Maintaining focus on the core object (ball): The ball-based intelligent broadcasting method ensures that the key object in the game, the ball, is always in the center of the screen. This is particularly important in fast-moving scenes. The ball is always in the center of the screen, allowing viewers to always focus on the key developments in the game, especially during fast counterattacks or attacks, and clearly see the ball's trajectory.

[0123] (3) Intelligent Adjustment of Player Distribution: The player-based intelligent broadcasting method uses K-Means cluster analysis to accurately capture high-density activity areas during the game, ensuring that when players are concentrated, the camera can automatically adjust the viewing angle to focus on the most intense or critical areas. This method is particularly suitable for precise broadcasting in areas where multiple players are active (such as offense, passing, and scrambling), allowing viewers to quickly grasp the climax and key moments of the game.

[0124] (4) Preventing image loss due to missed detection: On the court, rapid movement and occlusion are common. By smoothly transitioning through the historical frame queue and coordinate list, even if a frame fails to successfully detect the ball or player, the system can still smoothly transition to the next frame based on the previous historical coordinates, ensuring that the image does not lose focus or lose key information.

[0125] (5) Automation and intelligence: Through these two intelligent directing methods, the system can automatically determine and adjust the focus of the picture without manual intervention. Whether based on the movement of the ball or the distribution of players, the system can respond in a short time, greatly improving the efficiency and quality of directing and reducing the burden on manual directing.

[0126] This application utilizes advanced target detection technology and intelligent algorithms to analyze exciting gameplay footage in real time and automatically adjust camera angles based on dynamic changes in the game, generating a dynamic and rhythmic broadcast effect. Requiring no human intervention, this automated broadcast system significantly reduces the high costs of traditional broadcasting and greatly improves broadcasting efficiency. More importantly, this technological breakthrough will provide a new solution for the high-quality dissemination and popularization of soft hockey events, allowing them to be presented to a wider audience in a more professional, efficient, and precise manner, further promoting the promotion and development of soft hockey both domestically and internationally.

[0127] Based on the same inventive concept, embodiments of the present application also provide a device for automatically directing panoramic field hockey videos, configured to implement the aforementioned method for automatically directing panoramic field hockey videos. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the device for automatically directing panoramic field hockey videos provided below can be found in the aforementioned definition of the method for automatically directing panoramic field hockey videos, and will not be further elaborated here.

[0128] In an exemplary embodiment, a panoramic video automatic broadcasting device for soft hockey is provided, comprising:

[0129] A model construction module is configured to construct a target detection model based on an improved YOLOv11 model; the target detection model is configured to determine target detection results based on a panoramic image of a soft hockey game; the target detection results include: players, goalkeepers, balls, and referees; the improved YOLOv11 model is based on the YOLOv11 model and uses a backbone network that incorporates ConvNeXtV2, a C3k2 module that incorporates a CA attention mechanism, and an upsampling method that uses a dynamic upsampling operator;

[0130] A video stream acquisition module is used to acquire a panoramic video stream of soft hockey;

[0131] The intelligent broadcast guidance module is used to guide the panoramic video based on each frame in the soft hockey panoramic video stream, using a smooth transition mechanism of historical frame queues and coordinate lists and based on the target detection model.

[0132] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, memory, and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for automatically directing a panoramic video of soft hockey is implemented.

[0133] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0134] In an exemplary embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0135] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0136] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0137] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may include, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units based on quantum computing, and the like.

[0138] In this application, all actions to obtain signals, information or data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0139] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0140] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for automatically directing a panoramic video of soft hockey, characterized in that: The method for automatically directing a panoramic video of soft hockey includes: Based on an improved YOLOv11 model, a target detection model is constructed; the target detection model is used to determine target detection results based on a panoramic image of a soft hockey game; the target detection results include: players, goalkeepers, balls, and referees; the improved YOLOv11 model is based on the YOLOv11 model and uses a backbone network that introduces ConvNeXtV2, a C3k2 module that introduces a CA attention mechanism, and an upsampling method using a dynamic upsampling operator; Get the full video stream of soft hockey; According to each frame in the soft hockey panoramic video stream, a smooth transition mechanism of historical frame queue and coordinate list is adopted, and the panoramic video is directed based on the target detection model.

2. The method for automatically directing a panoramic video of soft hockey according to claim 1, characterized in that: The target detection model is constructed based on the improved YOLOv11 model, specifically including: Obtain panoramic videos of soft hockey at different time periods; Based on the soft hockey panoramic videos of different time periods, the OpenCV library technology is used to extract the soft hockey panoramic images frame by frame; Performing data preprocessing on the soft hockey panoramic image; the data preprocessing includes: image cropping and image screening; Perform object annotation on the preprocessed soft hockey panoramic images and construct a dataset; According to the dataset, an object detection model is built based on the improved YOLOv11 model.

3. The method for automatically directing a panoramic video of soft hockey according to claim 2, characterized in that: The data preprocessing of the soft hockey panoramic image specifically includes: The panorama image of the soft hockey is cropped based on a sliding window method to obtain an image of a set size; According to the image of the set size, the YOLOv8n pre-trained model is used to perform image screening; the YOLOv8n pre-trained model is used to identify whether there is an object in the image of the set size.

4. The method for automatically directing a panoramic video of soft hockey according to claim 1, wherein: The ConvNextV2 Block module in ConvNeXtV2 uses a combination of large kernel convolution and inverse bottleneck structure.

5. The method for automatically directing a panoramic video of soft hockey according to claim 1, wherein: The method uses a smooth transition mechanism of a historical frame queue and a coordinate list based on each frame in the soft hockey panoramic video stream, and directs the panoramic video based on a target detection model, specifically including: When directing based on spherical coordinates, determine the initial center position pre_position (x0, y0) and the corresponding initial coordinates; Get the coordinate list position and the history frame queue history_frame; and set the detection threshold n and the output video threshold t; the coordinate list position is initially empty; the detection threshold n is used to indicate that detection is performed once every n frames; the output video threshold t is used to indicate that the images in the history frame queue history_frame are output every t frames; Each frame of the soft hockey panoramic image in the soft hockey panoramic video stream is added to the history frame queue history_frame; when the value count of a frame of the soft hockey panoramic image satisfies count%n=0, the object detection model is used to detect the ball and the coordinates of the detected ball are added to the coordinate list position; When the length of the history frame queue history_frame reaches the output video threshold t, according to the tail coordinate (x t ,y t ) determining a displacement amount and a movement direction of the center coordinates of the i-th frame of the soft hockey panoramic image compared to the (i-1)-th frame of the soft hockey panoramic image, and determining the center coordinates of the i-th frame of the soft hockey panoramic image; Loop through each frame of the soft hockey panorama image in the history frame queue history_frame, and output the image with the center of the obtained soft hockey panorama image as the center; after the output is completed, update the initial position coordinate pre_position to (x t ,y t ), clear the coordinate list position and the history frame queue history_frame; return to the step of adding each frame of the soft hockey panoramic image of the soft hockey panoramic video stream to the history frame queue history_frame until the output of the soft hockey panoramic video stream is completed; According to the output images of the current soft hockey panoramic video stream, the broadcast is directed based on the ball coordinates in sequence.

6. The method for automatically directing a panoramic video of soft hockey according to claim 1, wherein: The method uses a smooth transition mechanism of a historical frame queue and a coordinate list based on each frame in the soft hockey panoramic video stream, and directs the panoramic video based on a target detection model, specifically including: When directing based on player coordinates, determine the initial center position pre_position(x0,y0) and the corresponding initial coordinates; Get the coordinate list positions and the history frame queue history_frame; and set the detection threshold n and output video threshold t; the coordinate list positions is initially empty; the detection threshold n is used to indicate that detection is performed once every n frames; the output video threshold t is used to indicate that the images in the history frame queue history_frame are output every t frames; Each frame of the soft hockey panoramic image in the soft hockey panoramic video stream is added to the history frame queue history_frame; when the value count of a frame of the soft hockey panoramic image satisfies count%n=0, the target detection model is used to detect the player and add the detected player's coordinates to the coordinate list person_positions; a K-Means clustering calculation is performed on the coordinate list person_positions to determine the center coordinates of the cluster where the players are relatively concentrated, and the center coordinates are added to the coordinate list positions; the coordinate list person_positions is initially empty; When the length of the history frame queue history_frame reaches the output video threshold t, according to the tail coordinate (x t ,y t ) determining a displacement amount and a movement direction of the center coordinates of the i-th frame of the soft hockey panoramic image compared to the (i-1)-th frame of the soft hockey panoramic image, and determining the center coordinates of the i-th frame of the soft hockey panoramic image; Loop through each frame of the soft hockey panorama image in the history frame queue history_frame, and output the image with the center of the obtained soft hockey panorama image as the center; after the output is completed, update the initial position coordinate pre_position to (x t ,y t ), clear the coordinate list position and the history frame queue history_frame; return to the step of adding each frame of the soft hockey panoramic image of the soft hockey panoramic video stream to the history frame queue history_frame until the output of the soft hockey panoramic video stream is completed; According to the output images of the current soft hockey panoramic video stream, the broadcast is directed based on the player coordinates in sequence.

7. A panoramic video automatic broadcasting device for soft hockey, characterized in that: The soft hockey panoramic video automatic broadcasting device includes: A model construction module is configured to construct a target detection model based on an improved YOLOv11 model; the target detection model is configured to determine target detection results based on a panoramic image of a soft hockey game; the target detection results include: players, goalkeepers, balls, and referees; the improved YOLOv11 model is based on the YOLOv11 model and uses a backbone network that incorporates ConvNeXtV2, a C3k2 module that incorporates a CA attention mechanism, and an upsampling method that uses a dynamic upsampling operator; A video stream acquisition module is used to acquire a panoramic video stream of soft hockey; The intelligent broadcast guidance module is used to guide the panoramic video based on each frame in the soft hockey panoramic video stream, using a smooth transition mechanism of historical frame queues and coordinate lists and based on the target detection model.

8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the automatic broadcasting method for panoramic video of soft hockey according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for automatically directing a panoramic video of soft hockey according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for automatically directing a panoramic video of soft hockey according to any one of claims 1 to 6 is implemented.