Competition playback image optimization method, device and equipment

By dynamically generating ROI regions and resolutions, and combining color channels and gradient variance for image quality restoration, the problems of resolution adaptation, distortion correction, and ROI region processing in sports video processing are solved, achieving efficient and low-cost sports replay video optimization.

CN121665073APending Publication Date: 2026-03-13GUANGDONG YAOLONG INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies for sports event image processing suffer from problems such as lack of dynamic resolution adaptation, low accuracy of image distortion correction, unreasonable ROI region processing, and resource waste and image blurring caused by reliance on single-frame information for image quality restoration.

Method used

By analyzing playback requirements and acquisition conditions, the ROI region and resolution are dynamically generated. Image quality is restored by combining color channels and gradient variance, fused image segments are generated, and the frame rate is adjusted according to the playback speed to optimize the event playback images.

Benefits of technology

It improves the accuracy of motion pixel recognition, reduces invalid data, lowers bandwidth and storage costs, ensures smooth playback, adapts to multiple terminals, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121665073A_ABST
    Figure CN121665073A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of sports management, in particular to a match playback image optimization method, device and equipment, and the method comprises the steps: analyzing match image data according to a preset fragment duration, a playback demand, a preset collection condition and a preset image encoder, so as to obtain a first ROI region and a playing resolution; performing image quality restoration on the first ROI region according to a preset color channel, a preset gradient variance and a preset adjustment coefficient set to obtain a restored ROI region; generating a fused image segment according to the restored ROI region, a preset high code rate parameter and a preset low code rate parameter; performing frame rate adjustment on the fused image segment according to a preset playback speed and a preset playing resolution to obtain a playback image segment; by accurately analyzing playback requirements, optimizing the ROI, repairing image problems, fusing different code rate parameters and adjusting the frame rate according to requirements, the method is integrally adaptive to multiple competition and multiple terminals, user experience is optimized, and competition playback is assisted to be efficiently presented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sports management technology, and in particular to a method, apparatus and equipment for optimizing event replay images. Background Technology

[0002] With the development of high-definition imaging, 5G, and AI technologies, on-site image acquisition and playback has been applied to various scenarios such as sports events and emergency rescue. However, existing technologies still have significant shortcomings: First, resolution adaptation lacks dynamism, often using a fixed resolution throughout (such as 4K). This results in wasted storage in low-complexity scenarios (such as empty meeting rooms) (4K stores approximately 80GB per hour, four times that of 1080P), while in high-complexity scenarios (such as players battling in a game), insufficient resolution leads to loss of detail. Second, image distortion correction accuracy is low, relying on data from a single calibration board without considering the influence of ambient light and temperature. Even after correction, there is still a 10-15 pixel deviation, and the distortion degree is not dynamically matched with the correction parameters, leading to slight over-correction and incomplete heavy correction. Third, ROI region processing is unreasonable, based on fixed coordinate division. High bitrate areas are over-covered, with bandwidth consumption exceeding 20Mbps, exceeding the 5G carrying capacity. Image quality restoration relies solely on single-frame information, and the image still contains blurry noise. Summary of the Invention

[0003] In order to overcome the shortcomings of the prior art, the purpose of this invention is to provide a method, apparatus and device for optimizing sports replay images.

[0004] The first aspect of this invention provides a method for optimizing sports event replay images, comprising: acquiring replay requirements and sports event image data; analyzing the sports event image data according to a preset segment duration, replay requirements, preset acquisition conditions, and a preset image encoder to obtain a first Region of Interest (ROI) and playback resolution; performing image quality restoration on the first ROI according to a preset color channel, a preset gradient variance, and a preset set of adjustment coefficients to obtain a restored ROI; generating a fused image segment according to the restored ROI, preset high bitrate parameters, and preset low bitrate parameters; and adjusting the frame rate of the fused image segment according to a preset playback speed and playback resolution to obtain a replay image segment.

[0005] Furthermore, the step of analyzing the event video data according to preset segment duration, playback requirements, preset acquisition conditions, and preset image encoder to obtain the first ROI region and playback resolution includes: filtering event video data segments according to segment duration and playback requirements to obtain a key segment sequence; extracting the first ROI region from the key segment sequence according to acquisition conditions, a preset background difference threshold, and a preset inter-frame difference threshold; analyzing the key segment sequence according to the image encoder and a preset text encoder to obtain a scene complexity index; and generating the playback resolution according to the scene complexity index.

[0006] Furthermore, the step of filtering event video data based on segment duration and playback requirements to obtain a key segment sequence includes: performing target identification on the event video data according to a preset deep learning model and playback requirements to obtain target type data, target location data, and target confidence data; calculating the target type data, target location data, and target confidence based on a preset weight set to obtain segment importance values; and filtering event video data based on segment duration and segment importance values ​​to obtain a key segment sequence.

[0007] Furthermore, the step of analyzing the key segment sequence based on the image encoder and the preset text encoder to obtain the scene complexity index includes: extracting features from the key segment sequence based on the image encoder to obtain an image feature vector; extracting features from the key segment sequence based on the text encoder to obtain a text feature vector; and generating the scene complexity index based on the image feature vector and the text feature vector.

[0008] Further, the step of generating a playback resolution based on the scene complexity index includes: analyzing the scene complexity index to obtain a complexity analysis result; when the complexity analysis result indicates that the scene complexity index is in a preset low complexity range, generating a first resolution based on the scene complexity index; generating a playback resolution based on the first resolution; when the complexity analysis result indicates that the scene complexity index is in a preset medium complexity range, generating a second resolution based on the scene complexity index; generating a playback resolution based on the second resolution; when the complexity analysis result indicates that the scene complexity index is in a preset high complexity range, generating a third resolution based on the scene complexity index; generating a playback resolution based on the third resolution; wherein the first resolution is less than the second resolution; and the second resolution is less than the third resolution.

[0009] Further, the step of extracting the first ROI region from the key segment sequence based on the acquisition conditions, a preset background difference threshold, and a preset inter-frame difference threshold includes: acquiring static frame data and reference frame data from the key segment sequence based on the acquisition conditions; performing feature extraction on the static frame data and reference frame data based on the background difference threshold to obtain candidate moving pixels; performing feature extraction on the static frame data and reference frame data based on the inter-frame difference threshold to obtain inter-frame moving pixels; performing intersection analysis on the candidate moving pixels and inter-frame moving pixels to obtain intersection moving pixels; and generating the first ROI region based on the intersection moving pixels.

[0010] Furthermore, the step of performing image quality restoration on the first ROI region based on preset color channels, preset gradient variance, and preset adjustment coefficient set to obtain a restored ROI region includes: analyzing the key segment sequence based on color channels, preset image gray levels, and gradient variance to obtain the inter-frame absolute difference and texture complexity; and performing image quality restoration on the first ROI region based on the inter-frame absolute difference, texture complexity, and adjustment coefficient set to obtain a restored ROI region.

[0011] Furthermore, the step of generating a fused image segment based on the repaired ROI region, preset high bitrate parameters, and preset low bitrate parameters includes: obtaining a second ROI region from the key segment sequence based on the high bitrate parameters; obtaining a third ROI region from the key segment sequence based on the low bitrate parameters; and generating a fused image segment based on the repaired ROI region, the second ROI region, and the third ROI region.

[0012] A second aspect of the present invention provides a sports event replay image optimization device, comprising: a data acquisition module for acquiring replay requirements and sports event image data; an analysis module for analyzing the sports event image data according to a preset segment duration, replay requirements, preset acquisition conditions, and a preset image encoder to obtain a first Region of Interest (ROI) and playback resolution; an image quality restoration module for performing image quality restoration on the first ROI according to a preset color channel, a preset gradient variance, and a preset set of adjustment coefficients to obtain a restored ROI; an image segment generation module for generating a fused image segment according to the restored ROI, preset high bitrate parameters, and preset low bitrate parameters; and a frame rate adjustment module for adjusting the frame rate of the fused image segment according to a preset playback speed and playback resolution to obtain a replay image segment.

[0013] A third aspect of the present invention provides a sports replay image optimization device, the sports replay image optimization device comprising: a memory and at least one processor, the memory storing instructions; at least one processor calling the instructions in the memory to cause the sports replay image optimization device to perform the various steps of a sports replay image optimization method as described in any one of the above.

[0014] In the technical solution of this invention, by analyzing playback requirements, the content, quality, and time requirements are accurately located, providing direction for subsequent processing. Combining segment duration, acquisition conditions, and the image encoder, a first ROI region and an adapted resolution are generated, improving the accuracy of motion pixel recognition, reducing invalid data, and balancing image quality and efficiency. The first ROI region is repaired through gradient variance and adjustment coefficients, resolving issues of blurring, loss of detail, and color cast, laying the foundation for fusion. Fusion repair of the ROI and high-bitrate and low-bitrate regions ensures image quality in the core area and controls costs in the background area, resulting in natural image transitions. Finally, the frame rate is adjusted according to playback speed and resolution to ensure smooth playback. Overall, the solution is compatible with multiple events and terminals, improving user experience, reducing bandwidth and storage costs, and facilitating efficient presentation and commercial application of event replays. Attached Figure Description

[0015] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which: Figure 1 This is a first flowchart of a method for optimizing sports replay images provided in an embodiment of the present invention; Figure 2 This is a second flowchart of a method for optimizing sports replay images provided in an embodiment of the present invention; Figure 3 This is a third flowchart of a method for optimizing sports replay images provided in an embodiment of the present invention; Figure 4 This is a fourth flowchart of a sports replay image optimization method provided in an embodiment of the present invention; Figure 5 A fifth flowchart of a method for optimizing sports replay images provided in an embodiment of the present invention; Figure 6 The sixth flowchart of a sports replay image optimization method provided in an embodiment of the present invention; Figure 7 The seventh flowchart of a sports replay image optimization method provided in an embodiment of the present invention; Figure 8 The eighth flowchart of a sports replay image optimization method provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of a sports replay image optimization device provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of a sports replay image optimization device provided in an embodiment of the present invention. Detailed Implementation

[0016] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0017] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of a method for optimizing sports replay images according to the present invention includes: 101. Obtain replay requirements and event video data; In this embodiment, playback requirements refer to the user's specific requirements for video content, playback quality, and time range. Specifically, they can be achieved by using a natural language processing model to parse user instructions, which can then guide the selection of key segments and parameter settings. 102. Analyze the event video data according to the preset segment length, playback requirements, preset acquisition conditions and preset image encoder to obtain the first ROI area and playback resolution. In this embodiment, the segment duration refers to the duration threshold of each image segment; by combining the segment duration, playback requirements (targeted content selection), acquisition conditions (ensuring frame data quality) and image encoder (extracting key features) to analyze the event images, the first ROI region (focusing on the core motion area) and the playback resolution are accurately generated, improving the accuracy of motion pixel recognition, reducing the false judgment rate, and reducing invalid data processing, balancing image quality and efficiency, and adapting to the playback requirements of multiple scenarios; 103. Perform image quality restoration on the first ROI region based on the preset color channels, preset gradient variance, and preset adjustment coefficient set to obtain the restored ROI region; In this embodiment, by combining preset color channels (such as YUV channels to accurately locate image quality defects), gradient variance (to determine texture complexity and identify blurred areas), and adjustment coefficient set (to dynamically adapt the repair intensity), the first ROI area is specifically repaired, effectively solving the problems of motion blur, loss of detail, and color shift, improving the detail recognition of the core area and the coordination of the image, laying a high-quality foundation for subsequent image fusion, and adapting to complex scenes of multiple events. 104. Generate fused image segments based on the repaired ROI region, preset high bitrate parameters, and preset low bitrate parameters; In this embodiment, a fused image is generated by combining the repaired ROI area (with optimized image quality in the core motion area), preset high bitrate parameters (focusing on densely detailed areas and preserving high-definition details), and low bitrate parameters (covering non-core background areas and controlling resource consumption). This ensures high-quality image quality in the core area while reducing resource consumption in non-critical areas. The image transition is natural and seamless, balancing the viewing experience with bandwidth and storage costs, and is suitable for multi-terminal event playback scenarios. 105. Adjust the frame rate of the merged image segments according to the preset playback speed and playback resolution to obtain the playback image segments; In this embodiment, by analyzing playback requirements, the content, quality, and time requirements are accurately identified, providing direction for subsequent processing. Combining segment duration, acquisition conditions, and the image encoder, a first ROI region and an adapted resolution are generated, improving the accuracy of motion pixel recognition, reducing invalid data, and balancing image quality and efficiency. The first ROI region is repaired using gradient variance and adjustment coefficients, resolving issues such as blurring, loss of detail, and color cast, laying the foundation for fusion. Fusion repair of the ROI and high-bitrate and low-bitrate regions ensures image quality in the core area while controlling costs in the background area, resulting in natural image transitions. Finally, the frame rate is adjusted according to playback speed and resolution to ensure smooth playback. Overall, the system is adapted to multiple events and terminals, improving user experience, reducing bandwidth and storage costs, and facilitating efficient presentation and commercial application of event replays.

[0018] Please see Figure 2 A second embodiment of a method for optimizing sports replay images in this invention includes: 201. Based on the segment length and playback requirements, the event video data is segmented to obtain a sequence of key segments; In this embodiment, the basic duration is set according to the type of event and the replay scenario. For example, the duration of a regular replay segment in a football match is set to 5-10 seconds (corresponding to 150-300 frames), and the duration of a segment of key events (such as goals and controversial decisions) is set to 10-15 seconds to ensure the complete presentation of the event process. The replay requirements are converted into quantifiable indicators, such as "tactical analysis requirements" corresponding to "player tactical movement density ≥ 80%" and "passes ≥ 5 times / segment", and "highlights requirements" corresponding to "high dynamic events (goals, saves) occurring ≥ 1 time" and "audience attention score ≥ 8 points (out of 10)". 202. Based on the acquisition conditions, the preset background difference threshold, and the preset inter-frame difference threshold, the first ROI region is extracted from the key segment sequence; In this embodiment, frame data quality is ensured by collecting data under certain conditions (e.g., sampling at 1 frame / second, using the initial static frame as a baseline, and filtering static frames based on a pixel change rate of <5% for three consecutive frames). Motion pixels are accurately extracted using background difference thresholding, and the first ROI region is generated through intersection analysis and connectivity algorithms. This improves the accuracy of motion pixel recognition, reduces the false positive rate, defines the scope for subsequent optimization, and adapts to multiple sporting events, thereby improving image processing efficiency and adaptability. 203. Analyze the key segment sequence based on the image encoder and the preset text encoder to obtain the scene complexity index; In this embodiment, a CNN-based encoder (such as ResNet50) is used to extract three layers of features from the key segment sequence: low-level features (edges, textures, color contrast), mid-level features (target density, scale differences), and high-level features (inter-frame motion speed, acceleration), forming a 512-dimensional image feature vector; a Transformer-based encoder (such as BERT) is used to encode the associated text of the segments (match rule labels "penalty kick" and "offside", action description "curved ball", event importance "key goal") to capture semantic information and form a 512-dimensional text feature vector; 204. Generate playback resolution based on scene complexity index; In this embodiment, the segment duration is set according to the event type and replay requirements (e.g., 5-10 seconds for regular football matches, 10-15 seconds for key events), and converted into quantitative indicators to screen key segments, accurately focusing on core content and improving the efficiency of viewers obtaining effective information; the first ROI region is extracted based on the collection conditions and background difference threshold to improve the accuracy of motion pixel recognition and reduce the false judgment rate, thus defining the scope for subsequent optimization; the scene complexity index is obtained by extracting bimodal features using CNN and Transformer encoders, and the resolution is dynamically generated in combination with the index to improve the detail recognition of high-complexity scenes, reduce resource consumption of low-complexity scenes, adapt to multiple terminals and events, reduce bandwidth and storage costs, improve processing efficiency, and balance the viewing experience and operational benefits.

[0019] Please see Figure 3 A third embodiment of a method for optimizing sports replay images in this invention includes: 301. Based on the preset deep learning model and playback requirements, target identification is performed on the event video data to obtain target type data, target location data, and target confidence data; In this embodiment, the deep learning model refers to a real-time target detection framework built on convolutional neural networks. Specifically, it can be implemented using YOLOv5 or Faster R-CNN architectures. It captures moving targets at different scales through a multi-level feature fusion mechanism. The model continuously detects the spatial location and category attributes of core elements such as athletes and spheres in the video stream according to playback requirements, and outputs a confidence value of 0-1. Targets with a confidence value ≥ 0.5 are considered valid detection results, and low-confidence targets that are falsely detected (such as misidentifying debris in the audience as a sphere) are removed. This provides a data foundation for dynamically dividing ROI regions. 302. Calculate the target type data, target location data, and target confidence based on a preset weight set to obtain the segment importance value; In this embodiment, for all valid targets (confidence ≥ 0.5) in a certain frame of image, the local score (Sᵢ) of each target is calculated. In the formula, The type weight of the i-th target. Let i be the position weight of the i-th target. Let be the confidence weight for the i-th target; 303. Based on the segment duration and segment importance value, the event video data is segmented to obtain the key segment sequence; In this embodiment, for example, the average importance score is calculated for all frames within a segment duration (e.g., 5 seconds, corresponding to 150 frames) as the final importance of the segment. , In the formula, N is the number of frames in the segment, m is the number of valid targets in a single frame (the maximum target score in a single frame is used to avoid duplicate calculations for multiple targets), and k indicates that the current calculation is for the k-th frame in the segment, and its value range is ( ); In this embodiment, YOLOv5 or Faster R-CNN architecture is adopted, and multi-level feature fusion is used to capture moving targets at different scales, improving the accuracy of target recognition. Confidence scores effectively reduce false detections by the detection mechanism, lowering the missed detection rate of key targets, and accurately outputting target type, location, and confidence data, laying a reliable foundation for subsequent processing. By dynamically quantifying the importance value of segments through weight sets, and combining target type, location adaptability, and confidence scores to calculate segment scores, accurate matching of users' personalized playback needs is achieved, improving the demand adaptation rate. Key segments are filtered according to segment length and importance, retaining the core content of the original image, reducing data processing volume, and improving processing efficiency. Furthermore, frame continuity optimization ensures the logical integrity of segments, improving audience attention concentration. It also supports multiple event types and multi-terminal adaptation, combining practicality and scalability, effectively solving the pain points of low accuracy, low efficiency, and poor adaptation in traditional filtering methods.

[0020] Please see Figure 4 The fourth embodiment of a sports replay image optimization method in this invention includes: 401. Extract features from the key segment sequence using the image encoder to obtain the image feature vector; In this embodiment, the image feature vector includes low-level features: such as basic visual information like edges, textures, and color distribution, quantifying indicators such as color contrast and light and shadow change rate of the image; mid-level features: such as spatial features like target density (e.g., number of players) and target scale differences (e.g., close-up and panoramic views); and high-level features: inter-frame motion features (e.g., ball flight speed and athlete displacement acceleration), quantifying the intensity of motion. 402. Extract features from the key segment sequence using the text encoder to obtain the text feature vector; In this embodiment, the text feature vector includes text data sources: covering event rule tags (such as "penalty kick", "offside"), action descriptions (such as "tomahawk dunk", "curve shot"), event importance annotations (such as "key goal", "regular pass"), and semantic associations (such as the strong association between "red card" and "serious foul"). 403. Generate a scene complexity index based on image feature vectors and text feature vectors; In this embodiment, the image feature vector and text feature vector are first standardized and aligned in dimensions according to [0,1]. Then, an attention mechanism is used to calculate the association weights, and the feature weight fusion vector is assigned according to the scene type. The initial index is mapped through a fully connected network, and the final index is obtained by combining visual and semantic verification and event type calibration. Finally, it is divided into low, medium, and high complexity levels according to [0,3), [3,7), and [7,10]. The scene complexity index includes three levels: In this embodiment, the image feature vector covers bottom, middle, and top-level information, while the text feature vector covers semantics such as competition rules and action descriptions. The dual-modal complementarity avoids the limitations of single-visual evaluation, reduces the complexity evaluation error of low-quality images (blurred, low-light), and improves the accuracy of distinguishing visually similar scenes such as "goal" and "shot off target." Through standardization, attention weight calculation, and scene adaptation fusion, the index accurately reflects the true complexity of the scene, providing a basis for resource optimization, reducing overall bandwidth consumption, and minimizing motion blur in key scenes. The three-level complexity division adapts to multiple needs, guiding live streaming resource allocation, content editing (high-complexity segments are marked as key materials to improve editing efficiency), and training analysis. It also supports calibration for multiple competitions such as football and basketball, shortening the adaptation cycle and combining accuracy, practicality, and scalability.

[0021] Please see Figure 5 The fifth embodiment of a sports replay image optimization method in this invention includes: 501. Analyze the scene complexity index to obtain the complexity analysis results; 502. When the complexity analysis result shows that the scene complexity index is in the preset low complexity range, the first resolution is generated based on the scene complexity index. In this embodiment, the low complexity range is [0,3), which corresponds to static or low dynamic scenes such as player positioning and slow-paced passing. The human eye has low sensitivity to detail clarity, so setting a lower resolution can meet the basic viewing needs. The first resolution = 720P × (0.8 + scene complexity index / 7.5). 503. Generate playback resolution based on the first resolution; 504. When the complexity analysis result shows that the scene complexity index is in the preset medium complexity range, a second resolution is generated based on the scene complexity index. In this embodiment, the medium complexity range is [3,7), covering common dynamic scenarios such as ordinary attack and defense and tactical cooperation. It is necessary to balance clarity and smoothness, so a medium resolution is set. The second resolution = 1080P×(0.9+(scene complexity index-3) / 40). 505. Generate playback resolution based on the second resolution; 506. When the complexity analysis result indicates that the scene complexity index is in the preset high complexity range, a third resolution is generated based on the scene complexity index. In this embodiment, the high complexity range is [7,10], which includes key / high dynamic scenes such as the moment of scoring and multiple players tackling. The audience pays close attention to details (such as player movements and ball trajectory), and the highest resolution is required to ensure the experience. The third resolution = 2.7K×(1+(scene complexity index-7) / 30). 507. Generate playback resolution based on the third resolution; The first resolution is smaller than the second resolution; The second resolution is smaller than the third resolution; In this embodiment, resolutions are generated based on low complexity, medium complexity, and high complexity ranges, corresponding to 720P, 1080P, and basic resolutions (2.7K / 4K). This is combined with exponential fine-tuning to ensure details in high-complexity key scenes (such as goals) to meet the high attention demands of viewers, while reducing resource consumption in low-dynamic scenes and improving the subjective experience satisfaction of the human eye. The first, second, and third resolutions increase sequentially, and a smooth transition is achieved through formulas to avoid image quality gaps. At the same time, it is compatible with multiple terminals, reduces mobile data consumption, improves the detail recognition of large screens, optimizes resource utilization, reduces bandwidth consumption, reduces storage costs, reduces the load on encoding servers, and supports multi-event adaptation, shortening the access cycle for new events, and balancing experience, cost, and scalability.

[0022] Please see Figure 6 The sixth embodiment of a sports replay image optimization method in this invention includes: 601. Based on the acquisition conditions, static frame data and reference frame data are acquired from the key segment sequence; In this embodiment, the acquisition conditions include frame sampling frequency, reference frame selection rules, and static frame filtering conditions. Frame sampling frequency: Considering the high dynamic characteristics of the event images, the sampling frequency is set to 1 frame / second (sampling 1 frame per second can balance data volume and motion continuity). Reference frame selection rules: Select the static frame at the beginning of the key segment sequence (such as the initial standing frame where the player has not moved) as the reference frame to ensure that the reference frame is free from motion interference and has pure background information. Static frame filtering conditions: Judged by the pixel change rate between frames, when the pixel change rate of 3 consecutive frames is <5%, it is determined to be a static frame to avoid misjudging low dynamic frames as static frames. 602. Based on the background difference threshold, feature extraction is performed on static frame data and reference frame data to obtain candidate moving pixels; In this embodiment, the step of extracting candidate moving pixels includes: Differential calculation: Calculate the grayscale difference between the static frame and the reference frame using the following formula: D_background(x,y)=|I_static(x,y)-I_base(x,y)| Where I_static(x,y) is the grayscale value of the static frame at the (x,y) coordinates, I_base(x,y) is the grayscale value of the base frame at the corresponding coordinates, and D_background(x,y) is the background difference result; Threshold filtering: Background difference threshold (usually set to 20-30, grayscale value range 0-255). When D_background(x,y)≥threshold, the pixel is determined to be a "candidate moving pixel" and is considered to belong to the background change area (i.e., there may be moving targets). In another embodiment, region morphological processing involves performing dilation (3×3 structuring element) and erosion (3×3 structuring element) operations on candidate moving pixels to fill pixel gaps, eliminate isolated noise, and generate continuous candidate moving pixel regions. 603. Based on the inter-frame difference threshold, feature extraction is performed on static frame data and reference frame data to obtain inter-frame motion pixels; In this embodiment, the step of extracting inter-frame moving pixels includes: Differential calculation: Select two adjacent static frames (such as static frames extracted from the 1st second and the 2nd second), and calculate the grayscale difference between them. The formula is: D_frame(x,y)=|I_static1(x,y)-I_static2(x,y)| Where I_static1(x,y) and I_static2(x,y) are the gray values ​​of two adjacent static frames, respectively, and D_frame(x,y) is the inter-frame difference result; Threshold filtering: Inter-frame difference threshold (usually set to 15-25, slightly lower than the background difference threshold, because the differences in motion between frames are more subtle). When D_frame(x,y)≥threshold, the pixel is determined to be an "inter-frame motion pixel" and is considered to belong to the dynamic change area between frames. 604. Perform intersection analysis on candidate moving pixels and inter-frame moving pixels to obtain the intersection moving pixels; In this embodiment, the coordinates of candidate moving pixels and inter-frame moving pixels are matched one by one, and pixels that simultaneously satisfy "background difference threshold ≥ 20-30" and "inter-frame difference threshold ≥ 15-25" are selected as intersection moving pixels. 605. Generate the first ROI region based on the intersection of moving pixels; In this embodiment, connected region analysis is performed on the intersecting moving pixels. The 8-neighborhood connectivity algorithm (pixels in 8 adjacent directions are considered connected) is used to aggregate the scattered intersecting moving pixels into continuous connected regions. The minimum bounding rectangle is drawn for each connected region, and finally the first ROI candidate region is obtained by merging them, which provides a clear range for subsequent image quality restoration and high bitrate encoding. In this embodiment, sampling is performed at 1 frame / second, with the initial static frame as the baseline frame, and static frames are filtered based on a change rate of <5% for three consecutive frames. This ensures that the frame data is pure while maintaining efficiency, laying a reliable foundation for subsequent detection. By combining background subtraction (threshold 20-30) and inter-frame subtraction (threshold 15-25), along with morphological processing and an 8-neighbor connectivity algorithm, the accuracy of motion pixel recognition is improved, the false positive rate is reduced, and lighting and noise interference are accurately eliminated. The core moving targets such as players and balls are fully captured, forming the first ROI region, making resource utilization more efficient. At the same time, it adapts to multiple event scenarios, shortens the adaptation cycle for new events, reduces positioning errors in complex scenes, provides a clear range for image quality restoration and high bitrate encoding, and helps improve the quality of event images.

[0023] Please see Figure 7 The seventh embodiment of a sports replay image optimization method in this invention includes: 701. Analyze the key segment sequence based on color channels, preset image gray levels, and gradient variance to obtain the inter-frame absolute difference and texture complexity. In this embodiment, the color channel is the YUV color channel. The Y color channel (luminance channel) of the YUV color space is selected as the core analysis object because motion blur and detail loss in sports images are mainly reflected in luminance changes, and the Y channel can more accurately reflect image quality defects. At the same time, the U and V color channels are referenced to help determine the degree of color shift, avoiding omissions in color restoration caused by relying solely on luminance. By presetting 256 levels of image grayscale (0-255), the sequence of key color segments is converted into a grayscale image, simplifying the calculation while ensuring that the grayscale gradient changes can fully reflect the brightness and darkness details of the image. (e.g., player jersey texture, ball surface lighting); By setting the gradient calculation window to 3×3, the Sobel operator is used to calculate the horizontal (Gx) and vertical (Gy) gradients of the image. The gradient variance reflects the dispersion of the gradient values; the larger the variance, the richer the image edges and textures, and vice versa (possibly blurry). This allows for the analysis of texture complexity features. Three consecutive frames (e.g., frame n, frame n+1, and frame n+2) are selected from the key segment sequence. The calculation focuses on the intra-frame region corresponding to the key segment sequence. For the Y-channel data, the absolute difference between corresponding pixels in adjacent frames is calculated using the following formula: In the formula, Let (x, y) be the frame pixel value at coordinate (x, y) of the nth frame. This represents the frame pixel value at coordinates (x, y) of the (n+1)th frame. This is the absolute difference between two frames; 702. Perform image quality restoration on the first ROI region based on the absolute difference between frames, texture complexity, and adjustment coefficient set to obtain the restored ROI region; In this embodiment, the degree of blurring is determined according to the absolute difference between frames, the blur repair coefficient is adjusted, and Wiener filtering and optical flow are used to deblur the first ROI region; the detail enhancement coefficient is set according to the texture complexity, and the details of the first ROI region are enhanced using an unsharpened mask; the first ROI region is color-complemented with reference to the color channel differential adjustment coefficient; the edges of the first ROI region are blurred with a smoothing coefficient, and the repaired ROI region is obtained after fusion. In this embodiment, the Y color channel in the YUV color channel is the core, with the U and V color channels as auxiliary channels. Combining 256 gray levels and a 3×3 window Sobel operator, the absolute difference between frames (for locating motion blur) and texture complexity (for recognizing detail loss) are accurately extracted, improving the accuracy of defect recognition and avoiding color shift omissions. The blur repair coefficient is adjusted according to the absolute difference between frames, and Wiener filtering and optical flow are used to remove blur and improve the edge sharpness of moving targets. The detail enhancement coefficient is set according to the texture complexity, and the unsharpened mask improves the detail recognition. Color correction coefficients are used to complement colors, and smoothing coefficients are used to eliminate edge breaks, resulting in a natural image that is suitable for multiple events and complex scenes, balancing image quality and efficiency.

[0024] Please see Figure 8 The eighth embodiment of a sports replay image optimization method in this invention includes: 801. Obtain the second ROI region from the key segment sequence based on the high bitrate parameter; In this embodiment, the high bitrate parameter focuses on areas with "dense details and high user attention". The core parameters include bitrate threshold (e.g., 15-20Mbps, which is 50% higher than the normal bitrate), detail density threshold (texture complexity σ²>1200), and target priority (weight of core targets such as spheres and player faces>0.8). 802. Obtain the third ROI region from the key segment sequence based on the low bitrate parameter; In this embodiment, the low bitrate parameter is for areas with "sparse details and non-core focus". The core parameters include the bitrate threshold (e.g., 3-5 Mbps, which is 40% lower than the regular bitrate), the detail density threshold (texture complexity σ² < 500), and the target priority (background elements such as the audience seats and billboards have a weight < 0.3). 803. Generate fused image fragments based on the repaired ROI region, the second ROI region, and the third ROI region; In this embodiment, the second ROI region (detail-dense area) is located using high bitrate parameters (15-20Mbps, σ²>1200, core target weight>0.8), and the third ROI region (non-core area) is determined using low bitrate parameters (3-5Mbps, σ²<500, background weight<0.3). Combined with the repaired ROI region, precise region division is achieved, and resource allocation is more targeted. When generating images through fusion, the core area image quality is optimized and highlighted, the high bitrate of the second ROI region preserves details, and the repaired ROI region solves problems such as blurring, improving the recognition of key information. The low bitrate of the third ROI region controls costs, reduces overall bitrate consumption, and reduces storage requirements. It is compatible with multiple terminals and complex networks, shortens the access cycle for new events, expands the audience, and balances viewing experience, resource efficiency, and scene adaptability, providing support for the dissemination and commercialization of event images.

[0025] The above describes a method for optimizing sports replay images according to an embodiment of the present invention. The following describes a device for optimizing sports replay images according to an embodiment of the present invention. Please refer to [link / reference]. Figure 9 One embodiment of the sports replay image optimization device of the present invention includes: Data acquisition module 1 is used to acquire playback requirements and event video data; Analysis module 2 is used to analyze the event video data according to the preset segment duration, playback requirements, preset acquisition conditions and preset image encoder to obtain the first ROI area and playback resolution. Image quality restoration module 3 is used to restore the image quality of the first ROI region according to the preset color channels, preset gradient variance and preset adjustment coefficient set, so as to obtain the restored ROI region. Image segment generation module 4 is used to generate fused image segments based on the repaired ROI region, preset high bitrate parameters, and preset low bitrate parameters. The frame rate adjustment module 5 is used to adjust the frame rate of the fused image segment according to the preset playback speed and playback resolution to obtain the playback image segment. In this embodiment, by analyzing playback requirements, the content, quality, and time requirements are accurately identified, providing direction for subsequent processing. Combining segment duration, acquisition conditions, and the image encoder, a first ROI region and an adapted resolution are generated, improving the accuracy of motion pixel recognition, reducing invalid data, and balancing image quality and efficiency. The first ROI region is repaired using gradient variance and adjustment coefficients, resolving issues such as blurring, loss of detail, and color cast, laying the foundation for fusion. Fusion repair of the ROI and high-bitrate and low-bitrate regions ensures image quality in the core area while controlling costs in the background area, resulting in natural image transitions. Finally, the frame rate is adjusted according to playback speed and resolution to ensure smooth playback. Overall, the system is adapted to multiple events and terminals, improving user experience, reducing bandwidth and storage costs, and facilitating efficient presentation and commercial application of event replays.

[0026] Figure 10 This is a schematic diagram of the structure of a sports replay image optimization device 900 provided in an embodiment of the present invention. This sports replay image optimization device 900 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 910 (e.g., one or more processors) and a memory 920, and one or more storage media 930 (e.g., one or more mass storage devices) storing application programs 933 or data 932. The memory 920 and storage media 930 can be temporary or persistent storage. The program stored in the storage media 930 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the sports replay image optimization device 900. Furthermore, the processor 910 may be configured to communicate with the storage media 930 and execute a series of instruction operations in the storage media 930 on the sports replay image optimization device 900 to implement the steps of the sports replay image optimization method provided in the above-described method embodiments.

[0027] A sports replay image optimization device 900 may further include one or more power supplies 940, one or more wired or wireless network interfaces 950, one or more input / output interfaces 960, and / or one or more operating systems 931, such as Windows Server, MacOSX, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 10 The illustrated structure of a sports replay image optimization device does not constitute a limitation on a sports replay image optimization device. It may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0028] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system, device, or unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0029] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for optimizing sports replay images, characterized in that, include: Acquire replay requirements and event video data; The event video data is analyzed based on the preset segment length, playback requirements, preset acquisition conditions, and preset image encoder to obtain the first ROI region and playback resolution. The image quality of the first ROI region is restored based on the preset color channels, preset gradient variance, and preset set of adjustment coefficients to obtain the restored ROI region. Generate fused image fragments based on the repaired ROI region, preset high bitrate parameters, and preset low bitrate parameters; The frame rate of the merged video clips is adjusted according to the preset playback speed and playback resolution to obtain the playback video clips.

2. The method for optimizing sports replay images as described in claim 1, characterized in that, The process of analyzing the event video data based on preset segment duration, playback requirements, preset acquisition conditions, and preset image encoders to obtain the first ROI region and playback resolution includes: The event video data was filtered based on the segment length and playback requirements to obtain key segment sequences; The first ROI region is extracted from the key segment sequence based on the acquisition conditions, the preset background difference threshold, and the preset inter-frame difference threshold. The scene complexity index is obtained by analyzing the key segment sequence based on the image encoder and the preset text encoder. The playback resolution is generated based on the scene complexity index.

3. The method for optimizing sports replay images as described in claim 2, characterized in that, The process of filtering event video data based on segment duration and playback requirements to obtain key segment sequences includes: Based on the preset deep learning model and playback requirements, target identification is performed on the event video data to obtain target type data, target location data, and target confidence data. The importance value of a segment is calculated based on a preset set of weights for target type data, target location data, and target confidence. The event video data was filtered based on segment duration and segment importance to obtain key segment sequences.

4. The method for optimizing sports replay images as described in claim 2, characterized in that, The step of analyzing key segment sequences based on an image encoder and a preset text encoder to obtain a scene complexity index includes: The key segment sequence is feature extracted based on the image encoder to obtain the image feature vector; The key segment sequence is feature extracted using a text encoder to obtain a text feature vector; The scene complexity index is generated based on image feature vectors and text feature vectors.

5. The method for optimizing sports replay images as described in claim 2, characterized in that, The process of generating playback resolution based on scene complexity index includes: The scene complexity index is analyzed to obtain the complexity analysis results; If the complexity analysis result indicates that the scene complexity index is within the preset low complexity range, then the first resolution is generated based on the scene complexity index. The playback resolution is generated based on the first resolution. If the complexity analysis result indicates that the scene complexity index is within the preset medium complexity range, then a second resolution is generated based on the scene complexity index. The playback resolution is generated based on the second resolution. If the complexity analysis result indicates that the scene complexity index is in the preset high complexity range, then a third resolution is generated based on the scene complexity index. The playback resolution is generated based on the third resolution. The first resolution is smaller than the second resolution; The second resolution is smaller than the third resolution.

6. The method for optimizing sports replay images as described in claim 2, characterized in that, The step of extracting the first ROI region from the key segment sequence based on the acquisition conditions, a preset background difference threshold, and a preset inter-frame difference threshold includes: Static frame data and reference frame data are collected from the key segment sequence according to the acquisition conditions; Feature extraction is performed on static frame data and reference frame data based on background difference threshold to obtain candidate moving pixels; Feature extraction is performed on static frame data and reference frame data based on the inter-frame difference threshold to obtain inter-frame motion pixels; Intersection analysis is performed on candidate moving pixels and inter-frame moving pixels to obtain the intersection moving pixels; The first ROI region is generated based on the intersection of moving pixels.

7. The method for optimizing sports replay images as described in claim 1, characterized in that, The step of performing image quality restoration on the first ROI region based on a preset color channel, a preset gradient variance, and a preset set of adjustment coefficients to obtain a restored ROI region includes: The key segment sequence is analyzed based on color channels, preset image gray levels, and gradient variance to obtain the inter-frame absolute difference and texture complexity. Image quality restoration is performed on the first ROI region based on the absolute difference between frames, texture complexity, and adjustment coefficient set to obtain the restored ROI region.

8. The method for optimizing sports replay images as described in claim 2, characterized in that, The step of generating fused image segments based on the repaired ROI region, preset high bitrate parameters, and preset low bitrate parameters includes: The second ROI region is obtained from the key segment sequence based on the high bit rate parameter; The third ROI region is obtained from the key segment sequence based on the low bit rate parameter; A fused image fragment is generated based on the repaired ROI region, the second ROI region, and the third ROI region.

9. A sports event replay image optimization device, characterized in that, include: The data acquisition module is used to acquire playback requirements and event video data; The analysis module is used to analyze the event video data according to the preset segment duration, playback requirements, preset acquisition conditions and preset image encoder to obtain the first ROI region and playback resolution. The image quality restoration module is used to restore the image quality of the first ROI region based on the preset color channels, preset gradient variance and preset adjustment coefficient set, so as to obtain the restored ROI region. The image segment generation module is used to generate fused image segments based on the repaired ROI region, preset high bitrate parameters, and preset low bitrate parameters. The frame rate adjustment module is used to adjust the frame rate of the fused image segments according to the preset playback speed and playback resolution to obtain the playback image segments.

10. A sports event replay image optimization device, characterized in that, The event replay image optimization device includes: a memory and at least one processor, wherein the memory stores instructions; At least one of the processors invokes the instructions in the memory to cause the sports replay image optimization device to perform the steps of the sports replay image optimization method as claimed in any one of claims 1-8.