Video clip method, apparatus, and storage medium based on ball hit detection

By using a video editing method based on ball-hitting detection, this method identifies ball sports scenes, performs frequency domain enhancement processing, and aligns the beats. This solves the problems of lack of scene-specific audio enhancement, low accuracy in identifying the ball-hitting moment, and poor audio-visual rhythm adaptation in ball sports video editing, achieving efficient and accurate video editing results.

CN121585776BActive Publication Date: 2026-04-28SHENZHEN EMEET TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN EMEET TECH CO LTD
Filing Date
2026-01-23
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing technologies, video editing for ball sports relies on manual editing and background music, resulting in poor editing quality, long editing time, and easy deviations in the selection of segments.

Method used

The video editing method based on ball-hitting detection determines the core frequency response range by identifying ball sports scenes, performs frequency domain enhancement processing, extracts audio features, identifies the moment of hitting the ball, aligns it with the background music beat, and dynamically adjusts the playback speed of the hitting segment.

Benefits of technology

It improves the audio clarity, accuracy of ball-hitting moment recognition, and audio-visual rhythm matching of ball sports videos, thereby enhancing editing efficiency and the audio-visual adaptation effect of the finished video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121585776B_ABST
    Figure CN121585776B_ABST
Patent Text Reader

Abstract

The application discloses a video clip method and device based on ball hitting detection and a storage medium, relates to the technical field of videos, and comprises the following steps: in response to a clip trigger instruction, determining a core frequency response range according to a ball game scene recognition result of a to-be-processed video; performing frequency domain enhancement processing on original audio of the to-be-processed video to obtain audio data; extracting audio features of each evaluation dimension and determining whether the audio features all meet corresponding audio quality standards; if yes, determining a ball hitting moment in the to-be-processed video; aligning the ball hitting moment with the beat of background music and adjusting the playing speed of a ball hitting segment to obtain a target video. According to the application, the core frequency response range is determined according to the ball game scene, the audio is enhanced in the frequency domain, the ball hitting moment is recognized, and then the ball hitting moment is aligned with the beat of the background music and the speed is adjusted, so that the problem that the effect of ball hitting video clipping is poor is solved, and the to-be-processed video clipping efficiency, ball hitting recognition accuracy and finished product audio-visual adaptation effect are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video technology, and in particular to a video editing method, device and storage medium based on ball hit detection. Background Technology

[0002] Athletes, coaches, amateur enthusiasts, and other groups often use video recording equipment to record ball games. The recorded videos are usually tens of minutes to several hours long. These videos need to be edited to extract the highlights of the game and make them into compilations for sharing, thereby reducing the spread of irrelevant content and increasing the viewing value.

[0003] In related technologies, manual editing and music addition are usually relied upon. This method requires a lot of time. It involves manually playing images and video materials repeatedly to identify the moment of impact, then manually cutting out the corresponding segments and adding music. This can easily lead to deviations in the selection of segments, resulting in poor editing of the impact video.

[0004] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main objective of this application is to provide a video editing method, device, and storage medium based on ball-hitting detection, aiming to solve the technical problem of poor video editing results.

[0006] To achieve the above objectives, this application proposes a video editing method based on ball hit detection, the method comprising:

[0007] In response to the editing trigger command of the video to be processed, the core frequency response range is determined based on the ball sports scene recognition result of the video to be processed;

[0008] Based on the core frequency response range, the original audio of the video to be processed is subjected to frequency domain enhancement processing to obtain audio data;

[0009] Extract the audio features of the audio data in each evaluation dimension, and determine whether the audio features all meet the corresponding audio quality standards;

[0010] If so, determine the moment of impact in the video to be processed based on the hit recognition rules;

[0011] Align the hitting moment with the beat of the background music to obtain a beat video;

[0012] The playback speed of the ball-hitting segments in the beat video is dynamically adjusted according to the interval between the beats to obtain the target video.

[0013] In one embodiment, in response to an editing trigger command for the video to be processed, key frames that are representative of the scene are extracted from the video to be processed uploaded by the user to form basic data;

[0014] The basic data is analyzed, and the site features and sports equipment features in the keyframes are identified based on the analysis results.

[0015] Based on the characteristics of the venue and the characteristics of the sports equipment, the ball game scenario is determined;

[0016] The mapping database between the scene and the core frequency response range is invoked, and the core frequency response range is determined based on the mapping relationship stored in the mapping database and the ball game scene.

[0017] In one embodiment, the original audio is separated from the video to be processed based on the core frequency response range;

[0018] Obtain pre-set spectrum energy weighting parameters, wherein the spectrum energy weighting parameters include a set gain weight corresponding to the signal within the core frequency response range, and a set attenuation weight corresponding to noise and redundant signals outside the core frequency response range;

[0019] According to the spectral energy weighting parameters, the energy values ​​of each frequency band of the original audio are adjusted in the time-frequency domain to obtain the audio data.

[0020] In one embodiment, candidate hitting points are located based on the transient energy peak detection results of the audio data;

[0021] By matching target frequency features, verifying the centroid of the spectrum, and analyzing the temporal morphology of the signal, interference signals in the candidate hitting points are eliminated, and the hitting time is determined.

[0022] In one embodiment, if not, output the dimensions that do not meet the quality standards and optimization suggestions;

[0023] The interactive options are output according to the aforementioned quality non-compliance dimensions and the aforementioned optimization suggestions, wherein the interactive options include at least one of re-uploading the video, adjusting the frequency domain enhancement parameters, and supplementing scene information;

[0024] Receive the user's selected interaction choice and execute the processing process corresponding to the selected interaction choice.

[0025] In one embodiment, the selected background music is subjected to rhythm analysis, and the timestamp corresponding to the beat is determined based on the analysis result;

[0026] Based on the deviation between the hitting time and the timestamp, the playback sequence of the hitting segment is adjusted to align the hitting time with the beat of the background music;

[0027] The video to be processed, with the playback timing of the ball-hitting segment adjusted, is used as the beat video.

[0028] In one embodiment, a weighted calculation result is obtained based on the transient peak significance and target frequency band energy ratio corresponding to the ball-hitting segment, as well as the preset weights of the transient peak significance and the target frequency band energy ratio;

[0029] From each of the aforementioned shot segments, a predetermined number of shot segments with the highest weighted calculation result value are selected as high-priority segments;

[0030] If the duration of the background music exceeds the duration of the high-priority segment, then extract the strong beat timestamps from the background music that are of high intensity and have the same number as the high-priority segment.

[0031] The high-priority segment's hit time is aligned with the strong shot timestamp one by one to obtain the beat video.

[0032] In one embodiment, the target video is pushed to an interactive interface, and feedback entry points for beat alignment accuracy, playback speed adaptation, and audio clarity are provided to obtain user input adjustment requirements;

[0033] Retrieve the preset technical parameter configuration mapping relationship corresponding to the adjustment requirement, and determine the technical parameter modification scheme that matches the adjustment requirement;

[0034] After optimizing the corresponding parameters according to the technical parameter modification scheme, the target video generation step is re-executed to obtain the optimized target video and push it to the interactive interface.

[0035] Furthermore, to achieve the above objectives, this application also proposes a video editing device based on ball-hitting detection, the video editing device based on ball-hitting detection comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the video editing method based on ball-hitting detection as described above.

[0036] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the video editing method based on ball hit detection as described above.

[0037] This application provides a video editing method based on ball-hit detection, including responding to an editing trigger command for the video to be processed, determining a core frequency response range based on the ball sports scene recognition result of the video to be processed, performing frequency domain enhancement processing on the original audio of the video to be processed based on the core frequency response range to obtain audio data, extracting audio features of the audio data in each evaluation dimension and determining whether these audio features all meet the corresponding audio quality standards. If so, determining the ball-hit moment in the video to be processed based on ball-hit recognition rules, aligning the ball-hit moment with the beat of the background music to obtain a beat video, and then dynamically adjusting the playback speed of the ball-hit segment in the beat video according to the interval between beats to obtain the target video. This application solves the technical problems of lack of scene specificity in audio enhancement, low accuracy of ball-hit moment recognition, and poor audio-visual rhythm adaptation in traditional ball sports video editing by using a collaborative technical solution of adaptive core frequency response range setting for ball sports scenes, audio quality standardization verification, accurate ball-hit moment recognition, beat alignment, and dynamic adjustment of segment playback speed. It improves the audio clarity, ball-hit moment recognition accuracy, and audio-visual rhythm matching degree of ball sports videos.

[0038] In summary, this application solves the technical problem of poor video editing effects by responding to editing instructions, determining the core frequency response range according to the ball sports scene, enhancing the audio in the frequency domain, identifying the moment of impact, and then aligning and speeding it with the background music beat. This improves the efficiency of video editing, the accuracy of impact recognition, and the audiovisual adaptation effect of the finished product. Attached Figure Description

[0039] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0041] Figure 1 This is a flowchart illustrating the first embodiment of the video editing method based on ball hit detection in this application;

[0042] Figure 2 This is a flowchart illustrating the fifth embodiment of the video editing method based on ball hit detection in this application;

[0043] Figure 3 This is a flowchart illustrating the eighth embodiment of the video editing method based on ball hit detection in this application;

[0044] Figure 4 This is the core block diagram of this application;

[0045] Figure 5 This is a schematic diagram of the video editing device based on ball-hitting detection in this application.

[0046] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0047] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0048] In related technologies, manual editing and music addition are usually relied upon. This method requires a lot of time. It involves manually playing images and video materials repeatedly to identify the moment of impact, then manually cutting out the corresponding segments and adding music. This can easily lead to deviations in the selection of segments, resulting in poor editing of the impact video.

[0049] This application provides a solution: First, in response to the editing trigger command of the video to be processed, the core frequency response range is determined based on the ball sports scene recognition result of the video to be processed. Then, the original audio of the video to be processed is subjected to frequency domain enhancement processing based on the core frequency response range to obtain audio data. Next, the audio features of the audio data in each evaluation dimension are extracted, and it is determined whether the audio features all meet the corresponding audio quality standards. If so, the hitting time in the video to be processed is determined based on the hitting recognition rules. Then, the hitting time is aligned with the beat of the background music to obtain a beat video. Finally, the playback speed of the hitting segment in the beat video is dynamically adjusted according to the interval between the beats to obtain the target video.

[0050] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device capable of performing the above functions, such as a video editing device based on ball impact detection. The following description uses a video editing device based on ball impact detection as an example to illustrate this embodiment and the subsequent embodiments.

[0051] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0052] This application provides a video editing method based on ball hit detection, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the video editing method based on ball hit detection in this application.

[0053] In this embodiment, the video editing method based on ball hit detection includes steps S10~S60:

[0054] Step S10: In response to the editing trigger command of the video to be processed, determine the core frequency response range based on the ball sports scene recognition result of the video to be processed.

[0055] In this embodiment, the editing trigger command refers to the operation signal that triggers the start of the video editing process. The ball sports scene recognition result refers to identifying the type of ball sports and related sound characteristics corresponding to the video. The core frequency response range refers to the main frequency distribution range of the target sound in the scene.

[0056] As an optional implementation, the video to be processed is uploaded and edited. Upon triggering a command, a pre-defined mapping set of ball sports scene recognition results and core frequency response ranges is retrieved. This mapping set pre-stores typical frequency distribution information of target sounds corresponding to various common ball sports. By identifying the type of motion scene corresponding to the input video, the corresponding frequency distribution interval is directly matched from the mapping set. The matched intervals are then smoothed at their boundaries to eliminate frequency abrupt changes at the interval connections, ultimately determining the suitable core frequency response range. This method has a simple process, can quickly output results, and is highly efficient, ensuring processing efficiency in common scenarios.

[0057] As an alternative implementation, the video to be processed is uploaded and edited. After triggering the command, the original audio data is first separated from the input video. A full-band spectrum scan is performed on the original audio data to extract the signal energy value and energy percentage of each frequency band. Candidate frequency ranges with energy percentages higher than a preset ratio and matching the sound characteristics of the target ball game are selected. Overlapping areas of these candidate frequency ranges are merged, and invalid frequency bands are removed. Then, combined with the frequency continuity characteristics of the target sound, the range boundaries are dynamically adjusted to finally define the suitable core frequency response range. This method dynamically determines the frequency response range through real-time spectrum analysis of the original audio, adapting to various conventional and special sports scenarios, achieving higher matching accuracy, and improving the targeting and effectiveness of audio enhancement.

[0058] Step S20: Perform frequency domain enhancement processing on the original audio of the video to be processed based on the core frequency response range to obtain audio data.

[0059] In this embodiment, the original audio refers to the original sound signal stream contained in the video to be processed. Frequency domain enhancement processing refers to the operation of selectively enhancing and suppressing noise in different frequency bands of the audio signal around the core frequency response range. Audio data refers to the optimized sound signal obtained after frequency domain enhancement processing.

[0060] As an optional implementation, the original audio is extracted from the video to be processed and converted to the time-frequency domain to obtain the signal distribution of each frequency band. A fixed gain rule is set according to the core frequency response range, and signals within the core frequency response range are amplified with a uniform amplitude. Signals outside the core frequency response range are attenuated by a fixed proportion, while environmental vibration noise in the ultra-low frequency band and electromagnetic interference signals in the ultra-high frequency band are filtered out. After processing, the time-frequency domain signal is restored to the time-domain signal to obtain the audio data. This method has a simple operation process, low resource consumption, and meets basic audio enhancement requirements.

[0061] As an alternative implementation, after extracting the original audio from the video to be processed, a full-band spectrum analysis is performed on the original audio to obtain the signal energy value and signal-to-noise ratio of each sub-band within the core frequency response range. Differential gain coefficients are dynamically generated based on the energy distribution differences of each sub-band. Higher gain is applied to sub-bands with weaker energy but acceptable signal-to-noise ratios, while moderate gain is maintained for sub-bands with sufficient energy. Simultaneously, an adaptive noise threshold algorithm is employed to detect and dynamically attenuate interference signals outside the core frequency response range in real time, avoiding excessive suppression that could lead to audio distortion. After processing, the audio data is obtained through time-domain reconstruction. This method accurately adapts to the signal characteristics of each sub-band within the core frequency response range, resulting in more targeted enhancement, more thorough noise suppression, and improved clarity and signal-to-noise ratio of the audio data.

[0062] Step S30: Extract the audio features of the audio data in each evaluation dimension, and determine whether the audio features all meet the corresponding audio quality standards.

[0063] In this embodiment, the evaluation dimension refers to a pre-defined category of core indicators used to measure audio quality. Audio features refer to the specific attribute parameters of audio data presented under each evaluation dimension. Audio quality standards refer to the pre-defined criteria for judging the pass / fail status of audio features under each evaluation dimension.

[0064] As an optional implementation, audio features corresponding to three evaluation dimensions—global effective signal-to-noise ratio, transient peak significance, and target frequency band energy proportion—are first extracted synchronously from the audio data. Then, the audio quality standards corresponding to each of the three features are retrieved all at once, and each audio feature is compared with its corresponding standard in a fixed order, with the compliance status of each feature recorded in real time. Finally, the comparison results of all evaluation dimensions are summarized to determine whether each audio feature meets its corresponding audio quality standard. This method has a compact operation process, can quickly complete the synchronous verification of multi-dimensional features, and has high processing efficiency.

[0065] As an alternative implementation, a differentiated verification priority is first assigned to three evaluation dimensions—global effective signal-to-noise ratio, transient peak significance, and target frequency band energy proportion—based on the impact of audio quality on subsequent processing. Then, corresponding audio features are extracted from the audio data, and each feature is compared with its corresponding quality standard in descending order of priority. If a high-priority feature fails to meet the standard, the audio quality is directly determined to be unacceptable; if a high-priority feature meets the standard, the next lower priority feature is verified. This process continues until all dimensions have been compared. This method highlights the core role of key evaluation dimensions, reduces the interference of secondary features on the judgment results, and achieves higher accuracy.

[0066] Step S40: If yes, determine the moment of impact in the video to be processed based on the hit recognition rules.

[0067] In this embodiment, the ball-hitting recognition rule refers to the judgment logic used to distinguish the ball-hitting signal from the interference signal in the audio data and to locate the time of the ball-hitting event. The ball-hitting time refers to the specific time node corresponding to the ball-hitting action in the audio data.

[0068] As an optional implementation, a preset hit recognition rule is first retrieved. This rule uses the transient features of audio data as the core criterion for judgment. Transient peak features for each time period are extracted from the audio data corresponding to the video to be processed. The extracted features are then matched one by one with the audio feature benchmarks in the hit recognition rule. The time points corresponding to successfully matched time periods are marked as the hit times. All marked time points are then summed to obtain the set of hit times in the video to be processed. This method relies solely on audio features for recognition, has a simple operation process, can quickly output hit time results, and has high processing efficiency.

[0069] As an alternative implementation, the method first retrieves the ball-hitting recognition rules, which include both audio and video features. Transient peak features are extracted from the audio data corresponding to the video to be processed, initially marking candidate ball-hitting moments. Then, the video frame features corresponding to each candidate ball-hitting moment are extracted, and the relative position and motion state of the racket and ball in the frame are analyzed. The analysis results of both audio and video features are simultaneously matched against the dual-feature benchmark in the ball-hitting recognition rules. Only candidate moments that successfully match both features are determined as valid ball-hitting moments. All valid ball-hitting moments are then aggregated to obtain the final result. This method combines audio and video dual-dimensional features for recognition, effectively eliminating false candidate moments caused by interference signals, resulting in higher recognition accuracy.

[0070] Step S50: Align the hitting moment with the beat of the background music to obtain a beat video.

[0071] In this embodiment, the corresponding beat of the background music is the beat position information extracted from the background music and used to match the moment of impact. Filling is the operation of placing the impact segments into the beat intervals. The beat video is the finished video formed by aligning the moment of impact with the beat of the background music and filling the beat intervals with impact segments.

[0072] As an optional implementation, all shot times are first arranged chronologically, while the corresponding beats of the background music are sorted in chronological order to ensure a one-to-one correspondence between shot times and beats. Then, the shot clip corresponding to each shot time is directly placed into the interval between the corresponding beats. After placing all shot clips, a beat video is obtained. This method has a simple operation process, high efficiency in beat alignment and clip filling, and can quickly generate basic beat videos, achieving the effect of rapidly generating basic beat videos and meeting the needs of common scenarios with high processing efficiency requirements.

[0073] As an alternative implementation, the duration of each shot segment corresponding to each shot moment is first calculated, while the interval duration between corresponding beats in the background music is measured, and the difference between the segment duration and the interval duration is compared. If the difference exceeds a preset range, the duration of the shot segment is fine-tuned without disrupting the smoothness of the shot, ensuring that the adjusted segment duration perfectly matches the beat interval duration. The adjusted shot segment is then placed into the corresponding beat interval, and after filling, the smoothness of the transition between adjacent segments is checked, correcting any abnormal transitions, ultimately resulting in a beat video. This method considers the adaptability of shot segment duration and beat interval, generating a beat video with smooth and natural audio-visual transitions, achieving a highly adaptable and smooth beat video effect, providing high-quality finished material for subsequent audio-visual collaborative optimization.

[0074] Step S60: Dynamically adjust the playback speed of the ball-hitting segments in the beat video according to the interval between the beats to obtain the target video.

[0075] In this embodiment, the hitting segment is a video clip containing the moment of hitting the ball. The interval between beats is the time difference between two adjacent corresponding beats. The playback speed is the speed at which the hitting segment is played. The target video is the final video with highly synchronized audio and video rhythm after dynamic adjustment of the playback speed.

[0076] As an optional implementation, the duration of all beat intervals is first statistically analyzed and the average duration is calculated. Then, using this average duration as a benchmark, the playback speed of all hitting segments in the beat video is uniformly adjusted so that the playback duration of each segment matches the average beat interval duration. The adjustment does not differentiate between the motion characteristics of different hitting segments or the duration differences of different beat intervals. After adjusting the speed of all segments, they are directly integrated to obtain the target video. This method has a simple adjustment logic, a fast operation process, and can quickly complete the speed adaptation of all segments. It has high processing efficiency and achieves the effect of quickly generating a basic target video, meeting the needs of common scenarios with high processing efficiency requirements.

[0077] As an alternative implementation, the specific duration of each beat interval is first measured. Then, a corresponding hitting segment is matched for each beat interval. Based on the difference between the original duration of the hitting segment and the corresponding beat interval duration, the playback speed of the segment is adjusted accordingly. During the adjustment process, the smoothness of the hitting motion is simultaneously checked to avoid distortion caused by excessive speed adjustments. After adjusting a single segment, other segments are processed sequentially. After all segments are adjusted, the rhythmic continuity of adjacent segments is checked, and abnormal deviations are corrected to obtain the target video. This method offers high adjustment accuracy, considers the adaptability of each segment and beat interval, and generates a target video with good audio-visual coordination, achieving a highly adaptable and smooth target video while ensuring the audiovisual quality of the target video.

[0078] For example, in a video editing scenario, in response to video editing instructions, the core frequency response range for the badminton scene (specific scene) is determined to be 2000 Hz to 8000 Hz (the main frequency band where the sound of the badminton racket hitting the shuttlecock is distributed). Based on this core frequency response range, a short-time Fourier transform is used to convert the video audio track to the time-frequency domain. A gain of 15 dB is applied to the 2000 Hz to 8000 Hz frequency band, and a attenuation of 10 dB is applied to the frequency bands below 2000 Hz and above 8000 Hz, respectively. Low-frequency noise and high-frequency interference from the environment are filtered out, and the audio data is restored to the time-domain signal. According to the preset three evaluation dimensions of signal-to-noise ratio, transient peak significance, and target frequency band energy ratio, the audio data features are extracted using a support vector machine model. The signal-to-noise ratio is 38 dB, the transient peak significance is 0.75, and the target frequency band energy ratio is 65%, all of which meet the preset standards (signal-to-noise ratio ≥ 35 dB, transient peak significance ≥ 0.7, target frequency band energy ratio ≥ 60%). Subsequently, based on the dynamic time warping algorithm combined with the energy peak detection hit recognition rules, the three hit times in the audio data were determined to be 0.8 seconds, 3.2 seconds, and 5.7 seconds. Background music with a tempo of 120 beats per minute (0.5-second intervals) was selected, and the hit times of the three hit segments were aligned with the timestamps of the 2nd, 7th, and 12th beats of the background music, respectively. The original 0.3-second hit segment was adjusted to 0.5 seconds, the 0.4-second segment to 0.5 seconds, and the 0.25-second segment to 0.5 seconds. These were then integrated in chronological order to obtain the target video.

[0079] By using scenario-based frequency domain enhancement, accurate ball-hitting recognition, and adaptive tempo adjustment, the problems of blind audio enhancement, inaccurate ball-hitting positioning, and poor audiovisual coordination in traditional ball game video editing have been solved, thus improving editing efficiency, ball-hitting recognition accuracy, and the professional quality of the target video.

[0080] Based on any of the above embodiments, in Embodiment 2 of this application, step S10 includes steps A11 to A13:

[0081] Step A11: In response to the editing trigger command of the video to be processed, extract key frames that are representative of the scene from the video to be processed uploaded by the user to form basic data.

[0082] In this embodiment, representative keyframes refer to video frames that can reflect the core scene characteristics of ball sports. Basic data refers to a dataset composed of selected keyframes arranged in chronological order.

[0083] As an optional implementation, after receiving video editing instructions, the system acquires the user-uploaded video to be processed and parses its frame sequence. Content features are extracted from each frame to identify whether it contains a moving subject, typical action postures, and core scene elements. Then, through inter-frame correlation analysis, the position of each frame in the action sequence or scene evolution is determined, and frames that can fully present the start, progress, and key changes of the core action are selected. Redundant frames with semantic repetition or that do not reflect core features are removed. The selected frames are then deduplicated, sorted, and their completeness is verified according to chronological order, ultimately forming the basic data. This method offers high targeting and accuracy in keyframe selection, and the basic data has strong scene representativeness, achieving the effect of obtaining high-quality, highly representative basic data.

[0084] Step A12: Analyze the basic data and identify the site features and sports equipment features in the keyframes based on the analysis results.

[0085] In this embodiment, field features refer to the appearance attributes of the sports field boundaries, ground markings, and area divisions presented in the keyframe. Sports equipment features refer to the core attributes of the shape, structure, and appearance of the equipment used by the sports subjects in the keyframe. Image recognition analysis refers to the operation of feature extraction, comparison, and judgment of the keyframe image content.

[0086] As an optional implementation, the site features and sports equipment features of each keyframe are extracted from the basic data. The extracted features are then compared one by one with a pre-defined feature template library for various ball sports scenes. This template library contains standard site appearance features and equipment shape features corresponding to various sports. The matching similarity between the extracted features and each template is calculated. The sports type corresponding to the template with the highest similarity and reaching a preset matching threshold is selected and directly identified as the ball sports scene corresponding to the video to be processed. This method has simple feature comparison logic, fast processing speed, and can quickly output scene recognition results.

[0087] As an alternative implementation, a layered feature extraction process is performed on keyframes in the basic data. First, shallow features such as the boundary shape, ground texture, and marker colors of the field are extracted. Then, deep features such as the structural proportions, contour details, and usage postures of the sports equipment are extracted. The shallow and deep features are then fused to form a comprehensive feature vector. Semantic association analysis is performed on this comprehensive feature vector using a pre-defined rule base for ball sports scenarios to determine its compatibility with various sports scenarios. Simultaneously, cross-validation of multiple keyframe features eliminates misjudgments caused by anomalies in single-frame features, ultimately determining the ball sports scenario corresponding to the video to be processed. This method offers comprehensive feature extraction, combines semantic analysis and cross-validation, achieves high recognition accuracy, and can adapt to scenarios with specific field or equipment styles.

[0088] Step A13: Call the mapping database between the scene and the core frequency response range, and determine the core frequency response range based on the mapping relationship stored in the mapping database and the ball game scene.

[0089] In this embodiment, the mapping database between scenes and core frequency response ranges refers to a database that pre-stores information on the association between various ball sports scenes and their corresponding adaptive core frequency response ranges.

[0090] As an optional implementation, a preset scenario-core frequency response range mapping database is invoked. This database stores the corresponding standard core frequency response ranges and applicable instructions categorized by ball sports scenarios. Based on the determined ball sports scenario name, a precise index search is performed in this mapping database to match the unique standard core frequency response range record corresponding to that scenario. The frequency interval data in this core frequency response range record is extracted, and the upper and lower boundaries of the intervals are format-validated to ensure data integrity and rationality. The frequency intervals that pass the validation are determined as the final core frequency response range. This method features simple query logic, fast data retrieval and validation processes, and rapid result output, ensuring processing efficiency in typical scenarios.

[0091] As an alternative implementation, a mapping database between scenes and core frequency response ranges is invoked. This database stores not only the basic core frequency response ranges for various scenes but also associated frequency fluctuation feature data of the target sound under different scenes. Based on a determined ball game scene, the corresponding basic core frequency response range is first retrieved, then the frequency fluctuation features associated with that scene are extracted. The upper and lower boundaries of the basic frequency response range are dynamically adjusted based on the fluctuation range. Simultaneously, historical adaptation and optimization records for that scene in the database are referenced to correct the adjusted interval, ensuring that the interval covers common frequency variations of the target sound in that scene, ultimately determining the core frequency response range. This method dynamically optimizes the frequency response range by incorporating scene frequency fluctuation features, resulting in stronger adaptability, higher accuracy, and improved effectiveness of audio enhancement.

[0092] For example, in a video editing scenario, noise reduction and enhancement are performed on the sound in the original video. Background noise in the environment is suppressed, highlighting the sound of the ball hitting the target. The core lies in dynamically adjusting the processing strategy according to the specific sports scenario (different ball types, courts, etc.). Preprocessing: Audio separation, separating the audio track from the video stream. Resampling and normalization, unifying the sampling rate, calculating the RMS (root mean square) value of the entire audio segment, and mapping it to a unified interval [-1, 1], eliminating volume differences caused by varying distances from the phone or microphone gain. Noise reduction and signal enhancement: The system presets the core frequency response range for different ball types in different scenarios, such as: Table tennis: crisp sound, high frequency (approximately 2kHz-5kHz). Tennis: muffled sound, including string vibration, low frequency and wide bandwidth (approximately 300Hz-1500Hz). The target frequency response range can be determined based on the input scenario. For tennis, different courts (clay, indoor, etc.) also correspond to different frequency response ranges. Input methods for ball game modes include, but are not limited to: user input via an interactive interface, and recognition via image classification algorithms (calling a lightweight image model to extract scene elements from the first frame or keyframe of the video). For different ball game modes, one method for noise reduction and signal enhancement is to set bandpass filter parameters using the frequency range corresponding to that ball game scene to suppress energy outside the core frequency response range. A second method uses STFT-based spectral energy weighting technology. The principle is to perform a short-time Fourier transform (STFT) on the signal to convert it to the time-frequency domain. The method is to scan frame by frame. If the energy in a certain frequency band is concentrated in the set "hitting frequency range," a gain coefficient >1.0 is given (amplification); if the energy is dispersed (e.g., white noise) or concentrated in the human voice frequency band (e.g., non-hitting harmonic structures in the 300Hz-3400Hz range), a gain coefficient <0.5 is given (attenuation). Finally, the sound is restored through inverse transform (iSTFT). A third method uses spectral fingerprint matching enhancement technology. The spectral fingerprint of the standard hitting sound (i.e., the power spectral density template under ideal conditions) is pre-stored. The cross-correlation function between the real-time audio stream and this fingerprint is calculated. When the frequency structure of the real-time signal closely matches the fingerprint, the correlation peak will be significantly enhanced; noise or unrelated impact sounds (such as rackets hitting the ground) are suppressed because their spectral structures do not match and their correlation is low.

[0093] By accurately locating motion scenes and mapping and matching core frequency response ranges through keyframe extraction and feature recognition, the problems of blind scene recognition and inaccurate frequency response range adaptation in traditional ball game video editing are solved, thus improving scene recognition efficiency and the targeting of subsequent audio enhancement.

[0094] Based on any of the above embodiments, in Embodiment 3 of this application, step S20 includes steps B11 to B13:

[0095] Step B11: Based on the core frequency response range, extract the original audio from the video to be processed.

[0096] As an optional implementation, the video to be processed is first split into audio and video streams to obtain the original audio signal containing all frequency bands. This original audio signal is then converted to the frequency domain to obtain the signal distribution of each frequency band. A filtering boundary is defined based on the core frequency response range, retaining all signal components within this range and eliminating frequency band signals outside the core frequency response range. Subsequently, the filtered frequency domain signal undergoes time-domain restoration processing, while simultaneously performing signal integrity verification to obtain the original audio. This method has a simple operation process, requires no complex signal analysis, has a fast processing speed, and low resource consumption.

[0097] As an alternative implementation, in the scene-adaptive audio enhancement module, after loading the initial core frequency response range corresponding to the motion scene, the frequency and energy fluctuation of the transient signals of ball impact in the audio stream are monitored in real time. When a preset number of transient signals of ball impact are detected within a first preset time period, the upper and lower boundaries of the initial core frequency response range are automatically expanded to fully cover the harmonic components of the ball impact sound during that period. When no effective transient signals of ball impact are detected within a second preset time period, the initial core frequency response range is automatically reduced to a preset proportion of the initial range, while the signal gain weight within the initial core frequency response range is increased to suppress environmental noise, thus obtaining the original audio. This method dynamically adapts to changes in the signal characteristics of the audio stream, improving the accuracy of audio enhancement at different times while maintaining scene adaptability.

[0098] Step B12: Obtain the preset spectrum energy weighting parameters, wherein the spectrum energy weighting parameters include the preset gain weights corresponding to signals within the core frequency response range, and the preset attenuation weights corresponding to noise and redundant signals outside the core frequency response range.

[0099] In this embodiment, redundant signals are signals outside the core frequency response range that are unrelated to the hitting event. Attenuation weights refer to the weighting configuration used to reduce noise and redundant signal energy outside the core frequency response range. Spectral energy weighting parameters refer to the set of parameters used to adjust the signal energy of each frequency band after integrating gain weights and attenuation weights.

[0100] As an optional implementation, after determining the upper and lower boundaries of the core frequency response range, all signals within this range are uniformly assigned the same gain weight according to fixed rules to ensure overall energy enhancement of the core frequency band. Simultaneously, all frequency bands below and above the core frequency response range are uniformly assigned the same attenuation weight to achieve overall suppression of noise and redundant signals outside the range. All the assigned gain and attenuation weights are then arranged and integrated in frequency band order to form a complete spectral energy weighting parameter. This method features simple weight setting logic, a fast operation process, rapid parameter generation, and high processing efficiency.

[0101] As an alternative implementation, the signal within the core frequency response range is first divided into sub-bands, and the signal energy and signal-to-noise ratio (SNR) of each sub-band are analyzed. Sub-bands with weaker energy but acceptable SNR are assigned higher gain weights, while sub-bands with sufficient energy are assigned moderate gain weights to ensure balanced signal energy enhancement within the core frequency range. Next, the interference signal intensity of each band outside the core frequency response range is analyzed. Bands with high noise intensity are assigned higher attenuation weights, while bands with weak noise are assigned moderate attenuation weights. Finally, the gain weights of each sub-band are correlated and verified with the attenuation weights outside the range to ensure smooth weight transitions, and these are integrated to form a spectral energy weighting parameter. This method's weighting configuration closely matches the actual characteristics of the signal, offering high targeting and accuracy, enabling fine-grained energy adjustment, and improving the clarity and SNR of the audio signal.

[0102] Step B13: According to the spectral energy weighting parameters, adjust the energy values ​​of each frequency band of the original audio in the time-frequency domain to obtain the audio data.

[0103] In this embodiment, the time-frequency domain is a signal space that simultaneously reflects the changes in the time dimension and the frequency distribution of the original audio. The original audio refers to the unoptimized initial audio signal of the video to be processed. The energy value of each frequency band refers to the signal energy of the original audio in different frequency ranges.

[0104] As an optional implementation, the spectral energy weighting parameters are first retrieved to determine the weighting scheme for each frequency band. Then, the original audio is converted to the time-frequency domain and divided into different frequency bands. The energy values ​​of each frequency band are adjusted synchronously according to the weighting parameters: the energy of core frequency bands is increased using a gain-type weighting, while the energy of non-core frequency bands is decreased using an attenuation-type weighting. During the adjustment process, the energy ratio of each frequency band remains relatively stable. After the energy adjustment of all frequency bands is completed, the time-frequency domain signal is converted back to the time domain to obtain the audio data. This method has a direct adjustment logic, can quickly complete the energy optimization of the entire frequency band, and has high processing efficiency.

[0105] As an alternative implementation, the original audio is first decomposed into a time-frequency domain signal, capturing the instantaneous frequency and energy fluctuation characteristics of the signal segment by segment. Then, combined with spectral energy weighting parameters, the energy of each frequency band at different times is dynamically adjusted. During energy peak periods, the gain weight is appropriately reduced to avoid distortion, while during energy trough periods, the gain of the core frequency band is strengthened. Simultaneously, the smoothness of energy transitions between frequency bands after adjustment is verified, and energy abrupt change areas are corrected to ensure overall signal stability. After completing the dynamic adjustment across all time periods, the signal format is converted to obtain the audio data. This method adapts to instantaneous signal changes, has high adjustment accuracy, avoids energy imbalance problems, achieves high-precision, low-distortion audio optimization effects, and ensures the integrity and clarity of the target signal.

[0106] As an alternative implementation, the adjusted time-frequency domain signal is divided into signal blocks according to the original time division. Fine-grained time-domain reconstruction is performed on each signal block, with real-time monitoring of signal phase and amplitude changes during the reconstruction process. Subsequently, a transition processing method is applied to the reconstruction results of adjacent signal blocks, eliminating inter-block transition traces through signal superposition and fusion. Then, global noise reduction optimization is performed on the overall reconstructed signal to further remove residual weak interference signals. Finally, signal integrity and smoothness are verified; after passing the verification, enhanced audio data is output. This method offers high reconstruction accuracy, preserves core signal details to the greatest extent, and produces higher quality output audio data.

[0107] For example, in the context of video editing, the audio effectiveness assessment module is the core risk control checkpoint of this invention. Its function is to perform a "check-up" on the enhanced signal before the time-consuming fine-grained positioning calculation. If the signal quality does not meet the standard, the process is directly interrupted and feedback is given to the user to avoid outputting substandard results. Feature extraction unit: Global Effective Signal-to-Noise Ratio (SNR), which measures the ratio of the potential target signal strength to the background noise level in the entire audio segment. Calculation method: First, the audio signal is processed by frame segmentation, and the short-time energy E(i) of each frame is calculated. The signal power P_signal is defined as the average energy of the first K% (e.g., the first 5%) frames (representing the suspected hitting moment); the noise power P_noise is defined as the average energy of the last M% (e.g., the last 20%) frames (representing the ambient noise level). The SNR calculation formula is as follows: SNR_global=10*log10(P_signal / (P_noise+ε)), where ε is a minimum value to prevent the denominator from being zero. Transient peak significance measures the prominence of a local peak in a time-domain waveform relative to its surrounding background, used to distinguish between a "sharp hit sound" and "gentle environmental fluctuations." It evaluates the average prominence of all transient sounds throughout the video. Calculation method: A sliding window of length W is defined. For a local maximum point x(t_p) within the window, its significance P_prom is calculated: P_prom = |x(t_p)| - μ_local, or the Peak-to-Average Ratio (PAR) can be used: PAR = 20 * log10(|x(t_p)| / μ_local), where μ_local is the mean or median of the signal amplitude excluding the peak region within the sliding window. P_prom will still be used hereafter. P_prom can be replaced with PAR in the following descriptions. The hit sound is a very short pulse, and the P_prom value should be significantly higher than non-transient noise such as human voice or wind noise. Subsequently, the statistical distribution (e.g., 90th percentile or mean μ_prom) of P_prom for all identified local maxima points in the entire audio stream is calculated. The final output is an indicator representing overall sharpness: the mean μ_prom of P_prom. The target frequency band energy proportion measures the concentration of signal energy in the frequency domain, verifying whether the current signal conforms to the spectral characteristics of the preset ball-playing pattern. It evaluates the average concentration of sound energy within the target ball-playing frequency range throughout the entire video. Calculation method: Perform a Fast Fourier Transform (FFT) on the suspected hitting frames to obtain the spectrum X(k). Based on the core frequency response range [f_min, f_max] determined by the aforementioned "scene adaptation module" (e.g., 2kHz-5kHz for table tennis), calculate the ratio of energy within this frequency band to the total energy of the entire frequency band.

[0108] R_spectral=Σ(|X(k)|^2)[k_min to k_max] / Σ(|X(k)|^2);

[0109] Where k_min and k_max correspond to the frequency indices of the lower and upper frequency limits f_min and f_max, respectively, and Σ represents summation. Judgment criteria: Valid ball-hitting sounds have specific formants and a higher R_spectral value; while broadband noise (such as wind noise) or erroneous sound sources (such as clapping) have a more dispersed energy distribution and a lower R_spectral value. Subsequently, the statistical distribution of all these R_spectral values ​​(e.g., 80th percentile or mean μ_spectral) is calculated, and the final output is an index representing the overall frequency domain purity: the mean μ_spectral of R_spectral. Judgment and decision unit: When the global effective signal-to-noise ratio is below the threshold, the result is directly output indicating that the audio quality does not meet the requirements (judging whether the overall acoustic environment is poor); when the signal-to-noise ratio is not less than the threshold, the weighted result of transient peak significance and target frequency band energy proportion is further compared with the threshold; if it is below the threshold, the audio quality is considered unacceptable. Deep learning-based audio quality discrimination methods: Mel-Frequency Cepstral Coefficients (MFCC) or Log-Mel Spectrogram are calculated using a sliding window. Lightweight Convolutional Neural Networks (CNNs), such as MobileNetV3 or EfficientNet-Lite, are used with a classification head to classify audio for each window.

[0110] Specifically, for the characteristics of "highlight extraction" in ball sports, a simple average cannot be used (because there may be a large amount of invalid time spent retrieving the ball in the video). Instead, the "effective segment ratio method" is adopted. Effective frame count: The number of slices in the sequence with a score exceeding the high confidence threshold (e.g., 0.7), denoted as N_good. Usability Ratio calculation: Ratio = N_good / N_total. Final decision logic: If Ratio > T_ratio (e.g., 0.1): Conclusion: PASS (quality meets the standard). Explanation: As long as 10% of the time in the video is clear and usable (highlights can be extracted), we consider the video valid and can proceed with further processing. If Ratio <= T_ratio: Conclusion: FAIL (quality does not meet the standard). Explanation: The entire video is almost entirely noise, and there are not enough high-quality segments to support highlight generation.

[0111] By using scenario-based weight configuration and precise frequency domain energy adjustment, the problems of weak audio enhancement and incomplete noise suppression in traditional video editing are solved, improving the clarity and signal-to-noise ratio of ball sports audio, and providing high-quality signal support for subsequent ball-hitting moment recognition.

[0112] Based on any of the above embodiments, in Embodiment 4 of this application, step S40 includes steps C11~C12:

[0113] Step C11: Locate the candidate hitting point based on the transient energy peak detection result of the audio data.

[0114] In this embodiment, the transient energy peak detection result is a set of records formed after detecting the signal characteristics of a sudden increase in instantaneous energy in the audio data. The candidate hitting point is the time node that is preliminarily determined to be the time point where the corresponding hitting action occurred.

[0115] As an optional implementation, the transient energy peak detection results corresponding to the audio data are first retrieved. The occurrence time and energy value of all transient energy peaks in the results are extracted. A uniform peak filtering threshold is set, and the time points corresponding to all transient peaks with energy values ​​exceeding this threshold are directly marked as candidate hitting points. Then, all marked candidate hitting points are sorted in chronological order to form a complete set of candidate hitting points. This method has a simple and intuitive operation process, requires no complex segmentation processing, and can quickly complete the batch location of candidate hitting points, resulting in high processing efficiency.

[0116] As an alternative implementation, the transient energy peak detection results corresponding to the audio data are first divided into consecutive time periods in chronological order. The distribution characteristics and energy fluctuation range of the transient energy peaks within each time period are extracted, and differentiated dynamic screening thresholds are set for each time period based on the characteristics of each time period. Then, transient peaks with energy values ​​exceeding the corresponding dynamic thresholds within each time period are marked as preliminary candidate hitting points. Subsequently, the time interval characteristics between each preliminary candidate hitting point and its adjacent peaks are verified. Preliminary candidate points whose interval characteristics do not conform to the hitting action pattern are eliminated, and the remaining preliminary candidate points are determined as final candidate hitting points and sorted by time to form a set. This method fully considers the differences in energy distribution across different time periods. Through dual screening using dynamic thresholds and interval verification, it effectively reduces the inclusion of invalid candidate points, resulting in higher positioning accuracy and achieving high-precision positioning of candidate hitting points. This provides high-quality candidate data support for the subsequent verification of effective hitting moments.

[0117] Step C12 involves eliminating interference signals from the candidate hitting points and determining the hitting time by matching target frequency features, verifying the spectral centroid, and analyzing the signal time-domain morphology.

[0118] In this embodiment, target frequency feature matching refers to the process of comparing the frequency features of the signal corresponding to the candidate hitting point with the preset frequency features of the hitting signal. Spectral centroid verification is the operation of checking whether the spectral centroid features of the signal corresponding to the candidate hitting point conform to the standard for the spectral centroid of the hitting signal. Signal temporal morphological analysis is the process of analyzing the waveform, duration, and other morphological features of the signal corresponding to the candidate hitting point in the time dimension. The candidate hitting point is the time node where the suspected corresponding hitting action is initially located. Interference signals are signals mixed in with the candidate hitting points that are unrelated to the hitting action.

[0119] As an optional implementation, the signal features corresponding to all candidate hitting points are first extracted, and three operations are performed simultaneously: target frequency feature matching, spectral centroid verification, and signal time-domain morphology analysis. A preset unified judgment standard is invoked; candidate hitting points whose results meet all three standards are judged as valid hitting times, while those whose results do not meet any standard are directly marked as interference signals and eliminated. All valid hitting times that pass the three checks are summarized and sorted chronologically to form the final set of hitting times. This method performs the three checks in parallel, eliminating the need for step-by-step waiting, significantly shortening the overall verification time, and exhibiting high processing efficiency, thus meeting the requirements of preliminary screening scenarios with high processing efficiency requirements.

[0120] As an alternative implementation, target frequency feature matching is first set as the first priority verification method. Frequency features corresponding to candidate hitting points are extracted and matched with preset hitting frequency features. Candidate points that pass the matching enter the second priority spectral centroid verification stage to verify whether their spectral centroid features meet the standard. Verified candidate points then enter the third priority signal time-domain morphology analysis stage to analyze their waveform and duration characteristics in the time dimension. Only candidate hitting points that pass all three levels of progressive verification are determined as valid hitting moments; candidate points that fail any level of verification are judged as interference signals and eliminated. Finally, all valid hitting moments are summarized into an ordered set. This method, through hierarchical progressive verification, highlights the dominant role of core indicators, effectively avoiding misjudgments caused by deviations in non-core indicators. It achieves higher accuracy in interference signal elimination, providing accurate and reliable time node basis for subsequent core processes such as beat alignment.

[0121] As an alternative implementation, in the audio validity evaluation module, if the overall sound quality score does not reach a threshold, the distribution of candidate hitting points within the audio is further identified: if the number of candidate points is greater than or equal to a preset value, the sound quality of each candidate point's local time segment is scored separately, and the percentage of candidate points meeting the local score standard is calculated. If the percentage is greater than the remaining preset percentage, the audio is determined to be valid overall, and the user feedback mechanism is skipped to proceed to the next process. If the percentage is insufficient, user feedback is triggered again. This method avoids high-quality hitting segments being mistakenly judged as invalid due to noise outside the hitting period, thus improving the rationality of audio evaluation.

[0122] As an alternative implementation, in the ball-hitting event recognition and anti-interference module, after locating the candidate hitting point, the system recalls several frames of the video corresponding to that time point to identify the racket's trajectory and the ball's position change: if the racket is in a swinging trend in the previous frame and the distance between the ball and the racket is continuously decreasing, then the candidate point is determined to be a valid hit. If the racket has no swinging action or the distance between the ball and the racket does not change in the previous frame, then the candidate point is eliminated. This method combines video motion features for supplementary verification, reducing misidentification caused by relying solely on audio.

[0123] As an alternative implementation, feature parameters of each candidate hitting point are extracted in three dimensions: target frequency characteristics, spectral centroid, and signal temporal morphology, constructing a comprehensive feature vector. Multi-dimensional fusion judgment rules are then established, and an overall adaptability analysis is performed on the comprehensive feature vector. Candidate points whose single-dimensional features slightly deviate from the standard but whose overall characteristics conform to the hitting signal are retained. Simultaneously, through cross-validation of the feature correlation between adjacent candidate points, isolated interference points without reasonable feature correlation are eliminated. Finally, the candidate points that pass the fusion judgment and cross-validation are time-calibrated to determine the final hitting time. This method offers greater flexibility and comprehensiveness in verification, maximizing the retention of effective candidate points, accurately eliminating interference signals, and achieving higher positioning accuracy. Ultimately, it achieves high-precision determination of the hitting time, ensuring the audiovisual quality of the target video.

[0124] Exemplarily, in the scenario of video clip, the hitting event recognition and anti-interference module receives the audio signal passed through quality discrimination, aiming to accurately extract the hitting moment from a complex acoustic environment, remove interference, and avoid misrecognition of hitting events. Generate candidate hitting events: For the purpose of energy envelope preprocessing, the original audio with high-frequency oscillation is converted into a smooth curve reflecting the instantaneous power change to prepare for peak searching. Implementation method: Input the enhanced audio signal x(t) and calculate its energy envelope E(t). Specifically, the Hilbert Transform can be used to take the modulus, or the Rolling RMS can be adopted: E(t) = Sqrt(Mean(x(τ)^2)), where τ is a tiny window near t (such as 5 - 10 ms). Adaptive threshold peak searching: Considering different recording distances, a fixed threshold is not used. Calculate the noise baseline Baseline_noise within the current analysis window (such as 1 second). Set the dynamic threshold: H_threshold = Baseline_noise + Offset_dynamic. When the height of the local maximum > H_threshold, this point is recorded as a "candidate hitting point". Remove interference events: This module adopts a "feature cascade filtering" mechanism. For the candidate hitting points preliminarily screened out, they are sequentially passed through the following three checkers. If any step fails, it is regarded as interference signal and removed. The principle of the target frequency feature checker: Use the preset ball frequency range (such as 2k - 5k for table tennis) to remove the impact sounds with inconsistent timbres. It includes the following steps: Interception: Take a very short audio segment (such as 30 ms) centered on the candidate hitting point. Transformation: Perform Fourier Transform (FFT) on the intercepted segment to obtain the spectral energy distribution. Extract the dominant frequency: Find the frequency point with the maximum energy in the spectrum, that is, the dominant frequency (F_dom, DominantFrequency). Judgment: Obtain the preset frequency interval [F_min, F_max] of the current scene. If F_dom < F_min or F_dom > F_max: Judgment result = remove (reason: frequency mismatch, may be human voice or footsteps). If F_min <= F_dom <= F_max: Judgment result = retain. The principle of the source distance discriminator: Distinguish "hitting in this field" from "hitting in the adjacent field". Utilize the high-frequency attenuation characteristic of sound propagation. The transient characteristics of far-field sound are smoothed due to multiple reflections. Feature: RiseTime. It is defined as the time span required for the waveform envelope to rise from 10% to 90% of the peak. Hitting in this field is the direct sound, and the waveform jumps instantaneously; hitting in the adjacent field is reflected, and the jump is softer. Judgment: Calculate the time T_rise required for the envelope to rise from 10% to 90% of the peak. If T_rise > T_limit (such as 8 ms), it is regarded as far-field interference and removed. The principle of the spectral centroid checker: Use the "geometric centroid" of the spectrum to distinguish "crisp sound" from "dull sound".Solve the interference of dull overall sound perception even though the main frequency is within the range (such as the sound of a ball hitting the ground, the sound of a sneaker stamping). Feature: Spectral Centroid. It reflects the "brightness" of the sound. The centroid of the hitting sound is higher, and the centroid of the landing sound is lower. Calculation: Based on the spectrum in step (1), calculate the centroid C_spec: C_spec = Sum(f * A(f)) / Sum(A(f)), (where f is the frequency value, A(f) is the energy amplitude corresponding to that frequency, and Sum represents summation). Judgment: Set the centroid threshold F_center_limit (for example, 800Hz - 1000Hz). If C_spec < F_center_limit: Judgment result = Reject (Reason: The energy is concentrated in the low frequency, the sound is dull, and it is judged as the sound of landing or footsteps). If C_spec >= F_center_limit: Judgment result = Keep. Hitting event monitoring method based on deep learning: Different from the above method of generating candidate hitting events + proposing interference events, another feasible solution is a hitting event detection method based on CRNN (Convolutional Recurrent Neural Network). First, use CNN feature extraction, and then input the output of CNN into a bidirectional long short-term memory network (Bi-LSTM) or GRU. Output a probability curve P(t) aligned with the input time frame. P(t) represents the probability (0.0 - 1.0) of a hitting event occurring at time t. A threshold (for example, 0.8) can be set for the output probability curve P(t) of the model. When P(t) > 0.8, take the time point with the maximum probability within this section as the hitting moment. Considering that running a deep learning model on a mobile phone may consume power, a cascaded architecture of DSP coarse screening + DL fine detection, that is, "energy detection coarse screening + neural network fine classification" can be used as the preferred solution. First, use the DSP method with small computational complexity to locate the suspected time points, and then intercept tiny audio slices and input them into a lightweight network for secondary confirmation. This method takes into account both real-time performance and high accuracy. The specific explanation is as follows: First, use the method in the "Generate candidate hitting events" section to find possible candidate points, and then, centered on each candidate point, intercept a small segment of audio (for example, 100ms before and after), and send it into a miniature CNN classifier to determine whether it is a real hit. Even if a deep learning model is not used and only the above three verifiers are used, it is also done on the segments near the candidate points, which saves more time and power compared to applying the verifiers to the entire video.

[0125] Due to multi-dimensional feature verification and precise peak screening, the problem of vulnerable interference and inaccurate positioning in identifying the hitting moment in ball game video clips is solved, improving the accuracy and reliability of hitting moment identification, and providing a solid foundation for subsequent audiovisual collaborative editing.

[0126] Based on any of the above embodiments, in Embodiment 5 of the present application, refer to Figure 2 , Figure 2 This is a flowchart illustrating the fifth embodiment of the video editing method based on ball hit detection in this application. Following step S30, steps D11-D13 are also included:

[0127] Step D11: If not, output the dimensions that do not meet the quality standards and optimization suggestions.

[0128] In this embodiment, the "non-compliant dimension" refers to the evaluation dimension that fails to meet the audio quality standard. The optimization suggestion refers to specific guidance proposed for improving audio quality for the non-compliant dimension.

[0129] As an optional implementation, after confirming that the audio features do not meet the audio quality standards, a feature source analysis is first performed on each non-compliant dimension to determine whether the non-compliance is due to signal acquisition issues, noise interference, or improper pre-processing parameters. Then, based on the analysis results, personalized optimization suggestions are customized for each non-compliant dimension, specifying the detailed operation steps, parameter adjustment ranges, and precautions. The impact weight of each non-compliant dimension is also marked, prioritizing dimensions with a greater impact on subsequent processing. The non-compliant dimension name, cause analysis, impact weight, and personalized optimization suggestions are integrated into structured feedback content and presented to the user in a logical order. This method accurately identifies the root cause of non-compliance, provides highly targeted and actionable optimization suggestions, and improves the efficiency and success rate of subsequent audio data optimization.

[0130] Step D12: Output interactive options according to the dimensions of substandard quality and the optimization suggestions, wherein the interactive options include at least one of re-uploading the video, adjusting the frequency domain enhancement parameters, and supplementing scene information.

[0131] In this embodiment, "re-upload video" refers to the user's option to replace the original video file. "Adjust frequency domain enhancement parameters" refers to the user's option to modify audio processing parameters such as core frequency response range and weight configuration. "Supplement scene information" refers to the user's option to refine the detailed description of the ball game scene. "Interactive selection" refers to the set of executable operation options provided to the user.

[0132] As an optional implementation, this method first summarizes the dimensions that fail to meet quality standards and their corresponding optimization suggestions. It then analyzes the compatibility between each failing dimension and the interactive options, matching a specific interactive option to each failing dimension while retaining general interactive options. These options are then sorted in order of specific options first, followed by general options, with each interactive option labeled with its corresponding compatible dimension and expected optimization. An interactive option list is generated and output. The specific problems of each failing dimension and the improvement directions after selecting the corresponding interactive option are simultaneously displayed, facilitating users to make precise selections based on the problems. This method offers highly targeted interactive selections, guiding users to quickly match suitable solutions, reducing selection difficulty, improving user efficiency, and achieving a precise, guided interactive selection output effect, thus meeting users' needs for efficiently locating optimization solutions.

[0133] As an alternative implementation, all dimensions failing to meet quality standards are first identified and integrated to form a problem overview. Then, core optimization directions are extracted, corresponding to three basic interactive options: re-uploading the video, adjusting frequency domain enhancement parameters, and supplementing scene information. Users can simultaneously select multiple interactive options for optimization combinations. The output provides suggested combinations of interactive options, indicating the optimization priority and potential quality effects of different combinations. A custom entry point is also provided to support users adding personalized needs. The failing dimensions are simultaneously associated with the suggested combinations for user reference and decision-making. This method supports multi-option combination optimization, offering high flexibility and adaptability to the optimization needs of complex quality problems. The reserved custom entry point enhances adaptability, enabling flexible combination-based interactive selection output effects to meet diverse and personalized quality optimization needs.

[0134] Step D13: Receive the user's selected interaction choice and execute the processing process corresponding to the selected interaction choice.

[0135] In this embodiment, the user-selected interaction choice is a specific operation instruction chosen by the user from the output optimization options to improve audio quality. The processing procedure is a pre-set technical process for each type of interaction choice, used to resolve substandard quality issues.

[0136] As an optional implementation, this method receives a single interactive selection command chosen by the user, identifies the type of the command, and if it is to re-upload the video, triggers the re-execution of the entire processing flow, starting from the beginning with all steps of scene recognition, core frequency response range determination, audio frequency domain enhancement, and audio quality verification. If it is to adjust the frequency domain enhancement parameters, the new parameter configuration is extracted, and frequency domain enhancement processing is performed on the original audio based on this parameter configuration, followed by audio quality verification and subsequent processes. If it is to supplement scene information, the method receives the user-supplied information and updates the scene data, redetermines the core frequency response range based on the updated scene data, and performs subsequent processing. The corresponding processing steps are advanced sequentially according to the command type, and the processing result is output upon completion. This method only supports the execution of a single command, has a simple process logic, and fast command recognition and process triggering speed. It can quickly respond to the user's basic optimization needs, achieve rapid response to single interactive selection processing effects, and meet the user's simple and direct quality optimization requirements.

[0137] As an alternative implementation, the method first receives one or more interactive selection commands selected by the user. These commands are then prioritized according to preset rules, with re-uploading the video set as the highest priority, supplementing scene information as the second highest priority, and adjusting frequency domain enhancement parameters as the lowest priority. The corresponding processing steps are executed sequentially from highest to lowest priority. If the user selects both re-uploading the video and other options, the complete process for re-uploading the video is executed first, embedding the operations for supplementing scene information and adjusting frequency domain enhancement parameters within this process. If the user selects both supplementing scene information and adjusting frequency domain enhancement parameters, the processing flow for supplementing scene information is executed first, followed by the re-enhancement operation for adjusting frequency domain enhancement parameters based on the updated scene data. After all command-related processes are completed, a unified processing result is output. This method supports the combined execution of multiple commands, has clear priority division, can adapt to diverse and complex user optimization needs, offers higher processing accuracy, achieves precise response to combined interactive selections, and meets users' comprehensive quality optimization demands.

[0138] For example, in a video editing scenario, the audio feature set generated based on enhanced badminton audio data has a global effective signal-to-noise ratio of 32 dB (preset standard ≥ 35 dB), a target frequency band energy ratio of 60% (preset standard ≥ 65%), and a transient peak significance of 0.72 (preset standard ≥ 0.7). The result is that the audio features do not meet the audio quality standards, and the user is given feedback on the specific non-compliant dimensions: "global effective signal-to-noise ratio, target frequency band energy ratio," with optimization suggestions: "enhance the signal energy of the core frequency response range and reduce environmental noise interference." Based on these non-compliant dimensions and optimization suggestions, an interactive option is provided: "Re-upload the video; adjust the frequency domain enhancement parameters; supplement scene information." The user selects "adjust the frequency domain enhancement parameters" and adjusts the core frequency response range (originally 1800-8000 Hz) to 2000-8500 Hz, increasing the gain weight by 20%. After receiving this interactive selection, the frequency domain enhancement processing is re-executed to obtain optimized audio data. Features were extracted again according to the evaluation dimensions to generate a feature set with a global effective signal-to-noise ratio of 38 dB, a target frequency band energy ratio of 68%, and a transient peak significance of 0.75. The feature set met the preset standards and qualified audio data was output for subsequent steps to determine the ball's striking moment.

[0139] By providing accurate feedback on substandard information and through interactive optimization, the problem of editing workflow interruptions caused by substandard audio quality has been solved, improving the audio data compliance rate and the continuity and flexibility of video editing.

[0140] Based on any of the above embodiments, in Embodiment Six of this application, step S50 includes steps E11 to E13:

[0141] Step E11: Perform rhythm analysis on the selected background music and determine the timestamp corresponding to the beat based on the analysis results.

[0142] In this embodiment, rhythm analysis is the process of analyzing and extracting features such as the rhythmic patterns, beat intervals, and dynamic distribution of the background music. The analysis result is a set of background music rhythmic features obtained through rhythm analysis. Timestamps are information used to mark the specific location of each beat on the background music timeline.

[0143] As an optional implementation, the complete audio stream of the selected background music is acquired, and a globally unified rhythm analysis is performed on the complete audio stream to extract its overall rhythmic periodic features. Based on these overall rhythmic periodic features, all beats in the complete audio stream that conform to a periodic pattern are identified, and each identified beat is directly labeled with its corresponding timestamp. During the labeling process, the interval between timestamps is kept consistent with the global rhythmic periodic features. After labeling all beats with timestamps, they are summarized in chronological order to form a complete set of beat timestamps. This method has a simple operation process, does not require audio segmentation, can quickly complete rhythm analysis and timestamp labeling, has high processing efficiency, and meets the requirements of conventional audio-visual matching scenarios with high requirements for rhythm analysis efficiency.

[0144] As an alternative implementation, the complete audio stream of the selected background music is first divided into multiple continuous segments based on changes in audio features, and each segment undergoes independent rhythm analysis. The rhythmic cycle features and intensity distribution features specific to each segment are extracted. Based on these features, the beats within each segment are identified, and a timestamp is assigned to each beat within that segment, ensuring that the timestamp intervals precisely match the rhythmic cycle features of the corresponding segment. After completing the timestamp annotation for each segment, the timestamp sets of all segments are integrated. Simultaneously, the continuity of beat timestamps between adjacent segments is verified, and annotation deviations in connecting areas are corrected to form the final beat timestamp set. This method fully adapts to the rhythmic changes of different segments of the background music, achieving higher accuracy in beat timestamp annotation. It effectively avoids annotation errors in rhythm switching areas, achieving high-precision acquisition of background music beat timestamps and meeting the needs of complex audio-visual collaboration scenarios with high requirements for rhythm matching accuracy.

[0145] Step E12: Adjust the playback sequence of the hitting segment according to the deviation between the hitting time and the timestamp, so as to align the hitting time with the beat of the background music.

[0146] In this embodiment, the deviation is the time difference between the moment of impact and the corresponding beat timestamp. The playback sequence refers to the order and temporal position of the impact segments in the video.

[0147] As an optional implementation, all hit times and corresponding background music beat timestamps are extracted. The deviation between each hit time and its corresponding timestamp is calculated, and all deviations are aggregated to calculate the average deviation. Based on this average deviation, the playback sequence of all hit segments is uniformly adjusted, shifting the time position of each hit segment as a whole according to the average deviation. After the shift is complete, all hit segments are arranged in chronological order to ensure that most hit times correspond to the beat after adjustment. Finally, the adjusted hit segments are integrated to form a preliminary timing matching result. This method has a simple operation process, does not require adjusting individual segments one by one, can quickly complete the overall timing calibration, has high processing efficiency, and achieves the effect of quickly aligning the hit times with the beat, meeting the needs of general processing scenarios with high requirements for timing adjustment efficiency.

[0148] As an alternative implementation, each shot moment is first matched with a corresponding background music beat timestamp, and the deviation between each individual shot moment and its corresponding timestamp is calculated. Based on the magnitude and direction of each deviation, the playback sequence of the corresponding shot segment is adjusted accordingly. For segments with larger deviations, the timing offset is appropriately increased; for segments with smaller deviations, the timing position is fine-tuned. After adjusting a single segment, the smoothness of the playback transition between that segment and its preceding and following segments is verified to avoid visual gaps caused by timing adjustments. The timing adjustment and transition verification of all shot segments are completed sequentially, and finally, all segments are integrated to form the final timing matching result. This method can accurately adapt to the deviation differences between each shot moment and the beat, achieving higher matching accuracy, effectively ensuring audio-visual coordination, and achieving high-precision alignment of shot moments and beats, meeting the needs of refined processing scenarios with high requirements for audio-visual matching quality.

[0149] Step E13: The video to be processed, with the playback timing of the ball-hitting segment adjusted, is used as the beat video.

[0150] As an optional implementation, the video to be processed, after adjusting the playback timing of the hitting segments, is first acquired. The video content is then scanned segment by segment to verify the accuracy of the timing match between each hitting moment and the corresponding background music beat, confirming that the time deviation is within a reasonable range. Next, the smoothness of the transitions between adjacent hitting segments and the consistency of audio-visual synchronization are verified. For segments with insufficient matching accuracy, the timing position is fine-tuned; for segments with abrupt transitions, transition frames are added for optimization. After all verification and correction operations are completed, the complete attribute information and verification and correction records of the video are extracted. The corrected video file is marked as a beat video, and a detailed verification report is generated as an auxiliary file for subsequent processes. This method, through multi-dimensional verification and targeted correction, ensures the audio-visual matching accuracy and playback smoothness of the beat video, generating high-quality videos and achieving the effect of generating high-quality, highly adaptable beat videos, meeting the needs of refined editing scenarios with high video quality requirements.

[0151] For example, in a video editing scenario, based on the precise hit timestamp output by the preceding module, the original video stream undergoes secondary spatiotemporal processing to extract the most visually appealing motion segments. The temporal segmentation unit's function is to extract video segments containing the complete hitting motion based on the hit point T_hit. The processing logic involves setting a forward buffer time Delta_t_pre (e.g., 1.0 seconds, covering the backswing) and a backward buffer time Delta_t_post (e.g., 1.5 seconds, covering the follow-through and ball trajectory). For each valid hit moment T_i, the segment is extracted as: Segment_i = [T_i - Delta_t_pre, T_i + Delta_t_post]. Anti-collision logic: if the interval between two hit points is too short (e.g., continuous rapid rallies), causing the two segments to overlap, the system automatically merges them into a single long segment to maintain viewing continuity. The spatial domain reconstruction unit's function: For original wide-angle videos, it automatically crops close-up shots centered on the subject using computer vision technology, adapting to the vertical screen viewing experience on mobile devices. Technical implementation: Human image detection, using lightweight object detection models (such as YOLO-Nano or MediaPipe Pose) to detect the athlete's bounding box in each frame of Segment_i. Dynamic tracking: Calculating the trajectory of the bounding box's center point. Applying a Kalman filter to smooth the trajectory, avoiding camera shake. Intelligent cropping: Using the smoothed center point as a reference, cropping the region of interest (ROI) according to a preset ratio (e.g., 9:16), generating a close-up video stream focusing on the hitting action. The music beat extraction module's function: Analyzing user-selected or system-recommended background music (BGM), it extracts perceptually salient rhythm points (Onsets) and stable beat sequences from the background music signal through time-frequency analysis. This includes the following steps: ODF construction, transforming the complex music signal into a curve that reflects energy abrupt changes. Spectral Flux Calculation: Perform a Short-Time Fourier Transform (STFT) on the music signal. Calculate the difference in amplitude spectra between two adjacent frames and retain only the positively increasing portion (i.e., the portion where energy suddenly increases, usually corresponding to drum beats). ODF(n) = Σ(H(|X(n,k)|-|X(n-1,k)|)), where H(x) is the half-wave rectified function (i.e., x is taken when x>0, otherwise 0), n is the frame index, and k is the frequency index. Adaptive Whitening: To prevent a particular instrument from consistently being too loud and masking the drum beats, logarithmic compression or dynamic range compression is applied to the ODF curve to highlight transient changes. Peak-Picking-based transient localization. Smoothing and Thresholding: Smooth the original ODF curve to remove minor jitter.A dynamic threshold `Threshold_local` is set, which is equal to a multiple of the mean ODF value within the current window. Maximum search: Locate all time points that satisfy ODF(n) > ODF(n-1), ODF(n) > ODF(n+1), and ODF(n) > Threshold_local. These points are marked as candidate onsets. Tempo induction for periodicity-based beat tracking: Calculate the inter-onset interval histogram. The highest peak in the histogram corresponds to the most likely song tempo (BPM, Beats Per Minute). Dynamic programming optimization: To select the sequence from the candidate points that best matches the BPM pattern, define an objective function to find a path such that: the selected points have the largest possible energy (representing true drum beats), and the time interval between the selected points is as close as possible to the theoretical period corresponding to the BPM. The Viterbi Algorithm is used to backtrack and output the final aligned, strictly periodic beat timestamp sequence. The audiovisual synchronization and synthesis unit is responsible for aligning and fusing the precisely identified shot segments with the rhythmic audio to generate the final target video with a musical beat. Core alignment principle: The system forces the center moment of each selected video segment Segment_i (i.e., the moment of the shot T_hit, i) to the selected musical beat point b_k. Smooth speed adjustment: If the length of the video segment does not perfectly match the interval between two beats, the system fine-tunes the video playback speed (between 0.9x and 1.1x) to ensure that the action and music are perfectly synchronized while maintaining a smooth visual effect. Adaptive content quantity: In practice, the number of shot segments (N_clips) and the number of musical beats (N_beats) often do not match. Video Redundancy: Prioritize and select only the N_beats video clips with the highest quality score (i.e., the weighted result of transient peak significance and target frequency band energy percentage) for alignment, and remove the remaining redundant clips. Music Redundancy: Trim the background music (BGM) to only retain the length that can cover all N_clips clips. Alternatively, only select the N_clips of the strongest beats in the BGM for alignment, skipping the weak beats. Video Insufficiency: Loop / Fill the video, looping or mirroring the clip with the highest quality score to fill the remaining beat nodes. Alternatively, trim the lowest intensity part of the BGM, retaining only the most exciting music segments. Final Compositing: Mix the processed video track with the BGM track, add transition effects (such as flash white, vibration), and output the final target video file.Among them, the weighted result of transient peak significance and target frequency band energy ratio plays a global role in the audio effectiveness evaluation module, evaluating the audio quality of the entire video; this function can be reused to evaluate the audio quality of each shot segment, in which case the scope of application changes from the entire video to the shot segment, and the quality of the shot segments can be ranked.

[0152] By precisely aligning the beats and dynamically adapting the adjustments, the problems of inconsistent audio-visual rhythm and abrupt transitions in ball game video editing have been solved, thus improving the audiovisual synergy and viewing experience of the target video.

[0153] Based on any of the above embodiments, in Embodiment 7 of this application, after step E11, steps F11 to F14 are further included:

[0154] Step F11: Based on the transient peak significance and target frequency band energy ratio corresponding to the ball-hitting segment, and the preset weights of the transient peak significance and the target frequency band energy ratio, a weighted calculation result is obtained.

[0155] In this embodiment, transient peak significance refers to the degree of distinction between the instantaneous peak signal and the stationary signal in the impact segment. Target frequency band energy proportion refers to the proportion of signal energy within the core frequency response range to the total energy of the entire frequency band. Weighted calculation result refers to the comprehensive quality score calculated after assigning preset weights to the two indicators.

[0156] As an optional implementation, two indicators—transient peak significance and target frequency band energy percentage—are extracted for all shot segments. A preset uniform weighting ratio is retrieved, which applies fixed values ​​to both indicators for all shot segments. Then, the two indicators for each shot segment are weighted according to the uniform weighting, and the calculated values ​​are directly recorded as the weighted calculation results for the corresponding shot segment. The results of all shot segments are then aggregated to form a complete result set. This method eliminates the need for differentiated analysis of different shot segments, has a simple and unified operation process, and can quickly complete the weighted calculation of a batch of shot segments. It has high processing efficiency and achieves the effect of quickly obtaining the basic weighted calculation results of a batch of shot segments, meeting the needs of conventional audio quality evaluation scenarios with high computational efficiency requirements.

[0157] As an alternative implementation, the transient peak salience and target frequency band energy proportion data for all shot segments are first extracted. The audio features of each shot segment are analyzed separately, and differentiated preset weights are assigned to the two indicators for different shot segments based on feature differences. For segments with prominent transient peak features, the weight of transient peak salience is increased; for segments with prominent target frequency band energy proportion features, the weight of target frequency band energy proportion is increased. Then, each shot segment is sequentially weighted according to the assigned differentiated weights. After calculation, the matching degree between the result and the actual audio features of the corresponding shot segment is verified. Abnormal results with substandard matching are removed, and the calculation is recalculated. All verified results are then combined to form the final weighted calculation result set. This method fully considers the differences in audio features among different shot segments, and the weighted calculation results can more accurately reflect the actual audio quality of the segments, achieving high-precision acquisition of weighted calculation results for shot segments, meeting the requirements of refined processing scenarios with high accuracy requirements for audio quality evaluation.

[0158] Step F12: Select a preset number of the ball-hitting segments with the highest weighted calculation result values ​​from each of the ball-hitting segments as high-priority segments.

[0159] In this embodiment, the weighted calculation result is a comprehensive quantitative score obtained by combining the transient peak significance of the shot segment, the energy proportion of the target frequency band, and the corresponding preset weights. The preset quantity is a pre-set specific number of high-priority segments that need to be selected. High-priority segments are the set of segments with the highest weighted calculation result and the best overall quality selected from all shot segments.

[0160] As an optional implementation, all shot segments and their corresponding weighted calculation results are extracted to establish a one-to-one correspondence between shot segments and values. Then, all weighted calculation results are globally sorted, arranging the corresponding shot segments in descending order. After sorting, the first preset number of shot segments are directly extracted and marked as high-priority segments. The entire filtering process uses the weighted calculation results as the sole criterion, and finally, the marked segments are summarized to form a high-priority segment set. This method has a simple and intuitive operation flow, requires no additional segmentation or classification processing, can quickly complete the filtering of batch segments, has high processing efficiency, and meets the needs of batch processing scenarios with high filtering efficiency requirements.

[0161] Step F13: If the duration of the background music exceeds the duration of the high-priority segment, extract the strong beat timestamps from the background music that are of high intensity and have the same number as the high-priority segment.

[0162] In this embodiment, "high intensity" refers to the part of the background music with the highest beat energy, and "the number is the same as the high priority segments" means that the number of strong beat timestamps is the same as the number of high priority segments. A strong beat timestamp refers to the time position information corresponding to the peak energy beat in the background music.

[0163] As an optional implementation, the time interval distribution of all high-priority segments is first calculated to determine the time distribution requirements that the appropriate beat timestamp should meet, and the minimum background music duration required to cover the high-priority segments is calculated. If the actual length of the background music exceeds this, the background music is divided into several corresponding time periods according to the minimum duration. Within each time period, the timestamp of the strongest beat is extracted, ensuring that each time period has strong beats selected, and that the time distribution of the selected strong beats matches the time intervals of the high-priority segments. After selecting the number of strong beat timestamps that match the number of high-priority segments, their intensity is checked to ensure that they are all within a preset high value range. If they meet the standard, they are determined as beat timestamps. This method takes into account both the strength of the strong beats and the adaptability of the time distribution, resulting in a more harmonious matching effect and providing excellent support for accurate beat matching of high-priority segments.

[0164] As an alternative implementation, the intensity distribution of background music throughout the entire time period is analyzed, all segments with the lowest intensity are marked, and their total duration is calculated. Based on the difference between the number of high-priority segments and the number of beat nodes, the duration of background music to be trimmed is calculated. The segments with the lowest intensity are precisely extracted and discarded, retaining only the core, exciting segments with high intensity and distinct rhythms. This ensures that the number of beat nodes to be aligned in the trimmed background music matches the number of high-priority segments. This method avoids repetitive playback of hitting segments, maintains the freshness and diversity of visual materials, and meets the needs of scenarios with high requirements for visual freshness.

[0165] As an alternative implementation, in the music beat extraction module, after extracting the strong beat timestamps of the background music, the distribution of the hitting time intervals of high-priority segments is first analyzed. When filtering strong beat timestamps, combinations of strong beats with an interval matching the hitting time interval at a preset ratio are prioritized. If the matching degree is insufficient, the selection range of some strong beats is fine-tuned, replacing them with adjacent second-strong beats, until the matching degree between the strong beat interval and the hitting interval meets the standard. This method improves the adaptability of strong beats to the hitting rhythm and reduces the magnitude of subsequent speed adjustments.

[0166] Step F14: Align the hitting time of the high-priority segment with the strong shot timestamp one by one to obtain the beat video.

[0167] As an optional implementation, high-priority segments are first sorted by quality score from highest to lowest, while the strong and weak beat attributes in the beat timestamps are marked. The hitting time of high-scoring segments is first matched with the strong beat timestamp, and then the remaining segments are matched with the weak beat timestamps, calculating the difference between the hitting time and the beat timestamp during matching. If the difference exceeds a preset range, the segment duration is fine-tuned to reduce the difference without disrupting the smoothness of the movement. After each matching is completed, the fit between the rhythm intervals of adjacent segments after alignment and the beat intervals of the background music is checked, and abnormal deviations are corrected to obtain the beat video. This method balances segment quality and beat attributes, achieving high alignment accuracy and good audio-visual coordination.

[0168] For example, in a video editing scenario, a "one-click" service for generating ball-sport sound effects is provided for ball sports scenes. First, the scene-adaptive audio enhancement module is activated. Based on user selection or image recognition, the specific sports scene (e.g., clay court tennis, indoor table tennis) is determined, and the corresponding preset core frequency response range is automatically loaded. While preserving the transient energy of the ball impact within this range, stable noise outside the range (e.g., wind noise) and irrelevant impact sounds (e.g., low-frequency footsteps) are actively suppressed, achieving targeted audio enhancement. Next, the "audio quality gating mechanism" of the audio effectiveness evaluation module performs multi-dimensional scoring on the enhanced audio, considering signal-to-noise ratio, transient saliency, and spectral purity. Subsequent processes are only initiated when the score exceeds a preset threshold or the proportion of effective segments meets the standard; otherwise, the user feedback mechanism is directly triggered. The process then proceeds to the ball-hitting event recognition and anti-interference module, employing a multi-level interference elimination mechanism based on "joint time-frequency features." First, candidate hitting points are located by peak finding in the time domain. Then, only small segments at the candidate point locations are extracted for on-demand Fast Fourier Transform (FFT) analysis. The authenticity of candidate events is verified through cascaded filtering of three-dimensional features: frequency domain dominant frequency, time domain morphology, and frequency domain energy distribution, reducing false identifications. Next, the ball-hitting segment editing module selects the high-priority hitting segments with the highest quality scores. Finally, the music beat extraction module analyzes the beat information of the selected background music. Finally, in the video and music synthesis module, the playback speed of the ball-hitting segment is dynamically fine-tuned (0.9x-1.1x) by combining a variable-speed frame interpolation mechanism, forcibly aligning the ball-hitting moment with the strong beat of the background music. At the same time, based on the matching of the number of high-priority segments and the music beat nodes, the segments are adapted by looping / mirroring or cropping the low-intensity parts of the background music, thus completing automated audiovisual synthesis and efficiently generating high-quality beat-matching videos. This not only solves the problem of time-consuming manual editing, but also reduces the probability of false detection through scene-based optimization, achieving a balance between low power consumption and high accuracy on mobile devices.

[0169] By using weighted scoring to select high-quality segments and precise extraction of strong beats, the problem of uneven segment quality and excessively long music causing beat alignment confusion in ball game video editing has been solved, improving the accuracy of beat alignment and the overall audiovisual quality of the final product.

[0170] Based on any of the above embodiments, in Embodiment Eight of this application, referring to Figure 3 , Figure 3 This is a flowchart illustrating the eighth embodiment of the video editing method based on ball hit detection of this application. Following step S60, steps G11-G13 are also included:

[0171] Step G11: Push the target video to the interactive interface and provide feedback on beat alignment accuracy, playback speed adaptation, and audio clarity to obtain user input for adjustment.

[0172] In this embodiment, the user interface is a visual interface for users to view videos and submit operation requests. Beat alignment accuracy refers to the accuracy of the match between the shot timing and the background music beat. Playback speed adaptation refers to the degree to which the playback speed of the shot segment matches the background music beat. Audio clarity refers to the clarity and distinguishability of sounds such as the shot sound and background music in the video. The feedback entry for adjustable dimensions is a dedicated entry point for receiving user requests to modify the above three dimensions. User-inputted adjustment requests are the specific modification requirements proposed by the user for the three adjustable dimensions.

[0173] As an optional implementation, the target video is pushed to the user interface in a drag-and-drop format. Simultaneously, three adjustable feedback entry points are displayed in the sidebar. Each entry point provides a parameter adjustment slider, a preset optimization template, and a custom input box. Users can fine-tune parameters using the slider, select a template, or enter text to describe their specific needs. The system generates a real-time preview of the adjustments for user confirmation, and the final adjustment request is recorded after user confirmation. This method offers highly refined and personalized adjustments, accurately capturing deep-seated user needs and improving target video satisfaction.

[0174] Step G12: Retrieve the preset technical parameter configuration mapping relationship corresponding to the adjustment requirement, and determine the technical parameter modification scheme that matches the adjustment requirement.

[0175] In this embodiment, the preset technical parameter configuration mapping relationship is a pre-established set of associated correspondences between adjustment requirements and corresponding technical parameter modification rules. The technical parameter modification scheme is a specific operational rule for adjusting technical parameters, generated based on the mapping relationship matching.

[0176] As an optional implementation method, the adjustment requirements are first broken down into core requirements and multiple sub-requirements. A hierarchical preset technical parameter configuration mapping relationship is retrieved, which is divided into a core layer and a sub-layer. First, the basic technical parameter modification schemes corresponding to the core requirements are matched, and then the parameter fine-tuning rules corresponding to each sub-requirement are matched one by one. The fine-tuning rules are superimposed onto the basic scheme, and the compatibility between the parameters after superposition is verified. Conflicting fine-tuning rules are eliminated, and suitable alternative rules are added. The final technical parameter modification scheme is then formed, and the complete adjustment logic and parameter details of the scheme are recorded. This method, employing a hierarchical matching and compatibility verification mechanism, can accurately adapt to complex adjustment requirements containing multi-dimensional demands. The scheme has higher fit and feasibility, achieving the effect of high-precision matching of technical parameter modification schemes corresponding to complex adjustment requirements, and meeting the processing scenarios of complex and diverse optimization needs.

[0177] Step G13: After optimizing the corresponding parameters according to the technical parameters, the target video generation step is re-executed to obtain the optimized target video and push it to the interactive interface.

[0178] In this embodiment, the optimized parameters are the target video generation parameters that have been adjusted and updated according to the technical parameter modification scheme. The target video generation process is a complete video production workflow, from audio enhancement and hit moment recognition to beat alignment and segment speed adjustment. The optimized target video is the final video product regenerated after parameter adjustment, and its effect meets the modification requirements.

[0179] As an optional implementation, the technical parameter modification scheme is retrieved to determine the types and adjustment ranges of parameters requiring optimization. The corresponding parameters in the target video generation process are then updated in batches. After the update, the complete target video generation process is directly initiated, sequentially executing all stages from the beginning, including audio frequency domain enhancement, audio feature extraction and quality verification, hit timing recognition, beat alignment, and hit segment speed adjustment. There is no need to break down or filter the process. Once the entire process is completed, the optimized target video is generated. The video file and parameter modification record are then packaged and pushed to the interactive interface. This method has a simple and direct operation logic, requires no analysis of the parameter influence range, and can quickly complete the entire process restart and video generation. It has high processing efficiency and achieves the effect of quickly generating and pushing optimized videos, meeting the needs of common optimization scenarios involving global parameter adjustments.

[0180] As an alternative implementation, the technical parameter modification scheme is first broken down, and the target video generation stage corresponding to each parameter to be optimized is analyzed to determine the scope of influence of the parameter adjustment. Only the parameters corresponding to the affected generation stages are updated in a targeted manner. After the update, there is no need to start the complete generation process; only the stages affected by the parameter adjustment are re-executed. The unaffected stages directly use the previously generated intermediate results. After the affected stages are completed, the newly generated intermediate results are integrated with the original results to generate the optimized target video. At the same time, a parameter impact analysis report is generated and pushed to the interactive interface for users to view. This method significantly reduces redundant calculations, lowers resource consumption, and shortens generation time by selectively executing the affected stages. It also offers higher accuracy and meets the needs of fine-grained optimization scenarios requiring local parameter fine-tuning.

[0181] For example, in a video editing scenario, a target video (25 seconds long) of a tennis scene is pushed to the user interface. The interface has three adjustable feedback entry points at the bottom: beat alignment accuracy, playback speed adaptation, and audio clarity (each entry point includes a parameter slider and a text input box). The user selects "Needs fine-tuning" for beat alignment accuracy via the slider and inputs their adjustment needs: "The timing of the shot deviates slightly from the strong beat of the music, the playback speed is slightly fast, and the shot sound is not clear enough." The system extracts the core dimensions as beat alignment accuracy, playback speed adaptation, and audio clarity. Based on these core dimensions, a preset technical parameter configuration mapping relationship library is retrieved. Beat alignment accuracy is matched with the "deviation correction amplitude" parameter (optimization target threshold ≤ 0.1 seconds), playback speed adaptation is matched with the "speed adjustment coefficient" parameter (optimization target threshold 0.95x-1.05x), and audio clarity is matched with the "core frequency band gain" parameter (core frequency response range 2k-5kHz, optimization target threshold ≥ 10dB). Re-execute the corresponding sub-steps according to the parameters and thresholds: correct the original beat deviation of 0.2 seconds to 0.08 seconds, adjust the playback speed from 1.1x to 1.0x, increase the core frequency band gain to 12dB, generate the optimized target video and push it to the user interface. After the user views it, they report "the beat alignment is accurate, the speed is appropriate, and the audio is clear", confirming that no further adjustments are needed.

[0182] By employing an interactive and precise optimization mechanism, the problem of mismatch between automatically synthesized videos and users' personalized needs has been resolved, thereby improving user satisfaction and adaptation flexibility of the target videos.

[0183] Based on any of the above embodiments, in Embodiment Nine of this application, referring to Figure 4 , Figure 4 This is the core block diagram of this application.

[0184] As an optional implementation, after inputting the video to be processed, the system first enters the scene-adaptive audio enhancement module. This module automatically loads the corresponding core frequency response range based on the sports scene (e.g., clay court tennis, indoor table tennis). While preserving the transient energy of the hit within this range, it actively suppresses stable noise and irrelevant impact sounds outside the range to enhance the audio quality. Next, the system enters the audio validity evaluation module. Through an "audio quality gating mechanism," the enhanced audio is scored in multiple dimensions, including signal-to-noise ratio, transient saliency, and spectral purity. Only when the score exceeds a preset threshold or the proportion of valid segments meets the standard does the system proceed to the next step; otherwise, the user feedback mechanism is triggered directly. Afterward, the system enters the hit event recognition and anti-interference module. This module employs a multi-level interference removal mechanism based on "time-frequency joint features." First, candidate hit points are located by time-domain peak finding. Then, only a small segment at the candidate point location is extracted for Fast Fourier Transform analysis. The authenticity of the candidate event is verified sequentially through three dimensions: frequency domain dominant frequency, time domain morphology, and frequency domain energy distribution, thus achieving accurate identification of the hit moment. The process then moves to the shot clip editing module. Based on the weighted calculation results of the transient peak significance and the target frequency band energy proportion of each shot clip, a preset number of clips with the highest quality scores are selected as high-priority clips. Simultaneously, background music is input, and the music beat extraction module analyzes its beat timestamps and adjacent interval information. If the background music length exceeds the requirement to cover the high-priority clips, the strongest beat timestamps with the highest intensity and the same number as the high-priority clips are extracted as the beat timestamps. If the number of high-priority clips is less than the number of beat nodes that the background music needs to align with, the highest quality high-priority clips are played in a loop / mirrored manner, or the lowest intensity part of the background music is trimmed for adaptation. The shot timestamps of each high-priority clip are then matched one by one with the beat timestamps to obtain the beat-aligned shot clips. A variable-speed frame interpolation mechanism is used to dynamically adjust the playback speed of the clips to maintain the visual smoothness of the shot action and achieve audiovisual synergy. Finally, the video and music synthesis module is entered, where the adjusted shot clip is mixed with the background music to obtain the target video. This target video is then pushed to the user interface, providing feedback on adjustable dimensions such as beat alignment accuracy, playback speed adaptation, and audio clarity. After obtaining the user's adjustment requirements, the module retrieves the corresponding preset technical parameter configuration mapping relationship based on the core dimensions of the requirements, matches and determines the specific technical parameters that need to be modified and their optimization target thresholds, re-executes the corresponding video processing sub-steps to generate the optimized target video, and pushes it to the user interface again until the user confirms that no further adjustments are needed, and finally outputs the finished product.

[0185] This application provides a video editing device based on ball hit detection. The video editing device based on ball hit detection includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the video editing method based on ball hit detection in the first embodiment described above.

[0186] The following is for reference. Figure 5 This document illustrates a structural schematic diagram of a video editing device based on shot detection suitable for implementing embodiments of this application. The video editing device based on shot detection in these embodiments may include, but is not limited to, mobile terminals such as mobile phones, laptops, intelligent shot capture cameras, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), and integrated editing devices, as well as fixed terminals such as end-to-end automatic shot video processing platforms and desktop computers. Figure 5 The video editing device based on ball hit detection shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0187] like Figure 5As shown, a video editing device based on ball-hit detection may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The random access memory 1004 also stores various programs and data required for the operation of the video editing device based on ball-hit detection. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the ball-hit-detection-based video editing device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows a ball-hit-detection-based video editing device with various systems, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems can be implemented alternatively.

[0188] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0189] The video editing device based on ball-hit detection provided in this application, employing the video editing method based on ball-hit detection in the above embodiments, can solve the technical problem of poor video editing effect. Compared with the prior art, the beneficial effects of the video editing device based on ball-hit detection provided in this application are the same as the beneficial effects of the video editing method based on ball-hit detection provided in the above embodiments, and other technical features in the video editing device based on ball-hit detection are the same as the features disclosed in the method of the previous embodiment, and will not be repeated here.

[0190] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0191] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0192] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the video editing method based on ball-hitting detection in the above embodiments.

[0193] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0194] The aforementioned computer-readable storage medium may be included in a video editing device based on ball-hitting detection; or it may exist independently and not assembled into a video editing device based on ball-hitting detection.

[0195] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a video editing device based on ball-hit detection, cause the video editing device based on ball-hit detection to: respond to an editing trigger command for the video to be processed, determine a core frequency response range based on the ball sports scene recognition result of the video to be processed; perform frequency domain enhancement processing on the original audio of the video to be processed based on the core frequency response range to obtain audio data; extract audio features of the audio data in each evaluation dimension, and determine whether the audio features all meet the corresponding audio quality standards; if so, determine the ball-hit moment in the video to be processed based on ball-hit recognition rules; align the ball-hit moment with the beat of the background music to obtain a beat video; and dynamically adjust the playback speed of the ball-hit segments in the beat video according to the interval between the beats to obtain the target video.

[0196] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0197] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0198] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0199] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described video editing method based on ball impact detection, thereby solving the technical problem of poor video editing results. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the video editing method based on ball impact detection provided in the above embodiments, and will not be repeated here.

[0200] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A video editing method based on ball hit detection, characterized in that, The method includes: In response to the editing trigger command of the video to be processed, the core frequency response range is determined based on the ball sports scene recognition result of the video to be processed; Based on the core frequency response range, the original audio of the video to be processed is subjected to frequency domain enhancement processing to obtain audio data; Extract the audio features of the audio data in each evaluation dimension, and determine whether the audio features all meet the corresponding audio quality standards; If so, the candidate hitting point is located based on the transient energy peak detection result of the audio data; By matching target frequency features, verifying the centroid of the spectrum, and analyzing the temporal morphology of the signal, interference signals in the candidate hitting points are eliminated to determine the hitting time. Align the hitting moment with the beat of the background music to obtain a beat video; The playback speed of the ball-hitting segments in the beat video is dynamically adjusted according to the interval between the beats to obtain the target video.

2. The video editing method based on ball hit detection as described in claim 1, characterized in that, The step of determining the core frequency response range based on the ball sports scene recognition result of the video to be processed in response to the editing trigger command of the video to be processed includes: In response to the editing trigger command of the video to be processed, key frames that are representative of the scene are extracted from the video to be processed uploaded by the user to form basic data; The basic data is analyzed, and the site features and sports equipment features in the keyframes are identified based on the analysis results. Based on the characteristics of the venue and the characteristics of the sports equipment, the ball game scenario is determined; The mapping database between the scene and the core frequency response range is invoked, and the core frequency response range is determined based on the mapping relationship stored in the mapping database and the ball game scene.

3. The video editing method based on ball hit detection as described in claim 1, characterized in that, The step of performing frequency domain enhancement processing on the original audio of the video to be processed based on the core frequency response range to obtain audio data includes: Based on the core frequency response range, the original audio is extracted from the video to be processed; Obtain pre-set spectrum energy weighting parameters, wherein the spectrum energy weighting parameters include a set gain weight corresponding to the signal within the core frequency response range, and a set attenuation weight corresponding to noise and redundant signals outside the core frequency response range; According to the spectral energy weighting parameters, the energy values ​​of each frequency band of the original audio are adjusted in the time-frequency domain to obtain the audio data.

4. The video editing method based on ball hit detection as described in claim 1, characterized in that, The audio features include global effective signal-to-noise ratio, transient peak significance, and target frequency band energy percentage.

5. The video editing method based on ball hit detection as described in claim 1, characterized in that, After the steps of extracting audio features from the audio data across various evaluation dimensions and determining whether all audio features meet the corresponding audio quality standards, the video editing method based on ball-hitting detection further includes: If not, output the dimensions where the quality is substandard and provide optimization suggestions; The interactive options are output according to the aforementioned quality non-compliance dimensions and the aforementioned optimization suggestions, wherein the interactive options include at least one of re-uploading the video, adjusting the frequency domain enhancement parameters, and supplementing scene information; Receive the user's selected interaction choice and execute the processing process corresponding to the selected interaction choice.

6. The video editing method based on ball hit detection as described in claim 1, characterized in that, The step of aligning the moment of impact with the beat of the background music to obtain a beat video includes: The selected background music is subjected to rhythm analysis, and the timestamp corresponding to the beat is determined based on the analysis results; Based on the deviation between the hitting time and the timestamp, the playback sequence of the hitting segment is adjusted to align the hitting time with the beat of the background music; The video to be processed, with the playback timing of the ball-hitting segment adjusted, is used as the beat video.

7. The video editing method based on ball hit detection as described in claim 6, characterized in that, After the step of performing rhythm analysis on the selected background music and determining the timestamp corresponding to the beat based on the analysis result, the video editing method based on ball hit detection further includes: The weighted calculation result is obtained based on the transient peak significance and target frequency band energy ratio corresponding to the ball-hitting segment, as well as the preset weights of the transient peak significance and the target frequency band energy ratio; From each of the aforementioned shot segments, a predetermined number of shot segments with the highest weighted calculation result value are selected as high-priority segments; If the duration of the background music exceeds the duration of the high-priority segment, then extract the strong beat timestamps from the background music that are of high intensity and have the same number as the high-priority segment. The high-priority segment's hit time is aligned with the strong shot timestamp one by one to obtain the beat video.

8. The video editing method based on ball hit detection as described in claim 1, characterized in that, After the step of dynamically adjusting the playback speed of the hitting segments in the beat video according to the interval between the beats to obtain the target video, the video editing method based on hit detection further includes: The target video is pushed to the interactive interface, and feedback entry points for beat alignment accuracy, playback speed adaptation, and audio clarity are provided to obtain user input adjustment requirements. Retrieve the preset technical parameter configuration mapping relationship corresponding to the adjustment requirement, and determine the technical parameter modification scheme that matches the adjustment requirement; After optimizing the corresponding parameters according to the technical parameter modification scheme, the target video generation step is re-executed to obtain the optimized target video and push it to the interactive interface.

9. A video editing device based on ball strike detection, characterized in that, The video editing device based on ball hit detection includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the video editing method based on ball hit detection as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the video editing method based on ball hit detection as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Method and system for editing billiard video based on artificial intelligence

    CN115499706A

  • Ball game analysis method and device, electronic equipment and storage medium

    CN120747811A