Video editing and artistic creation system based on AI intelligence

By constructing a dynamic correlation analysis of the action feature library and the audience feedback audio waveform data, the problems of low efficiency and insufficient artistic expression in opera video editing are solved, and accurate recognition and emotional matching of opera movements are achieved, and high-quality editing content that meets the short video platform is generated.

CN120416587AInactive Publication Date: 2025-08-01PACO VIDEO TECH (HANGZHOU) CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510538845.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing AI-based video editing system is inefficient and subjective in opera performances, and cannot accurately identify programmatic actions and emotional resonance clips. The editing content is misaligned with the audience's expectations, and lacks the support of the knowledge base in the field of opera.

Method used

A motion feature library is constructed, and dynamic correlation analysis is performed based on the audience's feedback audio waveform data to generate a short video sequence that conforms to the preset editing rules. Through multi-dimensional feature extraction and dynamic correlation analysis, accurate recognition and emotional matching of opera movements are achieved.

Benefits of technology

It improves the efficiency and artistic expression of opera performance video editing, ensures that the editing content resonates with the audience's emotional emotions, meets the requirements of the short video platform, and preserves the integrity and artistic value of opera movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416587A_ABST
    Figure CN120416587A_ABST
Patent Text Reader

Abstract

The invention discloses a video editing and artistic creation system based on AI intelligence, which relates to the technical field of video editing and comprises a data acquisition module, a feature analysis module, a feature library construction module, a rule matching module and a split mirror generation module, the data acquisition module is used for acquiring traditional opera stage video data and corresponding audience feedback audio waveform data, and the feature analysis module is used for extracting spatial-temporal feature vectors of stylized actions from the traditional opera stage video data. The feature library construction module is used for constructing an action feature library containing action type codes and music beat deviation level codes according to the spatio-temporal feature vectors; the method has the beneficial effects that dynamic association analysis is performed by constructing the action feature library and combining audience feedback audio data, the short video sequence conforming to the art rule is automatically generated, and the method has the advantage of improving the Chinese opera performance video editing efficiency and the art expressive force.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video editing, and particularly to a video editing and art creation system based on AI intelligence. Background Art

[0002] With the rapid development of artificial intelligence technology, the fields of video editing and art creation are gradually evolving towards automation and intelligence. In the prior art, AI-based video editing systems mostly focus on general scenarios, and realize shot switching and special effect generation through technologies such as object detection and action recognition. In Chinese opera performances, traditional editing relies on manual experience to screen highlight segments, which has problems of low efficiency and strong subjectivity. In addition, Chinese opera performances include stylized actions and specific art rules. Due to the lack of support from a domain knowledge base, existing general editing systems are difficult to accurately identify action and emotional resonance segments, resulting in a mismatch between the edited content and the audience's expectations. Summary of the Invention

[0003] In view of the above-mentioned prior art situation, the present application is proposed. Embodiments of the present application provide a video editing and art creation system based on AI intelligence, which has the advantages of improving the editing efficiency and artistic expressiveness of Chinese opera performance videos.

[0004] According to one aspect of the present application, there is provided a video editing and art creation system based on AI intelligence, including: a data acquisition module for acquiring Chinese opera stage video data and corresponding audience feedback audio waveform data; a feature analysis module for extracting spatio-temporal feature vectors of stylized actions from the Chinese opera stage video data, the spatio-temporal feature vectors including an amplitude dimension, a duration dimension, and a beat deviation dimension; a feature library construction module for constructing an action feature library containing action type codes and music beat deviation level codes according to the spatio-temporal feature vectors; a rule matching module for parsing a preset editing rule library of a short video platform, the editing rule library containing shot switching interval parameters and special effect triggering conditions; a storyboard generation module for calling the action type codes and beat deviation level codes in the action feature library, and performing dynamic correlation analysis in combination with the peak interval of the audience feedback audio waveform data to generate a short video sequence that conforms to the preset editing rule library.

[0005] As a preferred solution of a video editing and artistic creation system based on AI intelligence in this application, wherein the processing of the audience feedback audio waveform data includes: collecting the applause waveform in the audience through a directional microphone array to generate an emotion intensity model with timestamps; the calculation formula of the emotion intensity model is: emotion intensity value = waveform amplitude peak density × duration coefficient; wherein, the waveform amplitude peak density represents the number of peaks exceeding a preset amplitude threshold per unit time, when the applause duration is less than or equal to 5 seconds, the duration coefficient takes a value of 0.5, when the applause duration is greater than 5 seconds and less than or equal to 20 seconds, the duration coefficient = 1.0 + 0.05 * (applause duration - 5), when the applause duration is greater than 20 seconds, the duration coefficient takes a value of 2, and the peak interval is defined as the period when the emotion intensity value continuously exceeds the first preset threshold.

[0006] As a preferred solution of a video editing and artistic creation system based on AI intelligence in this application, wherein the data structure of the action feature library is a triple set of {action type encoding, amplitude level encoding, beat deviation level encoding}, and the construction of the action feature library includes: obtaining the actor's bone movement trajectory through bone key point capture technology, and extracting the trajectory spatio-temporal features of the bone movement trajectory, the trajectory spatio-temporal features including the curvature and acceleration parameters of the movement trajectory; matching the trajectory spatio-temporal features of the actor's bone movement trajectory with a preset stylized action classification rule library to generate an action type encoding including a stylized action classification encoding, wherein the preset stylized action classification rule library stores the benchmark stylized action data of each opera genre, and the benchmark stylized action data includes the spatio-temporal features of the standard bone movement trajectory and its corresponding classification encoding and action type encoding; calculating the amplitude dimension of the stylized action as the maximum Euclidean distance of hand movement in a three-dimensional coordinate system, and comparing the Euclidean distance with a preset amplitude level threshold to generate an action amplitude level encoding; calculating the beat deviation dimension as the absolute value of the difference between the actual action time code and the accompaniment theoretical beat time code, and mapping the absolute value of the difference to a preset deviation interval to generate a beat deviation level encoding.

[0007] As a preferred solution of a video editing and artistic creation system based on AI intelligence in this application, wherein the matching of the preset stylized action classification rule library includes: calculating the similarity between the trajectory spatio-temporal features of the actor's bone movement trajectory and the trajectory spatio-temporal features of the standard bone movement trajectory in the stylized action classification rule library; when the similarity exceeds the first similarity threshold, it is determined that the matching is successful.

[0008] As a preferred solution of a video editing and artistic creation system based on AI intelligence in this application, it further includes a real-time data enhancement module, which is configured to continuously obtain the latest opera performance video data stream and extract the spatio-temporal features of the real-time skeletal movement trajectory; perform a difference analysis on the spatio-temporal features of the real-time skeletal movement trajectory and the standard features in the stylized action classification rule library to generate a feature drift coefficient; when the feature drift coefficient exceeds a preset drift threshold, trigger an incremental update operation of the feature library construction module: write the real-time spatio-temporal features that meet the second similarity threshold into the stylized action classification rule library and generate corresponding new action type codes; recalculate the beat deviation level codes in the action feature library based on the updated rule library.

[0009] As a preferred solution of a video editing and artistic creation system based on AI intelligence in this application, the parsing of the preset editing rule library includes: obtaining the shot transition interval parameters of popular opera videos from a short video platform; establishing a mapping rule between the special effect trigger condition and the action amplitude, and activating the particle special effect generation instruction when it is detected that the amplitude dimension exceeds the second preset threshold.

[0010] As a preferred solution of a video editing and artistic creation system based on AI intelligence in this application, the dynamic correlation analysis by combining the peak interval of the audience feedback audio waveform data to generate a short video sequence that meets the preset editing rule library includes: calculating the overlap degree between the time code of the peak interval and the duration dimension of the spatio-temporal feature vector; when the overlap degree exceeds the preset overlap degree threshold, intercepting the start and end time codes of the corresponding video segment; and performing shot recombination on the intercepted segment according to the shot transition interval parameter.

[0011] As a preferred solution of a video editing and artistic creation system based on AI intelligence in this application, before intercepting the start and end time codes of the corresponding video segment, it further includes: constructing an opera action knowledge graph, the nodes of which include the association relationship between the action type and the costume color feature, and the edge relationship includes the spatio-temporal correspondence between the climax action segment and the applause peak interval; when it is detected that the matching degree between the current frame color feature and the preset color template of the climax action in the opera action knowledge graph meets the preset matching threshold and the amplitude dimension is lower than the second preset threshold, activating the particle special effect generation instruction to generate corresponding particle special effects in the overlapping segment between the climax action segment and the applause peak interval.

[0012] As a preferred solution of a video editing and art creation system based on AI intelligence in this application, wherein the storyboard recombination includes: performing duration compression processing on the intercepted video segment so that the total duration matches the target duration in the preset editing rule library; the duration compression processing adopts a non-uniform time scaling algorithm, maintaining the original playback rate in the segment that preserves the integrity of the stylized actions, and adopting accelerated playback processing in the transition segment.

[0013] As a preferred solution of a video editing and art creation system based on AI intelligence in this application, wherein the execution conditions of the non-uniform time scaling algorithm include: when it is detected that the action amplitude dimension of the current video segment is lower than the third preset threshold and the beat deviation dimension is less than the preset number of digits, starting the accelerated playback processing; the rate of the accelerated playback processing is inversely proportional to the emotional intensity value.

[0014] Compared with the prior art, by using a video editing and art creation system based on AI intelligence according to an embodiment of this application, it is possible to perform dynamic correlation analysis by constructing an action feature library and combining the audience feedback audio data, and automatically generate a short video sequence that conforms to the art rules, which has the advantages of improving the editing efficiency and artistic expressiveness of the opera performance video. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] By describing the embodiments of this application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of this application will become more obvious. The drawings are used to provide a further understanding of the embodiments of this application, and constitute a part of the specification, and are used to explain this application together with the embodiments of this application, and do not constitute a limitation to this application. In the drawings, the same reference numerals generally represent the same components or steps.

[0016] Figure 1 It is a block diagram of a video editing and art creation system based on AI intelligence of the present invention.

[0017] Figure 2 It is a working flowchart of the real-time data enhancement module of a video editing and art creation system based on AI intelligence of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Next, exemplary embodiments according to this application will be described in detail with reference to the drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited by the exemplary embodiments described herein.

[0019] APPLICATION OVERVIEW

[0020] In the traditional existing video editing system for Chinese opera performances, general AI technology relies on object detection and action recognition algorithms for shot transition decisions, but lacks the ability to perform multi-dimensional modeling of the spatio-temporal characteristics of stylized actions. Such systems cannot analyze the coupling relationship among the amplitude, duration, and beat deviation of Chinese opera actions, resulting in the inability to accurately identify action segments with artistic value. At the same time, the existing data processing flow does not establish a dynamic association mechanism between the audience's emotional feedback waveform data and the video content, causing the mismatch between the editing decision and the audience's emotional resonance point.

[0021] For example, in the processing scenario of the martial arts segment of the Peking Opera "Crossroads at Sancha", the traditional system extracts the trajectory of the actor's skeletal key points through the OpenPose algorithm, but only calculates the two-dimensional plane displacement parameters and does not construct a three-dimensional spatio-temporal feature vector including the amplitude level and beat deviation. When processing the cloud hand movement, the system mistakenly identifies the maximum value of the arm swing amplitude as the key frame, ignoring the matching requirement of the action duration and the rhythm of the gongs and drums. At the data flow level, there is no synchronous analysis channel established between the applause waveform data in the audience seats and the video time code, resulting in a large time delay deviation between the peak interval of the applause and the climax segment of the stylized action.

[0022] If the above problems are not solved, the video generation system will not be able to meet the requirements of short video platforms for accurate scene segmentation of Chinese opera content, resulting in the loss of the integrity of the genre-specific actions in the edited output segments. The decoupled state of the audience's emotional feedback data and the video content will reduce the user interaction rate, causing the probability of popular recommendation of the generated video to drop by more than 23%.

[0023] In addition, the lack of the beat deviation dimension will lead to incorrect trigger timing of special effects, causing damage to the coordination between the visual presentation and the music rhythm, and ultimately affecting the promotability of the technical solution in the vertical art field.

[0024] When facing the above problems, this application first considers that the traditional video editing system fails to establish a collaborative analysis mechanism for the multi-dimensional characteristics of Chinese opera stylized actions and the audience's emotional feedback.

[0025] Through analysis, it is found that simply relying on the two-dimensional parameters of the skeletal trajectory cannot accurately reflect the artistic characteristics of Chinese opera actions, and the asynchronous processing of the applause waveform data and the video time code leads to editing decision deviations.

[0026] In response to this, the concept of this application is to try to introduce a three-dimensional spatio-temporal feature vector, taking amplitude, duration, and beat deviation as the core dimensions, and at the same time constructing an emotional intensity model to achieve the dynamic mapping between the audience feedback data and the video content.

[0027] During the exploration process, schemes such as solely relying on action recognition algorithms and static rule library matching were compared, and it was found that multi-dimensional feature fusion combined with dynamic correlation analysis can effectively solve the problems of missing action integrity and emotional misalignment. Finally, it was determined to achieve accurate shot segmentation by constructing a composite data architecture of an action feature library and a clip rule library.

[0028] Exemplary system

[0029] A video editing and artistic creation system based on AI intelligence according to an embodiment of the present application includes a data acquisition module, a feature analysis module, a feature library construction module, a rule matching module, and a shot segmentation generation module.

[0030] The working process and principle of the present application are as follows: The data acquisition module acquires opera stage video data and audience feedback audio waveform data. The feature analysis module extracts spatio-temporal feature vectors of stylized actions from the opera stage video data, including amplitude dimension, duration dimension, and beat deviation dimension. The feature library construction module constructs an action feature library containing action type codes and music beat deviation level codes according to the spatio-temporal feature vectors. The rule matching module parses the preset clip rule library of the short video platform, including shot transition interval parameters and special effect trigger conditions. The shot segmentation generation module calls the action type codes and beat deviation level codes in the action feature library, and performs dynamic correlation analysis in combination with the peak interval of the audience feedback audio waveform data to generate a short video sequence that conforms to the preset clip rule library.

[0031] Each module works together to achieve the intelligent editing and creation of opera videos. The data acquisition module provides basic data for subsequent analysis. The feature analysis module captures the artistic characteristics of opera actions through multi-dimensional feature extraction. The feature library construction module encodes the extracted features for subsequent matching. The rule matching module introduces the clip rules of the short video platform to ensure that the generated content meets the platform requirements. The shot segmentation generation module comprehensively utilizes the outputs of the foregoing modules to achieve dynamic correlation between action features and audience feedback, and generates a short video sequence that conforms to the rules.

[0032] The reasons for selecting the key technical features are as follows: The multi-dimensional design of the spatio-temporal feature vectors can comprehensively capture the artistic characteristics of opera actions; the construction of the action feature library realizes the encoded storage of features; the introduction of the preset clip rule library ensures the platform adaptability of the generated content; the dynamic correlation analysis mechanism realizes the accurate matching of video content and audience feedback. The selection of these features significantly improves the recognition accuracy of the system for opera actions and the accuracy of clip decisions.

[0033] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0034] The data acquisition module acquires opera stage performance videos through a high-definition camera, and simultaneously uses a directional microphone array to collect the applause and cheering sounds of the audience seats.

[0035] The feature analysis module uses computer vision algorithms to extract the actor's skeletal key points from the video, calculates the maximum Euclidean distance of hand movement as the amplitude dimension, records the duration of the action, and compares it with the theoretical beat time code of the accompanying music to obtain the beat deviation dimension.

[0036] The feature library construction module matches the extracted feature vector with the preset action type template to generate the action type code. At the same time, it maps the amplitude and beat deviation to the predefined level interval to generate the corresponding level code. These codes are combined to form the data structure of the action feature library.

[0037] The rule matching module obtains the editing parameters of popular opera videos from the short video platform through the API interface, such as lens switching frequency, special effects usage rules, etc. These parameters are parsed and stored in the system's editing rule library.

[0038] The storyboard generation module first processes the audience feedback audio waveform data, calculates the peak amplitude density and duration of the applause, and generates an emotion intensity model. Then, it performs a time alignment analysis on the emotion intensity peak interval and the encoding in the action feature library. Based on this dynamic association, the system selects video clips that meet the editing rules and reassembles the storyboards.

[0039] During the storyboard reconstruction process, the system adjusts the connection between video segments according to the shot switching interval parameters of the editing rule library. For special effects triggering, when it detects that the movement amplitude exceeds the preset threshold, the system will add particle special effects at the corresponding position. Finally, the system adjusts the duration of the reconstructed video sequence to ensure that it meets the duration requirements of the short video platform.

[0040] Through the above scheme, the present application realizes the intelligent editing and creation of opera performance videos. The system can accurately identify the artistic characteristics of the opera's stylized movements and establish a dynamic association with the audience's emotional feedback. This method significantly improves the matching degree between the editing content and the audience's expectations, and solves the problems existing in traditional editing systems when dealing with opera performances. Specifically, the system can retain complete stylized action clips, avoiding the loss of artistic value caused by incorrect segmentation. At the same time, through the precise synchronization of audience feedback and video content, the short videos generated by the system can better capture the wonderful moments of the performance and the emotional resonance points of the audience. In addition, the editing strategy based on platform-specific rules ensures the platform adaptability of the generated content and improves the dissemination effect of the video. This intelligent editing method not only improves the efficiency of opera video production, but also provides technical support for the promotion of opera art on short video platforms.

[0041] In some of the above solutions of the present application, it is proposed to associate video clips with the audio waveform data of the audience's feedback. However, the collection of applause in the auditorium lacks a quantitative evaluation mechanism and cannot accurately capture the climax period of emotional resonance, resulting in a deviation between the edited segment and the actual emotional fluctuations of the audience.

[0042] The present application further proposes to collect the applause waveform in the auditorium through a directional microphone array and generate an emotional intensity model with timestamps. The calculation formula of the emotional intensity model is: Emotional intensity value = Peak density of waveform amplitude × Duration coefficient. The peak density of waveform amplitude represents the number of peaks exceeding the preset amplitude threshold per unit time. When the applause duration is less than or equal to 5 seconds, the duration coefficient takes a value of 0.5. When the applause duration is greater than 5 seconds and less than or equal to 20 seconds, the duration coefficient = 1.0 + 0.05 * (applause duration - 5). When the applause duration is greater than 20 seconds, the duration coefficient takes a value of 2. The peak interval is defined as the period when the emotional intensity value continuously exceeds the first preset threshold.

[0043] Among them, the directional microphone array covers the auditorium area in a spatial distribution manner, suppresses environmental noise through the beamforming algorithm, and extracts the pure applause waveform. The peak density of waveform amplitude statistically counts the number of peaks through a sliding time window. The window length is set to 1 second, and the preset amplitude threshold is dynamically adjusted according to historical applause data. The calculation of the duration coefficient introduces a piecewise function, adopting a linear growth strategy in the interval of 5 seconds to 20 seconds and a fixed gain factor above 20 seconds. The identification of the peak interval adopts a sliding window accumulation algorithm. When the emotional intensity values of three consecutive windows exceed the first preset threshold, it is determined as a valid peak interval.

[0044] Specifically, during the video editing process, after the applause waveform in the auditorium is captured by the directional microphone array, the timestamps are strictly synchronized with the video frame sequence. The peak density of waveform amplitude statistically counts the number of peaks exceeding the dynamic threshold per unit time through a sliding time window, avoiding the problem of insufficient sensitivity caused by a single fixed threshold. The duration coefficient is dynamically adjusted according to the applause duration. Short applause uses a low weight, medium-duration applause uses a linearly increasing weight, and long applause uses a high weight, reflecting the emotional intensity differences of applause with different durations. Through the continuous threshold detection of the emotional intensity value, the time period when the audience's emotions burst out can be accurately identified, providing an objective quantitative basis for video segment extraction. The generated peak interval is spatially and temporally aligned with the video action features, ensuring that the edited segment meets both the artistic rules and the audience's emotional resonance requirements.

[0045] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0046] The processing of the audio waveform data of the audience's feedback includes the following steps:

[0047] First, the applause waveform in the auditorium is collected through a directional microphone array to generate an emotional intensity model with timestamps. The directional microphone array consists of multiple high-sensitivity microphones evenly distributed along the auditorium to capture applause signals from all directions.

[0048] Secondly, the calculation formula for the emotional intensity model is: Emotional intensity value = Peak density of waveform amplitude × Duration coefficient.

[0049] The peak density of waveform amplitude represents the number of peaks exceeding the preset amplitude threshold per unit time. For example, the preset amplitude threshold can be set to 60 decibels, and the number of peaks exceeding this threshold per second is counted.

[0050] The value of the duration coefficient is divided into three cases according to the duration of applause: when the duration of applause is less than or equal to 5 seconds, the duration coefficient takes the value of 0.5; when the duration of applause is greater than 5 seconds and less than or equal to 20 seconds, the duration coefficient = 1.0 + 0.05 × (duration of applause - 5); when the duration of applause is greater than 20 seconds, the duration coefficient takes the value of 2.

[0051] Finally, the peak interval is defined as the time period when the emotional intensity value continuously exceeds the first preset threshold. The first preset threshold can be set to 5. When the emotional intensity value exceeds 5 for more than 3 seconds continuously, it is recognized as a peak interval.

[0052] Through the above technical solutions, the present application realizes the precise quantification of the audience's feedback. By collecting the applause waveform through a directional microphone array and combining the peak density of waveform amplitude and the duration coefficient, an emotional intensity model with a time dimension is generated. This method not only considers the intensity of applause but also takes into account the influence of the duration, thus more comprehensively reflecting the emotional feedback of the audience. By setting the peak interval, the time period when the audience reacts most enthusiastically is further identified, providing a reliable reference basis for subsequent video editing. This quantification method overcomes the subjectivity and inconsistency of traditional manual judgment, improving the accuracy and efficiency of audience feedback analysis.

[0053] In some of the above solutions of the present application, when constructing the action feature library, the precise capture and analysis of the spatio-temporal features of the bone movement trajectory are lacking, resulting in the generation of action type codes depending on subjective classification criteria, making it difficult to accurately match the benchmark stylized actions of opera genres and affecting the accuracy of subsequent clip rule matching.

[0054] The present application further proposes that the data structure of the action feature library is a set of triples of action type encoding, amplitude level encoding, and beat deviation level encoding. The construction of the action feature library includes obtaining the actor's skeletal motion trajectory through the skeletal key point capture technology, extracting the trajectory spatio-temporal features of the skeletal motion trajectory, and the trajectory spatio-temporal features include the motion trajectory curvature and acceleration parameters; matching the trajectory spatio-temporal features of the actor's skeletal motion trajectory with the preset stylized action classification rule library to generate the action type encoding including the stylized action classification encoding, where the preset stylized action classification rule library stores the benchmark stylized action data of each opera genre, and the benchmark stylized action data includes the spatio-temporal features of the standard skeletal motion trajectory and its corresponding classification encoding and action type encoding; calculating the amplitude dimension of the stylized action as the maximum Euclidean distance of the hand movement in the three-dimensional coordinate system, and comparing the Euclidean distance with the preset amplitude level threshold to generate the action amplitude level encoding; calculating the beat deviation dimension as the absolute value of the difference between the actual action time code and the accompaniment theoretical beat time code, and mapping the absolute value of the difference to the preset deviation interval to generate the beat deviation level encoding.

[0055] Among them, the skeletal key point capture technology adopts a multi-sensor fusion scheme, synchronously collects the skeletal joint coordinate data through an infrared depth camera and an inertial measurement unit to generate a time-continuous three-dimensional motion trajectory. The motion trajectory curvature in the trajectory spatio-temporal features calculates the local curvature value of the trajectory curve through the cubic spline interpolation algorithm, and the acceleration parameter extracts the second derivative from the displacement data through the differential method. The spatio-temporal features of the standard skeletal motion trajectory stored in the preset stylized action classification rule library include the stylized action benchmark data of genres such as Peking Opera and Kunqu Opera. For example, the standard curvature range of the "cloud hand" action in Peking Opera is 0.15 - 0.25 radians / meter, and the acceleration threshold is 1.2 m / s². During the generation of the action amplitude level encoding, the preset amplitude level threshold is set to five intervals, namely, the tiny amplitude encoding (0 - 30 cm), the small amplitude encoding (30 - 60 cm), the medium amplitude encoding (60 - 100 cm), the large amplitude encoding (100 - 150 cm), and the extremely large amplitude encoding (>150 cm). In the mapping rule of the beat deviation level encoding, the preset deviation interval is divided into five levels with 0.1 second as the basic unit. For example, when the absolute value of the deviation is less than 0.05 second, it is mapped to level 0, and when it is 0.05 - 0.1 second, it is mapped to level 1.

[0056] Specifically, after the joint coordinate data obtained by the skeletal key point capture technology is denoised by Kalman filtering, high-precision three-dimensional motion trajectories are generated. The motion trajectory curvature calculation module uses a sliding window algorithm to extract a curvature feature sequence with a time window of 0.5 seconds, and determines the main curvature feature value through peak detection. When calculating the acceleration parameter, a second-order difference operation is performed on the displacement data and the moving average method is used to eliminate high-frequency noise. When matching the preset stylized action classification rule library, the similarity between the real-time trajectory and the standard trajectory is calculated through the dynamic time warping algorithm. When the similarity exceeds 85%, the classification code generation is triggered. During the generation of the action amplitude level code, the maximum displacement of the hand joint coordinates is calculated by the Euclidean distance between the starting point and the ending point in the three-dimensional space and compared with the preset genre-specific threshold. In the generation logic of the beat deviation level code, the accompaniment theoretical beat time code is extracted through the music rhythm analysis algorithm, and the actual action time code is determined based on the acceleration extreme points of the skeletal motion trajectory.

[0057] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0058] The data structure of the action feature library is a triple set of {action type code, amplitude level code, beat deviation level code}. The construction of the action feature library includes the following steps:

[0059] First, the actor's skeletal motion trajectory is obtained through the skeletal key point capture technology. Specifically, the system uses a multi-view camera array to capture the actor's performance process from different angles. Each camera records the actor's actions at a sampling rate of 60 frames per second, and extracts the three-dimensional coordinate data of 17 key skeletal points of the human body through a deep learning algorithm. These key points include positions such as the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. The system extracts the trajectory spatio-temporal features from these skeletal motion trajectories, and the trajectory spatio-temporal features include the motion trajectory curvature and the acceleration parameter.

[0060] Furthermore, the system calculates the displacement vector between each key point in consecutive frames to obtain the velocity data, and the acceleration parameter is calculated by the difference between adjacent frame velocity vectors. For the curvature parameter, the system fits the motion trajectory of each key point into a parametric curve in the three-dimensional space, and then calculates the curvature value at each point. These curvature values and acceleration values constitute the core data of the trajectory spatio-temporal features.

[0061] Secondly, based on the trajectory spatio-temporal features of the actor's skeletal motion trajectory, the preset stylized action classification rule library is matched to generate an action type code including the stylized action classification code. The preset stylized action classification rule library stores the benchmark stylized action data of each opera genre, and the benchmark stylized action data includes the trajectory spatio-temporal features of the standard skeletal motion and the corresponding classification code and action type code.

[0062] For example, the "posing" movement in Peking Opera is encoded as A001, the "running in a circle" movement is encoded as A002, and the "cloud hands" movement is encoded as A003. The system calculates the similarity between the spatio-temporal features of the actual movement trajectory captured and the standard features in the rule library. When the similarity exceeds a predetermined threshold, the corresponding action type code is assigned to the current action segment.

[0063] Next, the system calculates the amplitude dimension of the stylized movement. The amplitude dimension is defined as the maximum Euclidean distance of hand movement in a three-dimensional coordinate system. In specific implementation, the system tracks the position changes of the left and right wrist key points during the entire movement process, calculates the Euclidean distance between any two time points, and takes the maximum value as the quantization index of the amplitude dimension.

[0064] Finally, the system calculates the beat deviation dimension. The beat deviation dimension is the absolute value of the difference between the actual action time code and the theoretical beat time code of the accompaniment. The system extracts the beat points in the accompaniment through audio analysis to obtain the theoretical beat time code. At the same time, the system marks the start and end time points of the action to obtain the actual action time code. After calculating the absolute value of the difference between the two, the system maps the absolute value of the difference to a preset deviation interval to generate a beat deviation level code. For example, when the absolute value of the difference is less than 0.1 second, the code is T1 (accurate); when the difference is between 0.1 - 0.3 seconds, the code is T2 (slight deviation); when the difference is greater than 0.3 seconds, the code is T3 (significant deviation).

[0065] Through the above technical solutions, this application realizes the accurate quantification and encoding of opera movement features, establishes a structured action feature library. This feature library standardizes key dimensions such as action type, amplitude, and beat deviation through a triple data structure, providing a data basis for subsequent video editing. The system can automatically identify and classify complex opera stylized movements, reducing the workload of manual annotation, improving the accuracy of feature extraction. The motion trajectory data obtained through the bone key point capture technology has high precision and high time resolution, making the quantification of action features more objective and reliable. In addition, encoding the action amplitude and beat deviation simplifies the data processing flow, enhances the practicality and adaptability of the system. This structured action feature library provides professional opera domain knowledge support for the intelligent video editing system, solving the problem of inaccurate recognition caused by the lack of domain knowledge in general editing systems.

[0066] In some of the above solutions of this application, it is proposed to extract the spatio-temporal features of the actor's bone motion trajectory through the bone key point capture technology and match them with the standard features in the preset stylized action classification rule library. However, in the matching process, the similarity calculation method is not clear, resulting in possible errors in the matching results and affecting the accuracy of the generation of action type codes.

[0067] The present application further proposes to calculate the similarity between the spatio-temporal characteristics of the trajectory of the actor's skeletal movement and the spatio-temporal characteristics of the standard skeletal movement trajectory in the stylized action classification rule library; when the similarity exceeds the first similarity threshold, it is determined that the match is successful.

[0068] Among them, the calculation of the similarity adopts a mathematical modeling method based on the spatio-temporal characteristics of the trajectory, including the comprehensive calculation of the Euclidean distance of the curvature difference degree of the movement trajectory and the acceleration parameter; the first similarity threshold is set to a dynamic range between 85% and 95%, and the specific value is adaptively adjusted according to the differences in opera genres; the logic for determining a successful match includes a two-level verification mechanism. When the curvature difference degree in the first-level verification is lower than the preset curvature threshold, the second-level verification is triggered to perform time-domain alignment analysis on the acceleration parameter.

[0069] Specifically, during the process of skeletal movement trajectory matching, first, the Fourier transform is performed on the curvature parameter of the actor's skeletal movement trajectory to extract the frequency-domain feature vector, and the cosine similarity is calculated with the frequency-domain vector of the standard feature to obtain the preliminary similarity value; when this value reaches the lower limit of the first similarity threshold, further perform dynamic time warping processing on the time-domain sequence of the acceleration parameter to eliminate the error caused by the time-axis offset, and calculate the root mean square error of the regularized acceleration sequence; the final similarity is determined by the weighted calculation result of the frequency-domain similarity and the acceleration error. By setting double calculation conditions and dynamic thresholds, it effectively avoids misjudgment caused by single-feature matching errors, ensures the accurate establishment of the corresponding relationship between the action type coding and the standard actions of opera genres, and thus provides a reliable data basis for the construction of the subsequent action feature library.

[0070] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0071] The matching preset stylized action classification rule library includes:

[0072] Calculate the similarity between the spatio-temporal characteristics of the trajectory of the actor's skeletal movement and the spatio-temporal characteristics of the standard skeletal movement trajectory in the stylized action classification rule library;

[0073] When the similarity exceeds the first similarity threshold, it is determined that the match is successful.

[0074] Specifically, first, the skeletal movement trajectory data of the actor is obtained through the skeletal key point capture technology. Further, the spatio-temporal characteristics of the trajectory are extracted from the skeletal movement trajectory data, including the curvature and acceleration parameters of the movement trajectory. Among them, the curvature reflects the degree of bending of the action trajectory, and the acceleration parameter reflects the speed change of the action.

[0075] Thus, the extracted spatio-temporal features of the trajectory are compared with the standard skeletal motion trajectory features stored in the pre-established stylized action classification rule library. For example, the cosine similarity algorithm is used to calculate the similarity between two sets of feature vectors. As a preferred implementation, the first similarity threshold can be set to 0.85. When the calculated similarity is greater than 0.85, it is determined that the match is successful, and it is determined that the action belongs to the corresponding stylized action type in the rule library.

[0076] Through the above technical solutions, the present application realizes the accurate recognition and classification of stylized actions in opera performances. Since a special opera action rule library is established and a similarity matching method is adopted, the accuracy of action recognition is improved, and the limitations of general action recognition algorithms in the opera field are avoided. This provides reliable action type information for subsequent video editing and special effect generation, and helps to generate short video content that better conforms to the characteristics of opera art.

[0077] In some of the above solutions of the present application, the stylized action classification rule library relies on the preset standard spatio-temporal features of skeletal motion for action matching. When non-standard action features appear in actual performances, the original rule library cannot automatically identify and update, resulting in a deviation between the beat deviation level coding in the action feature library and the real-time performance data.

[0078] The present application further proposes a real-time data enhancement module, which is configured to continuously obtain the latest opera performance video data stream, extract the spatio-temporal features of the real-time skeletal motion; perform a difference analysis on the spatio-temporal features of the real-time skeletal motion and the standard features in the stylized action classification rule library to generate a feature drift coefficient; when the feature drift coefficient exceeds the preset drift threshold, trigger an incremental update operation of the feature library construction module: write the real-time spatio-temporal features that meet the second similarity threshold into the stylized action classification rule library, and generate corresponding new action type codes; recalculate the beat deviation level coding in the action feature library based on the updated rule library.

[0079] Among them, continuously obtaining the latest opera performance video data stream is realized through a dynamic video stream parsing interface, with a frame rate of no less than 30 frames per second. The difference analysis uses the dynamic time warping algorithm to calculate the morphological difference value between the real-time skeletal trajectory and the standard trajectory. The morphological difference value is normalized to generate a feature drift coefficient in the range of 0-1. The preset drift threshold is 0.6. When the feature drift coefficient exceeds this threshold, it indicates that the real-time action features have exceeded the coverage of the original rule library. The second similarity threshold is 85%. The cosine similarity algorithm is used to determine the matching degree between the real-time spatio-temporal features and the existing features in the rule library. The features that meet this threshold are marked as valid new data. The new action type code adopts the dynamic hash coding method, with a coding length of 16 bits. The first 8 bits inherit the original classification code, and the last 8 bits record the timestamp and genre identifier of the new features.

[0080] Specifically, the video stream parsing interface captures the key frame sequence of the performance video in real time, extracts the three-dimensional coordinates of the actor's joints in each frame through the skeletal key point detection model, and the dynamic time warping algorithm aligns the real-time skeletal trajectory with the standard trajectory in the rule library for non-equal-length time series. The morphological difference value of the trajectory curvature and acceleration parameters is calculated. When the feature drift coefficient exceeds 0.6, the incremental update mechanism is activated, and the system automatically screens the real-time features with a similarity of 85% and stores them in the rule library. After the rule library is updated, the feature library construction module traverses all the stored action data again, recalculates the beat deviation level encoding based on the updated standard features to ensure data consistency. The above process adopts a batch processing mechanism and is automatically executed during the idle period of the system to avoid affecting the operation efficiency of the real-time editing function.

[0081] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0082] The real-time data enhancement module continuously obtains the latest data stream of the opera performance video and extracts the spatio-temporal features of the real-time skeletal movement. The spatio-temporal features of the real-time skeletal movement are analyzed for the degree of difference from the standard features in the stylized action classification rule library to generate a feature drift coefficient. When the feature drift coefficient exceeds the preset drift threshold, an incremental update operation of the feature library construction module is triggered.

[0083] Specifically, the real-time data enhancement module first obtains the real-time video stream of the opera performance through the video acquisition device. Then, computer vision algorithms are used to extract the skeletal key points of the actor from the video frames, including positions such as the head, shoulders, elbows, wrists, hips, knees, and ankles. Based on the temporal changes of these key points, spatio-temporal feature parameters such as the trajectory curvature and acceleration of the skeletal movement are calculated.

[0084] Furthermore, the extracted spatio-temporal features of the real-time trajectory are compared with the standard features stored in the stylized action classification rule library. By calculating metrics such as the Euclidean distance or cosine similarity, the degree of difference between the real-time features and the standard features is quantified to obtain the feature drift coefficient.

[0085] Thus, when the feature drift coefficient continuously exceeds the preset threshold, the system determines that a new action type or a change in the performance style has occurred. At this time, the incremental update mechanism of the feature library is triggered, and the real-time spatio-temporal features that meet the second similarity threshold are written into the stylized action classification rule library and a new action type code is assigned to them.

[0086] For example, assume that a new hand movement is detected, and the trajectory curvature and acceleration features of it still have a 30% difference from the most similar action in the existing library. The system adds this new action feature to the rule library and assigns a new code such as "XQ-001".

[0087] Finally, recalculate the beat deviation level encoding of each action in the action feature library based on the updated rule library. This step ensures the timeliness of the feature library, enabling the system to adapt to new types of actions and performance styles that emerge in Chinese opera performances.

[0088] Through the above technical solutions, the present application realizes the adaptive update of the Chinese opera action feature library: the system can timely capture and learn newly emerged action types, improving the recognition accuracy of diverse and innovative actions in Chinese opera performances. At the same time, by continuously optimizing the beat deviation encoding, the system's adaptability to performance rhythm changes is enhanced, providing more accurate action feature data support for subsequent video editing.

[0089] In some of the above solutions of the present application, the parsing of the preset editing rule library lacks adaptation to the characteristics of Chinese opera videos, and fails to establish an accurate mapping relationship between special effect triggers and action amplitudes, resulting in a mismatch between the shot transition rhythm of the generated video and the rhythm of Chinese opera performances, and a deviation between the application timing of special effects and the artistic expression requirements.

[0090] The present application further proposes a specific implementation method for parsing the preset editing rule library, including: obtaining the shot transition interval parameters of popular Chinese opera videos from short video platforms; establishing a mapping rule between special effect trigger conditions and action amplitudes, and activating the particle special effect generation instruction when it is detected that the amplitude dimension exceeds the second preset threshold.

[0091] Among them, the acquisition of the shot transition interval parameters is realized by calling the open interface of the short video platform, and this interface returns a set of average shot duration data of Chinese opera videos within a preset time window. The mapping rule construction of the special effect trigger conditions adopts an association matrix between the action amplitude dimension and the special effect type, and each amplitude level in the matrix corresponds to a preset particle special effect parameter combination. The value range of the second preset threshold is obtained by analyzing the action amplitude statistical values corresponding to high-like-rate segments in historical video data, and a typical value is the normalized amplitude value of 0.8.

[0092] Specifically, the process of obtaining the shot transition interval parameters includes a data cleaning step, filtering non-Chinese opera videos and low-quality video data, and retaining sample data that conforms to the characteristics of Chinese opera genres. In the execution logic of the special effect trigger conditions, the activation of the particle special effect generation instruction is realized by real-time monitoring of the amplitude level encoding output by the action feature library, and when the encoded value reaches the preset threshold, a special effect parameter package is sent to the rendering engine. This solution realizes automated processing while ensuring artistic expressiveness by quantifying the corresponding relationship between the action amplitude of Chinese opera performances and special effect triggers, making the generated short videos conform to the platform's popular trends and retain the characteristics of Chinese opera art.

[0093] As a preferred embodiment, the solution of the present application is specifically implemented as follows: The parsing of the preset editing rule library includes two key steps:

[0094] The first step is to obtain the shot transition interval parameters of popular opera videos from short video platforms. The system connects to the API interfaces of mainstream short video platforms such as Douyin and Kuaishou, collects 500 short opera video samples with a playback volume of over 1 million, and extracts the timestamp data of the shot transition points in each sample. By calculating the differences between adjacent timestamps, a dataset of shot durations is obtained. Statistical analysis is performed on these data to generate a shot transition interval parameter table, which includes the minimum transition interval (1.5 seconds), the maximum transition interval (8 seconds), the average transition interval in the climax section (2.3 seconds), and the average transition interval in the transition section (5.7 seconds).

[0095] The second step is to establish a mapping rule between the special effect trigger condition and the action amplitude. When it is detected that the amplitude dimension exceeds the second preset threshold, the system activates the particle special effect generation instruction, which includes three key parameters: special effect type selection, special effect density parameter, and special effect duration. These parameters are dynamically adjusted according to the action amplitude value. For example, when it is detected that the amplitude of the water sleeve action exceeds the second preset threshold, the system will select the streamline particle special effect and set a higher particle density and a longer special effect duration.

[0096] Through the above technical solutions, this application realizes the automatic parsing and application of the opera short video editing rules: by extracting the shot transition interval parameters from popular opera short videos, the system can simulate the editing rhythm of professional editors, making the generated short videos conform to the aesthetic habits of the audience. This data-driven editing rule parsing method avoids the inconsistencies caused by subjective judgments in traditional manual editing and improves the standardization degree and visual expressiveness of opera short video production.

[0097] In some of the above solutions of this application, it is proposed to intercept video segments through overlap calculation and perform split-screen recombination. However, in the process of simply relying on time code matching, there may be a situation where non-climax action segments are misjudged as emotional peak intervals, resulting in a mismatch between the special effect trigger and the visual language logic of the opera performance, and generating a technical problem of dislocation between the edited segment and the artistic expression.

[0098] This application further proposes to perform dynamic correlation analysis by combining the peak interval of the audience feedback audio waveform data to generate a short video sequence that conforms to the preset editing rule library, including: calculating the overlap degree between the time code of the peak interval and the duration dimension of the spatio-temporal feature vector; when the overlap degree exceeds the preset overlap degree threshold, intercepting the start and end time codes of the corresponding video segment; and performing split-screen recombination on the intercepted segment according to the shot transition interval parameters.

[0099] Among them, the overlap degree calculation adopts the time window sliding comparison algorithm, and the preset overlap degree threshold is set to 80% time interval coverage. Before intercepting the video segment, a dual-channel verification mechanism needs to be executed. In the process of shot recombination, the knowledge graph of opera actions is introduced for semantic constraint. When it is detected that the matching degree between the clothing color feature and the climax action template in the knowledge graph reaches 90% and the action amplitude is lower than the second preset threshold, the particle effect generation instruction is only activated in the overlapping segment. When the non-uniform time scaling algorithm is executed, the accelerated playback rate is dynamically adjusted according to the emotional intensity value. For every 1 unit increase in the emotional intensity value, the playback rate is reduced by 0.2 times the speed.

[0100] Specifically, the time code alignment module inputs the applause peak interval and the action duration dimension into the time window comparator, and the Hamming distance is used to calculate the time interval matching degree. When the matching degree exceeds 80%, the video interception instruction is triggered and the corresponding RGB frame sequence is extracted for color histogram analysis. The cosine similarity between the color feature and the preset template in the knowledge graph is calculated. When the similarity reaches 90% and the action amplitude is lower than the preset threshold, the particle effect generator injects particle parameters at the starting frame of the video interception segment. During the shot recombination process, the shot transition interval parameter is dynamically adjusted by the sliding window algorithm, and the window length is corrected inversely according to the emotional intensity value. For every 10% increase in the emotional intensity value, the window length is shortened by 0.5 seconds. The duration compression process adopts the frame difference adaptive algorithm. When it is detected that the curvature change rate of the bone trajectory exceeds 15 degrees per second, the original speed is maintained for playback, and motion-preserving acceleration based on optical flow estimation is implemented in the transition segment. The acceleration multiple forms an inverse proportional function relationship with the emotional intensity value to ensure that the integrity of the stylized action is not affected by time compression.

[0101] As a preferred embodiment, the solution of the present application is specifically implemented as follows:

[0102] Calculate the overlap degree between the time code of the peak interval and the duration dimension of the spatio-temporal feature vector. Specifically, for each peak interval, calculate the overlapping part of its time range and the action duration in the corresponding video segment. For example, assume that the time range of a certain peak interval is [10s, 15s], and the corresponding action duration is [8s, 18s], then the overlapping part is [10s, 15s], and the overlap degree is (15 - 10) / (18 - 8) = 50%.

[0103] When the overlap degree exceeds the preset overlap degree threshold, intercept the start and end time codes of the corresponding video segment. Further, the overlap degree threshold can be set to 60%. If the calculated overlap degree is greater than 60%, it is considered that the peak interval is highly relevant to the action segment and should be intercepted and retained.

[0104] Perform shot recombination on the intercepted segment according to the shot transition interval parameter. Thus, the retained video segments can be recombined according to the preset shot transition rules.

[0105] For example, if the shot transition interval is set to 3 - 5 seconds, then during the recombination process, a shot transition is performed every 3 - 5 seconds, thereby forming a short video sequence with a compact rhythm.

[0106] Through the above technical solution, the present application realizes intelligent video editing based on audience feedback and action features, improves the generation efficiency and quality of short opera videos. Since the audio waveform data of audience feedback is used as a reference, the editing result is more in line with the audience's preferences. At the same time, by setting the overlap threshold and shot transition rules, the coherence and rhythm of the edited video are ensured, enhancing the viewing experience. In addition, this solution avoids the subjectivity of manual screening and improves the objectivity and consistency of editing.

[0107] In some of the above solutions of the present application, there may be a problem that the inherent correlation between the costume color and the action type in the opera performance may be ignored during the video segment interception, resulting in a mismatch between the special effect trigger logic and the artistic expression. At the same time, due to the lack of spatio - temporal correspondence between the climax action segment and the peak interval of applause, it is easy to insert special effects at non - key frames, affecting the emotional resonance of the edited content.

[0108] The present application further proposes to construct an opera action knowledge graph, whose nodes contain the correlation between action types and costume color features, and whose edge relationships contain the spatio - temporal correspondence between the climax action segment and the peak interval of applause. When it is detected that the matching degree between the color feature of the current frame and the preset color template of the climax action in the opera action knowledge graph meets the preset matching threshold, and the amplitude dimension is lower than the second preset threshold, the particle special effect generation instruction is activated, and the corresponding particle special effect is generated in the overlapping segment between the climax action segment and the peak interval of applause.

[0109] Among them, the node data of the opera action knowledge graph stores the RGB value mapping table of the action type and the costume color feature. The preset color template is matched using the hue threshold interval in the HSV color space. The edge relationship establishes the correspondence between the start time code of the climax action and the peak time code of the applause waveform through the time axis alignment algorithm. The execution condition of the particle special effect generation instruction includes a double - verification mechanism, that is, the color matching degree needs to reach more than 85% and the action amplitude dimension is lower than the threshold of 0.3 m / s.

[0110] Specifically, the system calls the color association rules in the knowledge graph during the storyboard generation stage, and compares the similarity between the clothing color histogram of the video frame and the preset color template in real time. When the color matching degree of five consecutive frames exceeds the 85% threshold, the climax action segment recognition mechanism is triggered. At this time, combined with the amplitude dimension sensor data, if the arm movement speed is detected to be lower than 0.3 meters per second, it is judged as a static appearance action, which meets the typical characteristics of the climax segment in the opera performance. The system simultaneously retrieves the applause peak interval database and calculates the overlap between the action segment and the applause interval through the dynamic time warping algorithm. When the overlapping duration exceeds 70%, the corresponding video segment is superimposed with a golden particle light effect. The particle density of this light effect is positively correlated with the emotional intensity value. In segments where the applause lasts for more than 20 seconds, the particle transparency is automatically enhanced to 80%, achieving precise synchronization between the visual effect and the audience's emotional feedback.

[0111] Through the above technical solution, this application effectively solves the problem of insufficient consistency between the timing of special effects triggering and artistic expression in opera video editing.

[0112] In some of the above-mentioned schemes in this application, there is a problem of insufficient accuracy in identifying climax action segments during the dynamic correlation analysis process. Traditional methods rely on movement amplitude and beat deviation parameters, and do not consider the strong correlation between costume color and movement type in opera performances, resulting in a mismatch between the timing of special effects triggering and the needs of artistic expression.

[0113] The present application further proposes to construct a knowledge graph of opera actions, whose nodes include the association between action types and costume color features, and whose edge relationships include the spatiotemporal correspondence between climax action segments and applause peak intervals; when it is detected that the matching degree between the current frame color features and the preset color template of the climax action in the opera action knowledge graph meets the preset matching threshold, and the amplitude dimension is lower than the second preset threshold, the particle special effect generation instruction is activated to generate the corresponding particle special effect in the overlapping segment of the climax action segment and the applause peak interval.

[0114] Among them, the nodes of the opera action knowledge graph use HSV color space histogram to quantify the color features of costumes, and preset color templates store the standard costume color range of roles such as Qingyi and Huadan; the edge relationship establishes a mapping relationship between the starting time code of the climax action and the applause peak interval through the time axis alignment algorithm; the amplitude dimension threshold is set to 30%-50% of the standard action amplitude to distinguish between regular actions and high-difficulty technical actions; the particle special effect generation instruction calls the GPU rendering engine to generate dynamic light effects according to the particle morphology parameters associated with the action type in the knowledge graph.

[0115] Specifically, in the stage of shot recombination, the cosine similarity between the HSV histogram features of video frames and the color templates in the knowledge graph is calculated. When the similarity exceeds 85%, a color matching event is triggered; the action amplitude dimension is calculated through the three-dimensional coordinate differences of the skeletal key points. If the amplitude is lower than the second preset threshold and the color matching is successful, it is determined as a climax paragraph that requires enhanced visual effects; the particle special effect generation instruction calls the preset particle parameters according to the action type encoding. For example, the water sleeve action is associated with fluid particles, and the martial arts action is associated with spark particles; the generated particle special effect time code is limited to the overlapping part of the applause peak interval and the climax action paragraph, and the real-time synthesis of the video paragraph is achieved through the FFmpeg filter chain. This solution improves the artistic expressiveness and editing accuracy while reducing the computational complexity by integrating visual features and spatio-temporal association rules.

[0116] As a preferred embodiment, the solution of the present application is specifically implemented as follows: In the process of video paragraph recombination, a non-uniform time axis adjustment method based on motion trajectory feature detection is adopted. First, the convolutional neural network is used to identify the integrity boundary of the stylized action, and the original playback rate is maintained within the range from the start frame to the end frame of the action. In the transition paragraph between adjacent actions, the dynamic time warping algorithm is used to sample the video stream. The acceleration coefficient of the transition paragraph is dynamically controlled by the quantization value output by the action amplitude detection module. The execution condition of the accelerated playback is realized by real-time monitoring of the hand movement distance in the three-dimensional space. When it is detected that the amplitude value is lower than the preset action intensity threshold and the beat time difference is less than 0.12 seconds, the time compression function is automatically activated, and the acceleration rate is adjusted inversely according to the real-time calculated audience emotion intensity value, which is specifically realized by establishing a linear mapping relationship between the emotion intensity and the playback rate, ensuring a higher acceleration ratio in the transition paragraphs during low emotion intensity periods.

[0117] Through the above technical solution, the present application effectively solves the dual problems of low efficiency of manual screening and damage to artistic integrity by the automated system in traditional Chinese opera video editing. Through the intelligent non-uniform time processing mechanism, while maintaining the artistic expressiveness of the core actions, the duration of non-key paragraphs is significantly shortened. Combined with the dynamic rate adjustment based on the audience's emotional feedback, the finally generated short video not only conforms to the artistic norms of Chinese opera performances but also can accurately adapt to the viewing expectations of different audience groups, achieving the optimal balance between artistic expression and dissemination efficiency.

[0118] In some of the above solutions of the present application, when the non-uniform time scaling algorithm performs accelerated playback processing in the transition paragraph, if the acceleration rate is not properly processed, it may damage the synchronization between the emotion intensity and the video rhythm, resulting in a deviation between the emotional expression of the edited paragraph and the audience feedback.

[0119] The execution conditions of the non-uniform time scaling algorithm further proposed in this application include: when it is detected that the action amplitude dimension of the current video segment is lower than the third preset threshold and the beat deviation dimension is less than the predetermined number of digits, accelerate the playback process; the acceleration rate of the playback process is inversely proportional to the emotional intensity value.

[0120] Among them, the action amplitude dimension is quantified by the maximum Euclidean distance of hand movement in a three-dimensional coordinate system. The third preset threshold is the lower limit value of the numerical range of the Euclidean distance, and the predetermined number of digits is set as the millisecond-level error range of the accompaniment theoretical beat time code. The emotional intensity value is calculated from the peak density of the waveform amplitude and the duration coefficient, and the acceleration rate is dynamically adjusted according to the reciprocal of this value.

[0121] For example, when the emotional intensity value is 2, the acceleration rate is 0.5 times the original playback rate; when the emotional intensity value is 1, the acceleration rate is 1 times the original playback rate.

[0122] Specifically, in the stage of shot recombination, the system first compares the action amplitude dimension with the third preset threshold. If it is detected that the action amplitude of the current segment is lower than this threshold and the beat deviation dimension is within the range of the predetermined number of digits, the accelerated playback processing module is triggered. The emotional intensity value acts on the rate calculation unit in real time to generate an acceleration coefficient inversely proportional to the emotional intensity. In the paragraphs with a higher emotional intensity corresponding to the peak interval of applause, the system reduces the acceleration rate to extend the display time of climax actions; in the non-peak intervals with a lower emotional intensity, the acceleration rate is increased to compress the duration of the transition paragraphs. This processing mechanism controls the time scaling ratio through quantization parameters, and while maintaining the integrity of the stylized actions, realizes the dynamic adaptation of the video rhythm to the emotional fluctuations of the audience.

[0123] As a preferred embodiment, the solution of this application is specifically implemented as follows: in the process of video duration compression processing, a dynamic speed regulation mechanism based on action features and emotional intensity is adopted, that is, when the bone movement trajectory analysis module detects that the maximum Euclidean distance of hand movement in the current video segment is less than the preset amplitude threshold and the absolute value of the beat time difference is less than 0.15 seconds, the non-uniform time scaling controller is activated.

[0124] The above non-uniform time scaling controller dynamically adjusts the video playback rate according to the emotional intensity value obtained by parsing the real-time applause waveform: when the emotional intensity value reaches the 80th percentile, set the playback rate to 1.2 times; when the emotional intensity value drops to the 40th percentile, adopt a playback rate of 1.8 times. In the complete display segment of the stylized actions, the system maintains the original playback rate to ensure the integrity of key actions such as sleeve throwing and cloud hands.

[0125] Through the above technical solutions, the present application effectively solves the problem that traditional video compression technologies damage the integrity of artistic movements. Through the linkage analysis of movement features and emotional feedback, intelligent duration optimization is achieved while retaining the essence of opera performances. Specifically: based on the dual detection mechanism of movement amplitude and beat deviation, the transitional paragraphs suitable for acceleration are accurately identified; combined with the inverse proportional speed regulation strategy of emotional intensity values, it is ensured that the emotional expression intensity of the climax paragraphs is not damaged by mechanical compression, and finally a clip work that meets the duration requirements of short video platforms and is faithful to the characteristics of opera art is generated.

[0126] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A video editing and art creation system based on AI intelligence, characterized in that, Including: A data acquisition module, which is used to obtain the video data of the opera stage and the corresponding audio waveform data of the audience feedback; A feature analysis module, which is used to extract the spatio-temporal feature vectors of the stylized actions from the opera stage video data, and the spatio-temporal feature vectors include an amplitude dimension, a duration dimension and a beat deviation dimension; A feature library construction module, which is used to construct an action feature library containing action type codes and music beat deviation level codes according to the spatio-temporal feature vectors; A rule matching module, which is used to parse the preset editing rule library of the short video platform, and the editing rule library contains shot switching interval parameters and special effect triggering conditions; A storyboard generation module, which is used to call the action type codes and beat deviation level codes in the action feature library, and perform dynamic correlation analysis in combination with the peak interval of the audience feedback audio waveform data to generate a short video sequence that conforms to the preset editing rule library.

2. The video editing and artistic creation system based on AI intelligence according to claim 1, characterized in that, The processing of the audience feedback audio waveform data includes: Collecting the applause waveform of the audience seats through a directional microphone array to generate an emotion intensity model with time stamps; The calculation formula of the emotion intensity model is: emotion intensity value = waveform amplitude peak density × duration coefficient; Wherein, the waveform amplitude peak density represents the number of peaks exceeding the preset amplitude threshold within a unit time. When the applause duration is less than or equal to 5 seconds, the duration coefficient takes a value of 0.

5. When the applause duration is greater than 5 seconds and less than or equal to 20 seconds, the duration coefficient = 1.0 + 0.05 * (applause duration - 5). When the applause duration is greater than 20 seconds, the duration coefficient takes a value of 2. The peak interval is defined as the time period when the emotion intensity value continuously exceeds the first preset threshold.

3. An AI intelligence-based video editing and art creation system according to claim 1, characterized in that, The data structure of the action feature library is a triple set of {action type code, amplitude level code, beat deviation level code}, and the construction of the action feature library includes: Obtaining the actor's bone movement trajectory through bone key point capture technology, and extracting the trajectory spatio-temporal features of the bone movement trajectory. The trajectory spatio-temporal features include the curvature of the movement trajectory and the acceleration parameters; Matching the trajectory spatio-temporal features of the actor's bone movement trajectory with the preset stylized action classification rule library to generate action type codes including stylized action classification codes. Among them, the preset stylized action classification rule library stores the benchmark stylized action data of each opera genre, and the benchmark stylized action data includes the spatio-temporal features of the standard bone movement trajectory and its corresponding classification codes and action type codes; Calculating the amplitude dimension of the stylized action as the maximum Euclidean distance of hand movement in a three-dimensional coordinate system, and comparing the Euclidean distance with the preset amplitude level threshold to generate an action amplitude level code; Calculating the beat deviation dimension as the absolute value of the difference between the actual action time code and the accompaniment theoretical beat time code, and mapping the absolute value of the difference to a preset deviation interval to generate a beat deviation level code.

4. An AI intelligence-based video editing and art creation system according to claim 3, characterized in that, The matching of the preset stylized action classification rule library includes: Calculating the similarity between the trajectory spatio-temporal features of the actor's bone movement trajectory and the trajectory spatio-temporal features of the standard bone movement trajectory in the stylized action classification rule library; When the similarity exceeds the first similarity threshold, it is determined that the matching is successful.

5. An AI-intelligent based video editing and art creation system according to claim 4, characterized in that, It further includes a real-time data augmentation module, which is configured to: Continuously obtain the latest opera performance video data stream and extract the spatio-temporal features of the real-time skeletal movement trajectory; Perform a difference analysis on the spatio-temporal features of the real-time skeletal movement trajectory and the standard features in the stylized action classification rule library to generate a feature drift coefficient; When the feature drift coefficient exceeds the preset drift threshold, trigger an incremental update operation of the feature library construction module: Write the real-time spatio-temporal features that meet the second similarity threshold into the stylized action classification rule library and generate corresponding new action type codes; Recalculate the beat deviation level code in the action feature library based on the updated rule library.

6. An AI-intelligent based video editing and art creation system according to claim 1, characterized in that, The parsing of the preset editing rule library includes: Obtain the lens switching interval parameters of popular opera videos from the short video platform; Establish a mapping rule between the special effect trigger condition and the action amplitude. When it is detected that the amplitude dimension exceeds the second preset threshold, activate the particle special effect generation instruction.

7. An AI intelligence-based video editing and art creation system according to claim 6, characterized in that, The dynamic correlation analysis in combination with the peak interval of the audience feedback audio waveform data to generate a short video sequence that meets the preset editing rule library includes: Calculate the overlap degree between the time code of the peak interval and the duration dimension of the spatio-temporal feature vector; When the overlap degree exceeds the preset overlap threshold, intercept the start and end time codes of the corresponding video segment; Reassemble the intercepted segment according to the lens switching interval parameter.

8. An AI intelligence-based video editing and art creation system according to claim 7, characterized in that, Before intercepting the start and end time codes of the corresponding video segment, it further includes: Construct an opera action knowledge graph, the nodes of which contain the association relationship between the action type and the costume color feature, and the edge relationship contains the spatio-temporal correspondence between the climax action segment and the applause peak interval; When it is detected that the matching degree between the current frame color feature and the preset color template of the climax action in the opera action knowledge graph meets the preset matching threshold and the amplitude dimension is lower than the second preset threshold, activate the particle special effect generation instruction to generate corresponding particle special effects in the overlapping segment between the climax action segment and the applause peak interval.

9. An AI intelligence-based video editing and art creation system according to claim 7 or 8, characterized in that, The reassembly of the shots includes: Perform a duration compression process on the intercepted video segment to make the total duration match the target duration in the preset editing rule library; The duration compression process uses a non-uniform time scaling algorithm, maintaining the original playback rate in the segment where the integrity of the stylized action is retained, and using an accelerated playback process in the transition segment.

10. An AI-intelligent based video editing and art creation system according to claim 9, characterized in that, The execution conditions of the non-uniform time scaling algorithm include: When it is detected that the action amplitude dimension of the current video segment is lower than the third preset threshold and the beat deviation dimension is less than the preset number of digits, start the accelerated playback process; The rate of the accelerated playback process is inversely proportional to the emotional intensity value.

Citation Information

Cited By

  • Guqin playing action and audio synchronous analysis method based on multi-modal fusion

    CN120635654A

  • A guqin playing action and audio synchronization analysis method based on multi-modal fusion

    CN120635654B

  • Short video intelligent editing method and system based on multi-modal analysis

    CN120935432A

  • A short video intelligent clipping method and system based on multi-modal analysis

    CN120935432B