System for generating lens by combining action and audio detection

Through the combination of deep learning and long-term memory networks, the generation of film and television lenses is automatically solved, and the problems of low creative efficiency, high artistic threshold and poor music adaptability in traditional film and television production are solved, achieving a high degree of coordination between the lens and content and improving aesthetic quality.

CN120279098APending Publication Date: 2025-07-08赵雅芝
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510477266.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In traditional film and television production, lens design relies on manual experience, has low creative efficiency, high artistic threshold, and poor music adaptability, making it difficult to achieve dynamic synchronization of lens movement with music rhythm and emotional changes.

Method used

A convolutional neural network based on deep learning is used for video key point detection and audio preprocessing, combined with a long and short-term memory network to generate a lens strategy, through the deep fusion of action frequency and music frequency, a lens sequence that conforms to artistic rules is automatically generated.

Benefits of technology

Significantly improve the coordination between the lens and the content, improve video production efficiency, increase creative diversity, optimize data acquisition and processing processes, and improve the aesthetic quality of the lens.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
Patent Text Reader

Abstract

The invention belongs to the field of computer vision and intelligent content generation, and provides a system for generating a lens in combination with action and audio detection. A data acquisition module obtains action clip videos and audios, features are extracted through key point labeling and audio preprocessing, and the resolution is adjusted in a self-adaptive mode. And the frequency analysis module calculates action frequency according to key point coordinate change, constructs a pattern library in a clustering manner, and updates by means of a time window sliding mechanism. The demand input module receives and encodes a lens generation demand of a user. The decision generation module fuses multi-source information, generates a shot strategy according to rules and requirements, and gives consideration to aesthetic rules. And the generation execution module edits and processes the shot according to a strategy, adds a special effect, and outputs an animation shot sequence meeting requirements. In addition, the invention also relates to a corresponding method, electronic equipment and a computer readable storage medium. According to the method, the lens and content coordination, the manufacturing efficiency and the creativity diversity are improved, the data processing flow is optimized, and the lens aesthetic quality is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and intelligent content generation, and particularly relates to a method for generating shots based on action and audio detection, which is particularly applicable to scenarios such as film and television animation previews, short video automated production, and intelligent film and television creation. Background Art

[0002] In the traditional film and television production process, shot design highly depends on manual experience. The director or storyboard artist needs to manually draw the storyboard script according to the script and music rhythm, which has the following defects: low creation efficiency: the design of a single shot requires repeated adjustment of the composition, camera movement trajectory, and alignment with the timeline, taking up to several hours; high artistic threshold: high-quality shot design requires mastering complex film and television language rules; poor music adaptability: existing automated tools are difficult to achieve dynamic synchronization between camera movement and music rhythm and emotional changes. Summary of the Invention

[0003] Object of the Invention The present invention aims to provide a system and method that can automatically generate a shot sequence that conforms to artistic rules and adapts to action and music characteristics, reducing the threshold of film and television creation and improving production efficiency.

[0004] Technical Solution

[0005] The core architecture of the present invention includes the following technical features.

[0006] Data Acquisition Module: This module can widely obtain video data of action clips of various styles and themes through a specially designed video acquisition interface.

[0007] The video acquisition interface supports multiple common video formats, such as MP4, AVI, FLV, etc., to ensure compatibility with video materials from different sources. After obtaining the video data, an advanced convolutional neural network object key point detection algorithm based on deep learning is used to accurately label the key points of the main body in the video. This convolutional neural network model adopts the Residual Network (ResNet) architecture, with multiple convolutional layers and pooling layers. Through training with a large amount of diverse data, it can accurately identify the key point coordinates of key parts. The training data covers action subject images under different postures, different perspectives, and different lighting conditions to improve the generalization ability of the model.

[0008] During the annotation process, the model processes each frame of the image and outputs the two-dimensional coordinate information of each key point. Meanwhile, the module synchronously collects the audio data of the corresponding segment and performs a series of preprocessing operations on it, such as noise reduction and filtering. Spectral subtraction is used for noise reduction, which effectively reduces background noise by estimating the noise spectrum and subtracting it from the original audio spectrum. A finite impulse response (FIR) filter is used for filtering, and the filter coefficients are designed according to the frequency characteristics of the audio to remove unnecessary high-frequency and low-frequency components. Subsequently, mature signal processing techniques such as Fourier transform are used to accurately extract music frequency features, such as key information like fundamental frequency, harmonic frequency, and rhythm changes. To improve the accuracy of frequency feature extraction, the short-time Fourier transform (STFT) is adopted, which divides the audio signal into multiple short time segments and performs Fourier transform on each segment to obtain the spectral information of each segment.

[0009] In actual operation, when the original video resolution is too high, resulting in an excessive processing burden on the system and causing the system to run stuck, the module can automatically adjust the video resolution to the optimal processing resolution based on the original video resolution and the current processing capacity of the system. During the specific adjustment process, appropriate resolution is dynamically calculated according to indicators such as the CPU usage rate, memory occupancy, and video processing speed of the system, and the bilinear interpolation algorithm is used for image scaling to ensure image quality.

[0010] Frequency analysis module: This module deeply analyzes the dynamic changes of the coordinates of object key points based on the annotation data generated by the data acquisition module.

[0011] Specifically, for each key point coordinate sequence, the displacement change amount between adjacent moments is obtained through differential operation, and then combined with the frame rate information of the video, the displacement amount per unit time is accurately calculated. Taking the key point of the arm as an example, the displacement between two adjacent frames is accurately calculated and divided by the frame interval time to obtain the displacement amount per unit time. To improve the accuracy of the calculation, the displacement amount is smoothed using the moving average filtering algorithm to remove outliers caused by noise or jitter. Then, Fourier transform is used to convert these displacement amount time series from the time domain to the frequency domain, and thus the detailed action frequency distribution is obtained.

[0012] When performing Fourier transform, the fast Fourier transform (FFT) algorithm is used to improve the calculation efficiency. The above operations are performed separately for different parts of the object, such as the head, limbs, torso, etc., to calculate the action frequencies of each part. On this basis, the clustering algorithm is used to systematically classify similar action frequency patterns, and a comprehensive and accurate action frequency pattern library is gradually constructed.

[0013] During the clustering process, it is first necessary to determine the number of clusters. By calculating the sum of squared errors for different numbers of clusters, the point where the downward trend of the sum of squared errors slows down is selected as the optimal number of clusters.

[0014] To enable the action frequency pattern library to reflect the dynamic changes of actions in real time, the module introduces a time window sliding mechanism. In a long animation, the actions of the action subject often undergo various changes. For example, a person gradually changes from walking slowly to running fast. The time window sliding mechanism can keenly capture these changes in action frequency patterns and update the action frequency pattern library in a timely manner. Specifically, a time window of a fixed size is set. As time goes by, the window slides forward continuously. After each slide, the action frequency within the window is recalculated and compared with the existing action frequency patterns. If the difference exceeds a certain threshold, the action frequency pattern library is updated, significantly improving the timeliness and accuracy of the action frequency pattern library and providing a more practical reference basis for subsequent shot generation decisions.

[0015] Requirement input module: This module provides a user interface through which users can input the generation requirements for animation shots. These requirements include but are not limited to specific emotional atmospheres, shot type preferences, etc. The requirement input module will parse and encode the text information input by users and convert it into numerical features that can be processed by a computer for subsequent fusion with action frequency patterns and music frequency features.

[0016] Generation decision module: This module uses a long short-term memory network to construct a joint frequency analysis model to deeply fuse the information in the action frequency pattern library with the music frequency features.

[0017] The LSTM network consists of an input layer, multiple LSTM hidden layers, and an output layer. The input layer receives the action frequency sequence and the music frequency sequence as input data. Each sequence is normalized to map the data to the [0, 1] interval to improve the training effect of the model. The LSTM hidden layer contains multiple LSTM units, and each unit has an input gate, a forget gate, and an output gate, which can effectively handle the long-term dependencies in time series data. During the training process, the stochastic gradient descent (SGD) algorithm is used for parameter update. By continuously adjusting the weights and biases of the model, the model can accurately learn the complex temporal synchronization relationship and subtle emotional association between the action frequency sequence and the music frequency sequence.

[0018] The model can accurately output the predicted values of various parameters required for shot generation, such as shot transition time, scene size, camera movement speed, etc. When generating the shot generation strategy, the module makes intelligent decisions based on preset rules. When the music tempo speeds up, that is, when the high-frequency components in the music frequency increase significantly, the model quickly matches the shots corresponding to the high-action frequency pattern in the action frequency pattern library, and then determines parameters such as quickly switching shots and showing the overall action with a large scene size to enhance the visual impact and rhythm. When the music is soothing and the low-frequency components in the music frequency dominate, the model matches the shots corresponding to the low-action frequency pattern and determines parameters such as slowly switching shots and showing subtle actions in close-up to create a delicate and soothing atmosphere.

[0019] When generating the shot generation strategy, the shot generation decision-making module fully considers the shot aesthetic rules. When showing a shot of the protagonist standing in a vast landscape, according to the rule of thirds, the picture is made more attractive and artistic, enhancing the visual beauty of the generated shot. To ensure that the generated shots conform to the aesthetic rules, during the decision-making process, each candidate shot is evaluated, its matching degree with the aesthetic rules is calculated, and the shot with the highest matching degree is selected as the final solution.

[0020] Shot generation execution module: According to the shot generation strategy formulated by the shot generation decision-making module, the shot generation execution module accurately clips the required shots from the original action segments and combines them.

[0021] During the clipping process, video editing algorithms are used to accurately locate the corresponding positions in the original video according to the shot transition time in the shot generation strategy for clipping operations. At the same time, to ensure the quality of the clipped video, video interpolation algorithms are used to adjust the frame rate and smooth the picture of the clipped video. When post-processing the selected shots, the emotion conveyed by the music is fully considered. When the music emotion is exciting, the module automatically increases the color saturation and contrast of the shots to highlight the enthusiastic atmosphere.

[0022] Specifically, the color and contrast are adjusted by adjusting the color space parameters of the image, such as the saturation and lightness values in the HSV space. When the music emotion is sad, the shot color is adjusted to a cold tone to strengthen the melancholy mood. The cold tone atmosphere can be created by shifting the hue of the image towards blue and green directions. At the same time, a variety of transition effects are added, such as fade-in / fade-out, rotation switch, zoom transition, etc., effectively enhancing the smoothness of the transition between shots. When adding transition effects, according to the camera movement method and emotion atmosphere in the shot generation strategy, appropriate transition effects are selected and the smooth transition of the effects is achieved through animation algorithms. Finally, an animation shot sequence that perfectly fits the music and action frequency is output.

[0023] The specific steps of the method flow are as follows.

[0024] (1) With the help of the data acquisition module, a wide range of different action segment videos and corresponding audio data are collected. During this process, the object key point detection algorithm is used to comprehensively detect and label the key points of the objects in the video, providing basic data for subsequent analysis. The specific steps include video format checking, audio-video synchronization, key point detection model loading and prediction, etc.

[0025] (2) The action frequency analysis module deeply analyzes the labeled data, carefully calculates the action frequencies generated by the coordinate changes of the key points of different parts of the object, and constructs a comprehensive and accurate action frequency pattern library through the clustering algorithm. This process includes multiple steps such as displacement calculation, frequency conversion, clustering analysis, and pattern library update.

[0026] (3) The generation requirement input module receives the generation requirements of the animation shots input by the user, and parses and encodes the requirements. The music frequency features are accurately extracted from the audio data, and the information of the action frequency pattern library, the music frequency features, and the user requirement encoding information are organically integrated and input into the joint frequency analysis model to generate a scientific and reasonable shot generation strategy. This involves steps such as feature extraction, data normalization, model training and prediction, and strategy evaluation and optimization.

[0027] (4) The generation execution module performs a series of operations such as carefully editing the original segments, meticulous image processing, and adding appropriate transition effects according to the generated shot generation strategy, and finally generates a high-quality animation shot sequence that conforms to the music, action frequency, and user requirements. This covers specific steps such as editing operations, color adjustment, special effect addition, and video synthesis.

[0028] The present invention relates to a computer-readable storage medium on which a specially developed computer program is stored. When the program is executed by a processor, it can fully implement each step of the above-mentioned animation shot generation method, providing reliable storage support for the implementation of the entire technical solution.

[0029] At the same time, it also includes an electronic device, which consists of a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it can efficiently implement all steps of the above-mentioned animation shot generation method, and through the coordinated work of the hardware device, ensure the stable operation of the entire technical solution. Beneficial effects

[0030] Significantly improve the coordination between shots and content: By accurately and deeply analyzing the object movement frequency and music frequency, the shots generated by the present invention can be highly closely matched with the action of the action subject and the background music. In fierce battle scenes, the fast and exciting music rhythm can perfectly match the high-frequency movements of characters or objects, prompting the system to generate fast-switching, dynamic shots, greatly enhancing the tense and exciting atmosphere, allowing the audience to immerse themselves more deeply in the plot.

[0031] Significantly improve video production efficiency: The automated lens generation process greatly reduces the heavy workload of manually designing lenses. With the technical solution of the present invention, lens work can be completed in just a few hours. This not only significantly improves the efficiency of animation production, but also greatly reduces production costs, bringing higher economic benefits to the animation production industry.

[0032] Effectively increase creative diversity: The organic combination of the action frequency pattern library and the joint frequency analysis model can deeply explore more novel and unique combinations of shots, actions, and music. These new combinations provide creators with rich creative inspiration, help break through the limitations of traditional creative ideas, create more animation shots with unique styles such as fantasy and mystery, enrich the expression of animation, and meet the increasingly diverse aesthetic needs of the audience.

[0033] Optimize data acquisition and processing flow: The adaptive resolution adjustment function of the data acquisition module ensures that available data can be efficiently collected when facing various original video resolutions and system performance conditions, greatly improving the applicability and stability of the entire system. The time window sliding mechanism of the action frequency analysis module enables the action frequency pattern library to reflect the action changes of the action subject in the video in real time and accurately, laying a solid foundation for the subsequent generation of more accurate lens strategies, and further improving the quality and accuracy of lens generation.

[0034] Significantly improve the aesthetic quality of shots: The shot generation decision module fully considers the aesthetic rules of shots, so that the generated shots not only excel in the coordination of action and music, but also reach a higher level in the aesthetics of visual composition. By following aesthetic principles such as the rule of thirds and symmetrical composition, the generated shots are more reasonable and beautiful in terms of screen layout and element arrangement, further improving the viewing and artistic value of the video, and bringing more pleasant visual enjoyment to the audience. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] To more clearly illustrate the technical solutions of the embodiments of the present application, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0036] Figure 1 is the overall system architecture diagram provided by the embodiment of the present invention.

[0037] Figure 2 is the flowchart of the lens generation method provided by the embodiment of the present invention. Specific embodiments

[0038] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0039] Embodiment 1 of the present invention provides a system for generating a lens by combining action and audio detection. The overall system architecture is as Figure 1 shown and includes the following modules.

[0040] The data acquisition module described in Embodiment 1 of the present invention is used to obtain video data of different action segments, perform key point annotation on the video subject using an object key point detection algorithm, generate annotation data containing a key point coordinate sequence, and at the same time collect the audio data of the corresponding segment and perform preprocessing to extract music frequency features.

[0041] The action frequency analysis module described in Embodiment 1 of the present invention calculates the action trajectories and displacement amounts of different parts based on the changes in the object key point coordinates in the annotation data, obtains the action frequencies of each part in combination with time information, classifies similar action frequency patterns through a clustering algorithm, and generates an action frequency pattern library.

[0042] The requirement input module described in Embodiment 1 of the present invention is used to receive the generation requirements of the user for the animation lens, and the requirements include but are not limited to specific emotional atmospheres, theme styles, lens type preferences, etc.

[0043] The generation decision module described in Embodiment 1 of the present invention integrates the information of the action frequency pattern library, the music frequency characteristics, and the requirements input by the user, establishes a joint frequency analysis model, and according to the preset rules and the user's requirements. For example, when the music rhythm speeds up, it matches the high-action-frequency-pattern corresponding shots, and when the music is soothing, it matches the low-action-frequency-pattern corresponding shots. At the same time, combined with the user's requirements for the emotional atmosphere and theme style, it generates a shot generation strategy, and determines parameters such as shot switching points, shot sizes, and camera movement methods.

[0044] The generation execution module described in Embodiment 1 of the present invention clips and combines shots from the original segments according to the shot generation strategy, performs image processing on the selected shots, such as adjusting colors and contrast to match the music emotion and the user's requirements, and adds transition effects, and finally outputs an animated shot sequence that meets the music, action frequency, and user's requirements.

[0045] Furthermore, the data acquisition module described in Embodiment 1 of the present invention is provided with acquisition interfaces specifically adapted to a variety of common video formats (such as MP4, AVI, FLV, etc.), and can widely obtain action segment video data from multiple channels such as professional animation material libraries and user uploads. Using the convolutional neural network object key point detection algorithm based on deep learning, taking the residual network (ResNet) architecture as an example, through training with a large amount of diverse data covering different postures, perspectives, and lighting conditions, it can accurately identify the key point coordinates of the key parts of the video subject (such as the head, limbs, torso, etc.). At the same time, it preprocesses the audio data using spectral subtraction for noise reduction and finite impulse response (FIR) filter filtering, and uses short-time Fourier transform (STFT) to extract music frequency characteristics, such as fundamental frequency, harmonic frequency, and rhythm changes. In addition, the module has an adaptive resolution adjustment function, and can dynamically adjust the video resolution using the bilinear interpolation algorithm according to indicators such as the system CPU usage rate, memory occupancy, and video processing speed.

[0046] Furthermore, the action frequency analysis module described in Embodiment 1 of the present invention obtains the displacement change amount at adjacent moments through differential operations for the key point coordinate sequence in the labeled data, calculates the displacement amount per unit time in combination with the video frame rate, and uses the moving average filtering algorithm to smooth the displacement amount to remove outliers. It uses the fast Fourier transform (FFT) to convert the displacement amount time series to the frequency domain to obtain the action frequency distribution. After calculating the action frequencies for different parts of the object respectively, it classifies the similar action frequency patterns to construct an action frequency pattern library, uses the elbow method to determine the number of clusters, and introduces a time window sliding mechanism, sets a fixed-size time window, and slides the window over time to update the action frequency pattern library in real time.

[0047] Furthermore, the demand input module described in the embodiments of the present invention provides an intuitive user interaction interface, where users can input demands such as specific emotional atmospheres, theme styles, and lens type preferences. The module uses natural language processing technology to extract keywords and perform semantic analysis on the user input text, and encodes the analysis results into numerical features that can be processed by a computer.

[0048] Furthermore, the generation decision module described in the embodiments of the present invention constructs a joint frequency analysis model using a long short-term memory network (LSTM). The model consists of an input layer, multiple LSTM hidden layers, and an output layer. The input layer receives the normalized action frequency sequence, music frequency sequence, and user demand encoding information. During the training process, the stochastic gradient descent (SGD) algorithm is used to update the model weights and biases to learn the complex relationships among the three. When generating a lens generation strategy, according to preset rules and user demands, such as when the music rhythm speeds up, it matches the high-action frequency mode corresponding to the lens, and the lens selection is further optimized in combination with the user's emotional atmosphere and theme style demands; the same applies when the music is soothing. At the same time, fully considering the lens aesthetic rules, following the rule of thirds, symmetric composition method, etc., and combining the user's lens type preference, the final lens scheme is determined by calculating the matching degree of the candidate lens with the aesthetic rules and user demands.

[0049] Furthermore, the generation execution module described in the first embodiment of the present invention accurately clips the required lenses from the original action segments and combines them reasonably according to the lens generation strategy. The video interpolation algorithm is used to adjust the video frame rate and smooth the picture after clipping. When processing the selected lenses, according to the music emotion and user demands, by adjusting the HSV space parameters of the image and adding transition effects such as fade-in / fade-out, rotation switching, and zoom transition, an animated lens sequence is generated.

[0050] The object key point detection algorithm described in the first embodiment of the present invention is essentially an algorithm that uses deep learning technology to accurately locate the coordinates of the key parts of the video subject. Compared with the traditional manual annotation method, it has the advantages of high efficiency, accuracy, and adaptability to complex scenes.

[0051] The joint frequency analysis model described in the first embodiment of the present invention is a model that integrates multiple data information, learns the associations between different data through a specific neural network structure to generate a lens generation strategy, and can be analogized to an intelligent decision-making engine that outputs the best lens generation scheme according to the input information.

[0052] The clustering algorithm in the action frequency analysis module described in the first embodiment of the present invention can also be implemented by hierarchical clustering algorithm, DBSCAN density clustering algorithm, etc. to classify similar action frequency patterns.

[0053] The adaptive resolution adjustment function of the data acquisition module described in Embodiment 1 of the present invention aims to balance data acquisition quality and system processing efficiency. When the original video resolution is too high, causing system lag, by adjusting the resolution, it can not only ensure the accuracy of key point detection but also enable the system to operate efficiently.

[0054] The time window sliding mechanism of the action frequency analysis module described in Embodiment 1 of the present invention can capture changes in the action frequency pattern in real time, ensuring that the action frequency pattern library always reflects the true action situation of the action subject in the animation and providing an accurate basis for shot generation decision-making.

[0055] The demand input module described in Embodiment 1 of the present invention allows users to input their expectations for animation shots into the system, enabling the final generated animation shots to meet user personalized needs and improving user satisfaction.

[0056] The joint frequency analysis model of the generation decision module described in Embodiment 1 of the present invention generates a scientific and reasonable shot generation strategy by learning the complex relationships among action frequency, music frequency, and user needs, ensuring a high degree of coordination and unity among the shots, actions, music, and user needs.

[0057] The image processing and transition special effect addition operations of the generation execution module described in Embodiment 1 of the present invention. Image processing can enhance the fit between the shot and the music emotion and user needs, such as adjusting colors to create a specific atmosphere; adding transition special effects makes the shot transition smoother and improves the overall visual effect of the animation.

[0058] The system described in Embodiment 1 of the present invention is characterized in that when the convolutional neural network model used in the object key point detection algorithm is trained, transfer learning technology is used, and the pre-trained model is a model trained on ImageNet.

[0059] The system described in Embodiment 1 of the present invention is characterized in that when the action frequency analysis module calculates the action frequency, for complex actions, wavelet transform is used to perform multi-resolution analysis on the displacement time series.

[0060] The system described in Embodiment 1 of the present invention is characterized in that the demand input module also supports voice input of demands and converts them into text information through speech recognition technology.

[0061] The system described in Embodiment 1 of the present invention is characterized in that when the generation decision module generates a shot generation strategy, it considers the influence of the subject characteristics on the shot performance.

[0062] Embodiment 2 of the present invention provides a system for generating shots by combining action and audio detection. The specific steps are as Figure 2 shown, and the following will be described in detail.

[0063] (1)Data collection and preprocessing. This method first uses the data collection module to comprehensively collect and preprocess the input action segment video and corresponding audio data. The data collection module uses a convolutional neural network model based on deep learning to annotate the key points of the video subject from the action segment video, generating annotation data containing the key point coordinate sequence. At the same time, collect the audio data of the corresponding segment, and perform preprocessing operations such as noise reduction and filtering, and use algorithms such as Fourier transform to extract the music frequency characteristics. These characteristics will provide an important data basis for subsequent shot generation. For example, when making a dance animation, collect the dance video segment of the dancer, and accurately identify the key point coordinates of each key part of the dancer's body through the convolutional neural network model. At the same time, collect the dance music, and extract its music frequency characteristics after preprocessing the audio.

[0064] (2)Action frequency analysis and pattern library construction. The frequency analysis module calculates the action trajectories and displacement amounts of different parts based on the changes in the key point coordinates of the object in the annotation data. Perform differential operations on each key point coordinate sequence to obtain the displacement, combine the frame rate to calculate the displacement amount per unit time, and convert the displacement amount time series to the frequency domain through Fourier transform to obtain the action frequency distribution. Classify similar action frequency patterns through clustering algorithms to generate an action frequency pattern library. When constructing the action frequency pattern library, introduce a time window sliding mechanism to dynamically update the action frequency pattern to adapt to the changes of actions over time, and improve the timeliness and accuracy of the action frequency pattern library.

[0065] (3)Shot generation strategy formulation. The generation decision module integrates the information of the action frequency pattern library, the music frequency characteristics, and the user input requirements to establish a joint frequency analysis model. Use the long short-term memory network, take the action frequency sequence, the music frequency sequence, and the user requirement encoded information as inputs, and through training to learn the time synchronization relationship and emotional association among the three, and output the predicted values of the shot generation parameters. According to the preset rules and user requirements, such as matching the shots corresponding to the high action frequency pattern when the music rhythm speeds up and matching the shots corresponding to the low action frequency pattern when the music is soothing, and at the same time combining the user's requirements for the emotional atmosphere and theme style, generate a shot generation strategy. When generating the shot generation strategy, consider the shot aesthetics rules, such as following the rule of thirds, symmetric composition method, etc., and at the same time combining the user's requirements for the shot type preference, determine the parameters such as the shot switching point, shot size, and camera movement method.

[0066] (4)Lens generation and processing. The generation execution module clips and combines lenses from the original segments according to the lens generation strategy. Image processing is performed on the selected lenses, such as adjusting color and contrast to match the music emotion and user requirements, and adding transition effects, and finally outputs an animated lens sequence that conforms to the music, action frequency, and user requirements. For example, when the music rhythm is exciting, select the fast-switching close-up lenses corresponding to the high-action frequency mode, and increase the color saturation and contrast of the picture; when the music is soothing, select the panoramic or medium-shot lenses corresponding to the low-action frequency mode, and adjust the picture tone to be softer.

[0067] Embodiment 3 of the present invention provides an electronic device, which mainly consists of a processor, a memory, and a computer program stored in the memory and executable on the processor.

[0068] When the processor executes the computer program, it can implement the above-mentioned animated lens generation method. In practical applications, the processor can select a high-performance multi-core CPU, and its powerful computing power can quickly and efficiently process complex computing tasks, ensuring that the system can generate animated lenses in real time and accurately.

[0069] The memory can adopt a high-speed solid-state drive, which has the characteristics of fast reading and writing, and can quickly store and read a large amount of data such as action segment videos, audio data, feature data, and lens generation strategies, providing strong support for the efficient operation of the system.

[0070] Embodiment 4 of the present invention provides a computer-readable storage medium for storing a computer program, which can implement the above-mentioned animated lens generation method when the program is executed by a processor.

[0071] Common computer-readable storage media include optical discs, USB flash drives, hard disks, etc. These storage media have the characteristics of large capacity and stable storage, which are convenient for users to store and transfer program files. The user only needs to connect the medium storing the program to the electronic device, and the processor can read the program therein and execute the corresponding method steps to implement the dynamic generation function of the animated lens.

[0072] The present invention deeply analyzes the action and audio data, combines the user requirements, and uses advanced algorithms and models to generate a highly coordinated and unified animated lens sequence. Compared with the traditional animation production method, it greatly improves the animation production efficiency, enhances the fit between the lens and the content, and brings more creativity and possibilities to animation creation.

Claims

1. A system for generating a shot by combining motion and audio detection, characterized in that, Including: A data acquisition module, which is used to obtain video data of different action segments, use an object key point detection algorithm to label the key points of the video subject, generate annotation data containing the key point coordinate sequence, and at the same time collect the audio data of the corresponding segment and perform preprocessing to extract the music frequency characteristics; A frequency analysis module, based on the changes of the object key point coordinates in the annotation data, calculates the action trajectories and displacement amounts of different parts, combines the time information to obtain the action frequencies of each part, and classifies the similar action frequency patterns through a clustering algorithm to generate an action frequency pattern library; A requirement input module, which is used to receive the generation requirements of the lens input by the user, and the requirements include but are not limited to specific emotional atmospheres, theme styles, lens type preferences, etc.; A generation decision module, which integrates the information of the action frequency pattern library, the music frequency characteristics and the requirements input by the user, establishes a joint frequency analysis model, and according to the preset rules and user requirements, such as matching the high-action frequency pattern corresponding lens when the music rhythm speeds up and matching the low-action frequency pattern corresponding lens when the music is soothing, and at the same time combines the user's requirements for the emotional atmosphere and theme style, generates a lens generation strategy, and determines parameters such as lens switching points, shot sizes, and camera movement methods; A generation execution module, according to the lens generation strategy, clips and combines the lenses from the original segments, performs image processing on the selected lenses, such as adjusting the color and contrast to match the music emotion and user requirements, adds transition effects, and finally outputs an animated lens sequence that meets the music, action frequency, and user requirements.

2. The system according to claim 1, wherein The object key point detection algorithm uses a convolutional neural network model based on deep learning. Through a large amount of data training, it can accurately identify the key point coordinates of the key parts.

3. The system according to claim 1, wherein When the frequency analysis module calculates the action frequency, it performs a difference operation on each key point coordinate sequence to obtain the displacement, combines the frame rate to calculate the displacement amount per unit time, and converts the displacement amount time series to the frequency domain through Fourier transform to obtain the action frequency distribution.

4. The system according to claim 1, wherein The joint frequency analysis model uses a long short-term memory network, takes the action frequency sequence, the music frequency sequence, and the user requirement coding information as inputs, and through training, learns the time synchronization relationship and emotional association among the three, and outputs the predicted values of the lens generation parameters.

5. The system according to claim 1, wherein When the data acquisition module obtains the video data of the action segment, it has an adaptive resolution adjustment function, which can automatically adjust the video resolution to the optimal processing resolution according to the original video resolution and the system processing ability, ensuring the balance between data acquisition quality and processing efficiency.

6. The system according to claim 1, wherein When the frequency analysis module constructs the action frequency pattern library, it introduces a time window sliding mechanism to dynamically update the action frequency pattern to adapt to the changes of the action over time, improving the timeliness and accuracy of the action frequency pattern library.

7. The system according to claim 1, characterized in that, When the generation decision module generates the lens generation strategy, it considers the lens aesthetic rules, such as following the rule of thirds, symmetric composition method, etc., and at the same time combines the user's requirements for the lens type preference, making the generated lens more aesthetically pleasing visually and meeting the user's expectations.

8. A lens generation method based on the system according to any one of claims 1-7, characterized in that, It includes the steps of: collecting different segment videos and corresponding audio data, and performing key point detection and annotation on objects; analyzing the changes in the coordinates of key points of different parts in the annotated data, calculating the action frequency, and constructing an action frequency pattern library; receiving the generation requirements of the user for the shot; extracting the music frequency features of the audio data, fusing the information of the action frequency pattern library, the music frequency features and the user requirements, and generating a shot generation strategy through a joint frequency analysis model; according to the shot generation strategy, editing and processing the segments to generate a shot sequence that meets the music, action frequency and user requirements.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method according to claim 8.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to claim 8.

Citation Information

Cited By

  • Data processing method and device

    CN120935380A