Sound and action synchronous optimization system and method thereof

Through modular design and multimodal data fusion technology, combined with deep learning and real-time feedback optimization, the accuracy and adaptability of the sound and action synchronization system in complex environments is solved, and high-precision and adaptive synchronization optimization effect is achieved, improving performance quality and user experience.

CN120296664APending Publication Date: 2025-07-11BEIJING FILM ACAD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510385656.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing sound and action synchronization systems are difficult to achieve high-precision synchronization in complex performance environments, especially in the face of dynamic changes and nonlinear relationships, and most systems lack real-time adaptability and multimodal data processing capabilities.

Method used

The modular system design is adopted, combined with multi-modal data fusion, deep learning and real-time feedback optimization technology, through the modular processing of sound and action signal acquisition, analysis, execution and rendering units, multi-scale cross-correlation analysis, adaptive delay compensation algorithm and deep reinforcement learning are used to achieve accurate synchronization and adaptive optimization of sound and action.

Benefits of technology

It realizes high-precision, adaptive sound and action synchronization, can adapt to complex performance scenarios, improve performance quality and audience experience, and supports applications in fields such as virtual reality and augmented reality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296664A_ABST
    Figure CN120296664A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of music systems, in particular to a sound and action synchronous optimization system and method. The analysis processing unit is electrically connected with the sound and action signal acquisition unit; the performance action execution unit is electrically connected with the analysis processing unit; driving the picture rendering unit to combine the synchronization adjustment signal with the picture; the picture rendering unit is electrically connected with the performance action execution unit; the user interaction unit is electrically connected with the analysis processing unit, the performance action execution unit and the picture rendering unit, and is used for receiving operation information input by a user; performing parameter setting on the sound and action synchronous optimization system; according to the method, the synchronous optimization effect is fed back to the user, high-precision, self-adaptive and multi-mode synchronous optimization is achieved through innovative architecture design and algorithm optimization, the quality of various performances and audience experience can be remarkably improved, and powerful technical support is provided for emerging fields such as virtual reality and augmented reality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of music systems, in particular to a system and method for optimizing the synchronization of sound and motion. Background Art

[0002] In the fields of contemporary performing arts, virtual reality, and human-computer interaction, the precise synchronization of sound and motion has always been a highly concerned and challenging technical problem. With the continuous advancement of technology, various sound and motion synchronization systems have emerged, but there are still many problems that need to be solved.

[0003] In the prior art, the most common synchronization method is simple alignment based on timestamps. Although this method is simple to implement, it is difficult to cope with complex performance environments. For example, in live music performances, there is a natural time difference between the sound propagation speed of musical instruments and the movements of performers on stage, and simple timestamp alignment cannot accurately capture this subtle time difference. In addition, this method cannot handle the nonlinear correspondence between sound and movement. For example, in some dance performances, the changes in movement may slightly lead or lag behind the changes in the rhythm of the music.

[0004] Another common method is to use cross-correlation analysis to find the best match between sound and motion signals. This method performs well when dealing with highly periodic signals, but it often fails when faced with complex, non-periodic performances. For example, in modern dance or improvisation, the relationship between movement and music may be highly abstract and non-linear, and simple correlation analysis is difficult to accurately capture this complex correspondence.

[0005] In recent years, with the development of machine learning technology, some researchers have tried to use deep learning methods to solve the synchronization problem. These methods usually learn the mapping relationship between sound and action through a large amount of training data. Although good results have been achieved in certain specific scenarios, such methods generally have the problem of insufficient generalization ability. When faced with a completely new performance style or an unseen action sequence, the performance of the system is often greatly reduced. In addition, these methods usually require a lot of computing resources and are difficult to apply in live performances with high real-time requirements.

[0006] Existing technologies also face a common challenge, which is that they are difficult to adapt to dynamically changing performance environments. For example, in a concert that lasts for several hours, the status of the performers may change over time, and environmental factors (such as temperature and humidity) may also affect the propagation characteristics of sound. Most existing systems lack real-time adaptive capabilities and cannot make timely adjustments to these dynamic changes.

[0007] In addition, existing synchronization systems generally suffer from the problem of single-modal limitations. They either focus only on sound signals or only on action signals, and few systems can comprehensively utilize information from multiple perceptual modalities. This single-modal processing method is difficult to fully capture the rich connotations of performances, thus affecting the accuracy and robustness of synchronization.

[0008] In view of the above problems, there is an urgent need for a sound and action synchronization optimization system that can achieve high precision and adaptability to meet the urgent needs in the fields of modern performing arts and human-computer interaction. Summary of the Invention

[0009] The present invention aims to solve the above technical problems and provides a sound and action synchronization optimization system and method, which can achieve precise synchronization of sound and action and have real-time adaptability and multi-modal data processing capabilities.

[0010] The present invention proposes a sound and action synchronization optimization system, including:

[0011] A sound and action signal acquisition unit, used for:

[0012] Acquiring the sound signals and action signals emitted by the performer;

[0013] Converting the sound signals and action signals into digital signals;

[0014] An analysis and processing unit, electrically connected to the sound and action signal acquisition unit, used for:

[0015] Receiving the digitized sound signals and action signals sent by the sound and action signal acquisition unit;

[0016] Performing feature extraction and synchronization analysis on the digitized sound signals and action signals;

[0017] Generating a synchronization adjustment signal executable by the performance action execution unit;

[0018] A performance action execution unit, electrically connected to the analysis and processing unit, used for:

[0019] Receiving the synchronization adjustment signal sent by the analysis and processing unit;

[0020] Executing the synchronization adjustment signal;

[0021] Driving the screen rendering unit to combine the synchronization adjustment signal with the screen;

[0022] A screen rendering unit, electrically connected to the performance action execution unit, used for:

[0023] Receiving the driving signal from the performance action execution unit;

[0024] Render the screen corresponding to the synchronization adjustment signal in real time;

[0025] Send the rendered screen to the user;

[0026] The user interaction unit is electrically connected to the analysis and processing unit, the performance action execution unit, and the screen rendering unit, and is used for:

[0027] Receive the operation information input by the user;

[0028] Set parameters for the sound and action synchronization optimization system;

[0029] Feedback the synchronization optimization effect to the user.

[0030] Preferably, the sound and action signal acquisition unit includes:

[0031] The sound processing module is used for:

[0032] Collect the sound signals emitted by the performer;

[0033] Convert the sound signals into digital signals to obtain A / D audio signals;

[0034] The action processing module is used for:

[0035] Collect the action signals emitted by the performer;

[0036] Convert the action signals into digital signals to obtain action signals;

[0037] Among them, the sound processing module and the action processing module are connected through a data bus to achieve synchronous data transmission.

[0038] Preferably, the analysis and processing unit includes:

[0039] The synchronization adjustment module is used for:

[0040] Receive the A / D audio signals and action signals;

[0041] Process the A / D audio signals and action signals;

[0042] Generate executable synchronization adjustment signals;

[0043] The action processing module is used for:

[0044] Analyze the action signals;

[0045] Generate executable action signals;

[0046] The output signal generation module is electrically connected to the synchronization adjustment module and the action processing module, and is used for:

[0047] Receive the synchronization adjustment signal and the executable action signal;

[0048] Generate a final executable signal.

[0049] Preferably, the synchronization adjustment module includes:

[0050] A synchronization analysis unit for:

[0051] Analyze the synchronization relationship between the A / D audio signal and the action signal;

[0052] Calculate the synchronization adjustment coefficient;

[0053] A synchronization adjustment unit, electrically connected to the synchronization analysis unit, for:

[0054] Generate an A / D audio delay adjustment coefficient according to the calculation result of the synchronization analysis unit;

[0055] Generate a music enhancement coefficient;

[0056] Achieve the synchronization of the A / D audio signal and the action signal.

[0057] Preferably, the sound and action signal acquisition unit further includes a preprocessing module, and the preprocessing module is used for:

[0058] Perform noise reduction processing on the A / D audio signal;

[0059] Perform filtering processing on the action signal;

[0060] Wherein, the preprocessing module is arranged between the sound processing module and the synchronization adjustment module, and the preprocessing module is also electrically connected to the synchronization adjustment module and the output signal generation module, for:

[0061] Transmit the denoised sound signal to the synchronization adjustment module and the output signal generation module;

[0062] Transmit the filtered action signal to the synchronization adjustment module and the output signal generation module.

[0063] Preferably, the signal denoising module, the filtering module, the synchronization analysis unit, the synchronization adjustment unit, and the output signal generation module are electrically connected to the user interaction unit;

[0064] The user interaction unit includes a tuning signal module, and the tuning signal module is used for:

[0065] Send adjustment parameters to the signal denoising module, the filtering module, the synchronization analysis unit, the synchronization adjustment unit, and the output signal generation module;

[0066] Tune their parameters;

[0067] Among them, the adjustment parameters include denoising parameters, filtering parameters, synchronization evaluation parameters, and output control parameters, and are specifically used for:

[0068] The preprocessing module performs denoising processing on the A / D audio signal according to the denoising parameters;

[0069] The preprocessing module performs filtering processing on the action signal according to the filtering parameters;

[0070] The synchronization analysis unit analyzes the synchronization relationship between the A / D audio signal and the action signal according to the synchronization evaluation parameters to calculate the synchronization adjustment coefficient;

[0071] The synchronization adjustment unit generates an A / D audio delay adjustment coefficient and a music enhancement coefficient according to the synchronization adjustment coefficient;

[0072] The output signal generation module splices the A / D audio signal and the action signal according to the output control parameters to generate an executable synchronization adjustment signal.

[0073] Preferably, it further includes an artificial intelligence learning module, which is electrically connected to the analysis and processing unit and is used for:

[0074] Receiving the historical synchronization analysis data of the analysis and processing unit;

[0075] Training the historical data based on machine learning algorithms;

[0076] Generating an optimized synchronization prediction model;

[0077] Feeding back the optimized synchronization prediction model to the analysis and processing unit to improve the accuracy and efficiency of synchronization analysis.

[0078] Preferably, it further includes a multi-modal data fusion module, which is electrically connected to the sound and action signal acquisition unit and the analysis and processing unit and is used for:

[0079] Receiving various modal data from the sound and action signal acquisition unit, including audio data, video data, motion capture data, and physiological signal data;

[0080] Performing time alignment and feature fusion on the various modal data;

[0081] Generating a fused multi-modal feature vector;

[0082] Transmitting the multi-modal feature vector to the analysis and processing unit to improve the comprehensiveness and robustness of synchronization analysis.

[0083] Preferably, it further includes a real-time feedback optimization module, which is electrically connected to the performance action execution unit, the screen rendering unit and the user interaction unit, and is used for:

[0084] Real-time monitor the execution effect of the performance action execution unit and the rendering effect of the screen rendering unit;

[0085] Receive real-time feedback information from the user interaction unit;

[0086] Based on the execution effect, rendering effect and user feedback information, calculate real-time optimization parameters;

[0087] Feed back the real-time optimization parameters to the analysis and processing unit for dynamically adjusting the synchronization optimization strategy.

[0088] A sound and action synchronization optimization method, applied to the sound and action synchronization optimization system described above, includes the following steps:

[0089] S1. Collect the sound signal and action signal emitted by the performer;

[0090] S2. Convert the sound signal and action signal into digital signals;

[0091] S3. Extract features and perform synchronization analysis on the digitized sound signal and action signal;

[0092] S4. Generate a synchronization adjustment signal executable by the performance action execution unit;

[0093] S5. Execute the synchronization adjustment signal and drive the screen rendering unit;

[0094] S6. Real-time render the screen corresponding to the synchronization adjustment signal and send the rendered screen to the user;

[0095] S7. Receive the operation information input by the user and set the parameters of the sound and action synchronization optimization system;

[0096] S8. Use the artificial intelligence learning module to train the historical synchronization analysis data to generate an optimized synchronization prediction model;

[0097] S9. Through the multi-modal data fusion module, perform time alignment and feature fusion on multiple modal data to generate a multi-modal feature vector;

[0098] S10. Use the real-time feedback optimization module to dynamically adjust the synchronization optimization strategy to achieve continuous optimized synchronization of sound and action.

[0099] The beneficial effects of the present invention are mainly reflected in the following aspects:

[0100] The sound and motion synchronization optimization system of the present invention has achieved significant technological breakthroughs in multiple aspects through innovative system architecture and algorithm design:

[0101] First of all, the present invention adopts a modular system design, clearly dividing functions such as the acquisition, analysis, execution, and rendering of sound and motion. This architecture not only improves the maintainability and scalability of the system but also provides flexibility for the optimization of each module. For example, the modular design of the sound and motion signal acquisition unit enables the system to easily adapt to different types of sensors, while the independence of the analysis and processing unit allows researchers to continuously improve the algorithm without modifying other parts.

[0102] Secondly, the system of the present invention breaks through the limitations of traditional single-modal processing by introducing multi-modal data fusion technology. By simultaneously analyzing various data such as audio, video, motion capture, and physiological signals, the system can more comprehensively understand the internal structure and external performance of the performance. This multi-modal fusion not only improves the synchronization accuracy but also enhances the system's adaptability to complex performance scenarios. For example, in an environment with insufficient light, the system can rely more on audio and motion capture data; while in a noisy scenario, video and physiological signals may play a more important role.

[0103] Furthermore, the system of the present invention introduces an artificial intelligence learning module based on deep reinforcement learning, enabling the system to have the ability of continuous learning and optimization. This design allows the system to continuously accumulate experience and gradually improve its understanding of different performance styles and scenarios. As the usage time increases, the performance of the system will continuously improve, which is especially valuable for professional performance venues with long-term use.

[0104] In addition, the real-time feedback optimization module of the present invention provides the system with dynamic adaptive capabilities. By real-time monitoring the execution effect and user feedback, the system can continuously adjust the synchronization strategy during the performance. This feature enables the system to cope with sudden changes during the performance, such as when the performer temporarily changes the rhythm or movement, thus maintaining a continuous high-quality synchronization effect.

[0105] Finally, the system of the present invention has made multiple innovations at the algorithm level, such as using multi-scale cross-correlation analysis for synchronization analysis, using an adaptive delay compensation algorithm to generate synchronization adjustment signals, and using attention mechanisms and graph neural networks for feature fusion. These algorithm innovations not only improve the accuracy and efficiency of the system but also enhance the system's ability to handle complex non-linear relationships, enabling it to cope with various complex performance scenarios.

[0106] In summary, through innovative architecture design and algorithm optimization, the sound and motion synchronization optimization system of the present invention achieves high-precision, adaptive, and multi-modal synchronization optimization. This system can not only significantly improve the quality of various performances and the audience experience, but also provide strong technical support for emerging fields such as virtual reality and augmented reality, with broad application prospects and important practical value. BRIEF DESCRIPTION OF THE DRAWINGS

[0107] Figure 1 It is a logic block diagram of the overall system of the present invention.

[0108] Figure 2 It is a logic block diagram of the sound and motion signal acquisition unit of the present invention.

[0109] Figure 3 It is a logic block diagram of the analysis and processing unit of the present invention.

[0110] Figure 4 It is a logic block diagram of the user interaction unit of the present invention.

[0111] Figure 5 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0112] Referring to Figures 1-5 , the present invention provides a sound and motion synchronization optimization system and its method, aiming to solve the problem of inaccurate sound and motion synchronization in the prior art. The following will describe the detailed implementation of the present invention.

[0113] The sound and motion synchronization optimization system of the present invention includes a sound and motion signal acquisition unit 1, an analysis and processing unit 2, a performance action execution unit 3, a screen rendering unit 4, and a user interaction unit 5.

[0114] The sound and motion signal acquisition unit 1 is used to collect the sound signals and motion signals emitted by the performer and convert these signals into digital signals. Preferably, this unit can include a high-sensitivity microphone array and a high-precision motion capture device to ensure that the collected signals have a high signal-to-noise ratio and accuracy. For example, the microphone array can use 8 omnidirectional microphones with a sampling rate set at 48 kHz to capture all-round sound information. The motion capture device can use an optical marker point tracking system with a sampling rate set at 120 Hz to achieve millimeter-level motion accuracy tracking.

[0115] The analysis and processing unit 2 is electrically connected to the sound and motion signal acquisition unit 1, and is used to receive the digitized sound signal and motion signal, extract features and perform synchronous analysis on these signals, and generate a synchronous adjustment signal executable by the performance action execution unit 3. In an embodiment of the present invention, the analysis and processing unit 2 adopts an innovative multimodal feature fusion algorithm, which combines deep learning and traditional signal processing technologies and can effectively extract the time-frequency features of sound and motion signals. Specifically, this algorithm first uses the short-time Fourier transform (STFT) to extract the time-frequency features of the sound signal, and then uses a convolutional neural network (CNN) to extract the spatio-temporal features of the motion signal. Finally, these two types of features are fused through an attention mechanism to obtain a more robust synchronous feature representation.

[0116] This multimodal feature fusion algorithm can be expressed as:

[0117] F = Attention(CNN(M), STFT(S)),

[0118] where F is the fused feature, M is the motion signal, S is the sound signal, CNN and STFT respectively represent the convolutional neural network and short-time Fourier transform operations, and Attention represents the attention mechanism.

[0119] The performance action execution unit 3 is electrically connected to the analysis and processing unit 2, and is used to receive and execute the synchronous adjustment signal, and at the same time drive the screen rendering unit 4. The system of the present invention adopts a high-precision servo motor and a hydraulic drive system, which can achieve a response time of the microsecond level, thereby ensuring the accuracy and real-time performance of action execution. For example, in a specific embodiment, the response time of the servo motor can reach 0.5 ms, and the torque accuracy can be controlled within 0.01 Nm, which is crucial for achieving fine action control.

[0120] The screen rendering unit 4 is electrically connected to the performance action execution unit 3, and is used to render the screen corresponding to the synchronous adjustment signal in real time and send the rendered screen to the user. The present invention adopts advanced real-time rendering technology, combined with GPU acceleration and optimized rendering algorithms, which can control the rendering delay within 16.7 ms (i.e., a refresh rate of 60 fps) while ensuring the screen quality, thereby providing a smooth visual experience for the user.

[0121] The user interaction unit 5 is electrically connected to the analysis and processing unit 2, the performance action execution unit 3, and the screen rendering unit 4, and is used to receive the operation information input by the user, set parameters for the sound and action synchronization optimization system, and feedback the synchronization optimization effect to the user. An innovation of the present invention is that the user interaction unit 5 adopts an adaptive user interface technology, which can dynamically adjust the interface layout and function display according to the user's usage habits and operation frequencies, thereby improving the user's operation efficiency.

[0122] Further, the sound and action signal acquisition unit 1 of the present invention includes a sound processing module 11 and an action processing module 12.

[0123] The sound processing module 11 is used to collect the sound signals emitted by the performer and convert the sound signals into digital signals to obtain A / D audio signals. The present invention adopts a high-precision 24-bit analog-to-digital converter with a sampling rate of up to 192 kHz and an effective number of bits reaching 20 bit, which ensures the high-fidelity conversion of sound signals. At the same time, in order to cope with different application scenarios, this system also supports dynamically adjusting the sampling rate and bit depth. For example, on resource-constrained mobile devices, the sampling rate can be reduced to 48 kHz and the bit depth can be reduced to 16 bit to balance the sound quality and computing resources.

[0124] The action processing module 12 is used to collect the action signals emitted by the performer and convert the action signals into digital signals to obtain action signals. The present invention uses a method that combines a high-precision inertial measurement unit (IMU) and an optical motion capture system. The sampling rate of the IMU can reach 1000 Hz, and the angular accuracy can reach 0.1 degree, while the spatial resolution of the optical system can reach the sub-millimeter level. This dual capture mechanism can effectively reduce the errors caused by factors such as occlusion and magnetic field interference, and improve the accuracy and stability of motion capture.

[0125] It should be noted that the sound processing module 11 and the action processing module 12 are connected through a high-speed data bus to achieve synchronous data transmission. The present invention adopts the PCI Express 4.0 protocol with a two-way bandwidth of up to 64 GB / s, which ensures the real-time transmission of a large amount of audio and action data and lays a foundation for subsequent synchronous analysis.

[0126] The analysis and processing unit 2 of the present invention includes a synchronization adjustment module 21, an action processing module 22, and an output signal generation module 23.

[0127] The synchronization adjustment module 21 is used to receive the A / D audio signal and the action signal, process these signals, and generate an executable synchronization adjustment signal. This module adopts an innovative time series alignment algorithm, which combines dynamic time warping (DTW) and long short-term memory network (LSTM), and can effectively handle the non-linear time correspondence relationship between the sound and action signals. Specifically, the algorithm can be expressed as:

[0128] A = LSTM(DTW(S, M)),

[0129] where A is the aligned sequence, S is the sound signal sequence, M is the action signal sequence, DTW represents the dynamic time warping operation, and LSTM represents the long short-term memory network.

[0130] The action processing module 22 is used to analyze the action signal and generate an executable action signal. This module adopts an action optimization algorithm based on deep reinforcement learning, which can generate an optimal action sequence according to the current action state and sound signal. The core of this algorithm is a policy network π θ , and its goal is to maximize the expected reward:

[0131]

[0132] where θ is the network parameter, τ is the action trajectory, R(τ) is the reward function, which measures the synchronization degree between the action and the sound.

[0133] The output signal generation module 23 is electrically connected to the synchronization adjustment module 21 and the action processing module 22, and is used to receive the synchronization adjustment signal and the executable action signal, and generate a final executable signal. This module adopts an adaptive weighted fusion algorithm, which can dynamically adjust the weights of the sound and action signals according to different scenarios, so as to achieve the optimal synchronization effect. The algorithm can be expressed as:

[0134] O = w s S + w m M,

[0135] where O is the output signal, S is the sound signal after synchronization adjustment, M is the optimized action signal, w s and w m are the weights of the sound and action signals respectively, and satisfy w s + w m = 1. The values of the weights are calculated in real time through a small neural network to adapt to different performance scenarios.

[0136] Through the above innovative algorithms and module designs, the sound and motion synchronization optimization system of the present invention can achieve a synchronization accuracy at the millisecond level, greatly enhancing the realism and immersion of the performance. At the same time, the adaptive characteristics of the system enable it to adapt to various complex performance environments, providing users with stable and reliable synchronization optimization services. Continuing to describe the technical solution of the present invention in depth,

[0137] The synchronization adjustment module 21 includes a synchronization analysis unit 211 and a synchronization adjustment unit 212.

[0138] The synchronization analysis unit 211 is used to analyze the synchronization relationship between the A / D audio signal and the motion signal, and calculate the synchronization adjustment coefficient. In a preferred embodiment of the present invention, the synchronization analysis unit 211 adopts an innovative multi-scale cross-correlation analysis algorithm. This algorithm first performs wavelet transforms on the audio signal and the motion signal to obtain time-frequency representations at different scales, and then calculates the cross-correlation function between these representations. This method can effectively capture the synchronization relationship at different time scales, thereby improving the accuracy and robustness of the analysis. Specifically, the calculation formula for the synchronization adjustment coefficient λ is as follows:

[0139]

[0140] where, and respectively represent the wavelet coefficients of the audio signal and the motion signal at the i-th scale, corr represents the cross-correlation function, w i is the weight of different scales, and N is the total number of scales. Preferably, the present invention uses 5 scales for analysis, and the weight of each scale is dynamically adjusted according to its corresponding frequency range. For example, for a performance mainly featuring vocals, the weight of the mid-frequency band (such as 500 Hz - 2 kHz) will increase accordingly.

[0141] The synchronization adjustment unit 212 is electrically connected to the synchronization analysis unit 211, and is used to generate an A / D audio delay adjustment coefficient and a music enhancement coefficient according to the calculation result of the synchronization analysis unit 211 to achieve the synchronization of the A / D audio signal and the motion signal. The system of the present invention adopts an adaptive delay compensation algorithm, which can dynamically adjust the audio delay according to the current synchronization state. The calculation formula for the A / D audio delay adjustment coefficient τ is as follows:

[0142] τ = τ0 + k·(λ - λ0),

[0143] where, τ0 is the base delay, λ0 is the synchronization adjustment coefficient under the ideal synchronization state, and k is the adjustment coefficient. Preferably, τ0 is set to 10 ms, λ0 is set to 0.95, and the value range of k is 0.1 - 1.0. The specific value is dynamically optimized by a machine learning algorithm according to historical data.

[0144] The music enhancement coefficient α is used to adjust the volume and timbre of the audio signal to better match the actions. Its calculation formula is as follows:

[0145] α = 1 + β·(1 - λ),

[0146] where β is the enhancement intensity parameter, and its value range is 0.1 - 0.5. When the synchronization degree is low, the system will appropriately enhance the audio signal to guide the synchronization of the actions.

[0147] Furthermore, the sound and motion signal acquisition unit 1 of the present invention further includes a preprocessing module 13. The preprocessing module 13 is used to preprocess the A / D audio signal and the motion signal, and it includes a signal denoising module 131 and a filtering module 132.

[0148] The preprocessing module 13 is arranged between the sound processing module 11 and the synchronization adjustment module 21, and is electrically connected to the synchronization adjustment module 21 and the output signal generation module 23 at the same time. This design can ensure that the signal is fully preprocessed before entering the core processing unit, improving the efficiency and accuracy of subsequent processing.

[0149] The signal denoising module 131 is used to denoise the A / D audio signal. In an embodiment of the present invention, a deep learning-based autoencoder denoising algorithm is adopted. This algorithm first inputs the noisy audio signal into a multi-layer convolutional neural network, maps the signal to a low-dimensional latent space through a non-linear transformation, and then reconstructs the denoised signal through a deconvolution operation. This method can effectively remove complex background noise while retaining the detailed features of the signal. The quality of the denoised signal is evaluated using the signal-to-noise ratio (SNR) index. Preferably, the SNR improvement threshold of this system is set to 6dB, that is, only when the SNR of the denoised signal is increased by more than 6dB compared to the original signal, the denoising result will be adopted.

[0150] The filtering module 132 is used to filter the motion signal. Considering that the motion signal usually contains high-frequency noise and low-frequency drift, the present invention adopts an adaptive band-pass filter. The passband range of this filter is dynamically adjusted according to the type of action. For example, for fast actions (such as dancing), the passband is set to 2 - 20Hz; for slow actions (such as Tai Chi), the passband can be adjusted to 0.5 - 5Hz. The order of the filter is automatically selected through the minimum description length (MDL) criterion to balance the filtering effect and computational complexity.

[0151] Preferably, the preprocessing module 13 further includes a signal quality evaluation unit (not shown in the figure) for real-time monitoring of the quality of the preprocessed signal. When the signal quality is lower than the preset threshold (for example, the signal-to-noise ratio is less than 15dB or the average jitter amplitude of the motion signal exceeds 1cm), the system will automatically adjust the acquisition parameters or trigger re-acquisition to ensure the input quality of subsequent processing.

[0152] The user interaction unit 5 of the present invention includes a tuning signal module 51. The tuning signal module 51 is used to send adjustment parameters to the signal denoising module 131, the filtering module 132, the synchronization analysis unit 211, the synchronization adjustment unit 212, and the output signal generation module 23 to adjust their parameters. This design makes the system highly flexible and adjustable, capable of adapting to different application scenarios and user requirements.

[0153] The adjustment parameters include denoising parameters, filtering parameters, synchronization evaluation parameters, and output control parameters.

[0154] The specific functions of these parameters are as follows:

[0155] 1. Denoising parameters: Used to control the denoising intensity and method of the signal denoising module 131. For example, the number of layers (3 - 7 layers) of the autoencoder network and the number of neurons in each layer (64 - 512) can be adjusted, as well as the denoising threshold (0.1 - 0.5).

[0156] 2. Filtering parameters: Used to adjust the filtering characteristics of the filtering module 132. Include filter type (such as Butterworth, Chebyshev, etc.), cut-off frequency (can be adjusted in the range of 0.1 - 50 Hz), and filter order (1 - 10 orders).

[0157] 3. Synchronization evaluation parameters: Used to control the analysis method and sensitivity of the synchronization analysis unit 211. For example, the number of scales (3 - 7) of wavelet transform, the time window size (0.5 - 5 seconds) of cross-correlation analysis, and the synchronization threshold (0.8 - 0.99) can be adjusted.

[0158] 4. Output control parameters: Used to adjust the signal synthesis strategy of the output signal generation module 23. Include the mixing ratio (0.1 - 0.9) of audio and action signals, the smoothing coefficient (0.1 - 0.9) of the output signal, etc.

[0159] According to the system analysis results, sound prompts are issued through a speaker or headphones. For example: When the deviation between the action and the music rhythm exceeds 50 milliseconds, a voice reminder "The rhythm is a bit slow, you can keep up closer!" will be heard. When the emotional expression is particularly in place, applause or encouraging sound effects will be played. These prompts are synchronized with the progress bar and score box in the picture to help users adjust quickly.

[0160] The synchronization error (such as DTW distance) between the sound and the action is calculated by the synchronization analysis unit 211, and prompts such as "Rhythm synchronization: 90%" or "Need to strengthen following the beat" are displayed in real time. Combined with the audio analysis of the sound processing module 11, "Emotional investment: 85 points (the passion value of the current paragraph meets the standard)" or "Breath control: 75 points (pay attention to the uniformity of breathing)" are dynamically displayed.

[0161] The action scoring module (linked with the action processing module 22) will display the action completion degree in real time above the user interface, such as "Action standard: 92 points (jump height meets the standard, arm trajectory needs to be adjusted)"; physical expression, comprehensive action rhythm, muscle tension (through IMU data) and breathing synchronization (combined with the physiological signal module 14) give "physical expression: 88 points (rhythm matching is high, but muscles can be appropriately relaxed)"; expression scoring is through facial expression analysis captured by the picture rendering unit 4, showing "expression liveliness: 80 points (eyes can be more focused)".

[0162] After the performance, the user can review the performance through the "summary interface" of the user interaction unit 5, and the system will analyze it in detail like a coach: correlation analysis between physiological data and emotions, emotional consistency score: the multimodal data fusion module 7 will combine physiological signals (such as heart rate, blood pressure) with the performance rhythm to generate a report: "In the climax section at the 30th second, the heart rate increased by 15%, the blood pressure fluctuated by 6mmHg, and the emotional investment was sufficient, plus 10 points"; "In the lyrical section at the 45th second, the heart rate decreased but the expression was not synchronized, and the emotional consistency was deducted by 5 points", summary interface: total score display, comprehensive real-time score (action, emotion, rhythm) and physiological data correlation, showing "total score: 89 / 100"; key segment playback: when clicking "climax segment analysis", the system will play back the screen and mark "at this time, the heart rate peak matches, but the movement amplitude is insufficient, it is recommended to strengthen the body explosiveness"; personalized suggestions: based on the weak links, "in the slow rhythm section, you can enable the 'follow-up training mode' to assist in practice." The voice module of the real user interaction unit 5 directly calls the error data of the synchronous analysis unit 211 and the scoring results of the action processing module 22 to generate voice feedback. The physiological signal acquisition module (such as a heart rate belt, a blood pressure sensor) is connected in parallel with the motion capture device, and synchronously transmits data to the multimodal fusion module 7 through the PCIe4.0 bus to ensure comprehensive analysis. Next to the original screen rendering area, a new scoring bar and voice prompt area are added, and users can switch between "real-time mode" and "replay mode" with one click.

[0163] The system of the present invention dynamically adjusts the working mode of each module according to these parameters. For example, the preprocessing module 13 performs denoising on the A / D audio signal according to the denoising parameters, and performs filtering on the action signal according to the filtering parameters. The synchronization analysis unit 211 performs synchronization relationship analysis on the A / D audio signal and the action signal according to the synchronization evaluation parameters to calculate the synchronization adjustment coefficient. The synchronization adjustment unit 212 generates the A / D audio delay adjustment coefficient and the music enhancement coefficient according to the calculated synchronization adjustment coefficient. Finally, the output signal generation module 23 splices the A / D audio signal and the action signal according to the output control parameters to generate an executable synchronization adjustment signal.

[0164] Preferably, the system of the present invention further includes a parameter optimization unit (not shown in the figure), which uses the Bayesian optimization algorithm to automatically adjust the above parameters according to historical performance data to achieve continuous optimization of the system performance. The optimization objective function considers multiple factors such as synchronization accuracy, computing efficiency, and user experience, and has the following form:

[0165] f(θ) = w1·sync(θ) + w2·eff(θ) + w3·exp(θ),

[0166] where θ represents the set of all adjustable parameters, sync, eff, and exp respectively represent the scoring functions of synchronization accuracy, computing efficiency, and user experience, and w1, w2, and w3 are the weights of each item.

[0167] Through this adaptive parameter adjustment mechanism, the voice and action synchronization optimization system of the present invention can maintain excellent performance in different usage environments and application scenarios, providing users with stable and reliable synchronization optimization services. In a preferred embodiment of the present invention, the system further includes an artificial intelligence learning module 6. This module is electrically connected to the analysis and processing unit 2, and is used to receive the historical synchronization analysis data of the analysis and processing unit 2, train these data based on machine learning algorithms, generate an optimized synchronization prediction model, and feedback this model to the analysis and processing unit 2 to improve the accuracy and efficiency of synchronization analysis.

[0168] The artificial intelligence learning module 6 adopts an innovative deep reinforcement learning algorithm, which can continuously optimize the synchronization strategy of the system. Specifically, this module models the voice and action synchronization optimization problem as a Markov decision process (MDP), where the state space S includes the current audio features, action features, and synchronization status, the action space A includes possible synchronization adjustment operations, and the reward function R is based on synchronization accuracy and user experience scoring. Preferably, the present invention uses the Double Q-learning algorithm to solve this MDP problem, and its core update formula is as follows:

[0169]

[0170] where Q(s t ,a t ) represents the value function of taking action a t in state s t , α is the learning rate (value range 0.01 - 0.1), γ is the discount factor (value range 0.9 - 0.99), and r t is the immediate reward.

[0171] Through this continuous learning and optimization mechanism, the system of the present invention can adapt to different types of performances and performers with different styles, and continuously improve the synchronization effect. For example, the system can learn the best synchronization strategies under different music styles (such as classical, jazz, rock, etc.), or optimize the synchronization parameters for different performance forms (such as solo, ensemble, dance, etc.).

[0172] Furthermore, the system of the present invention further includes a multi-modal data fusion module 7. This module is electrically connected to the sound and motion signal acquisition unit 1 and the analysis and processing unit 2, and is used to receive various modal data from the sound and motion signal acquisition unit 1, including audio data, video data, motion capture data, and physiological signal data, perform time alignment and feature fusion on these various modal data, generate a fused multi-modal feature vector, and transmit this feature vector to the analysis and processing unit 2 to improve the comprehensiveness and robustness of synchronization analysis.

[0173] The multi-modal data fusion module 7 adopts a method that combines an innovative attention mechanism and a graph neural network (GNN). First, the self-attention mechanism is used to process the data of each modality to capture the long-range dependencies within the modality. Then, the feature representations of different modalities are constructed into a heterogeneous graph, where each node represents the feature of a modality and the edges represent the relationships between modalities. Finally, the graph attention network (GAT) is used to process this heterogeneous graph to obtain the fused multi-modal feature representation. This process can be expressed as:

[0174] H = GAT(SelfAttention(X1), SelfAttention(X2),..., SelfAttention(X n ))

[0175] where, X i represents the original feature of the i-th modality, SelfAttention represents the self-attention operation, GAT represents the graph attention network, and H is the finally fused feature representation. Preferably, the system of the present invention adopts different sampling rates and preprocessing methods for data of different modalities. For example, the audio data sampling rate is 48 kHz, the video data frame rate is 60 fps, the motion capture data sampling rate is 120 Hz, and the physiological signal (such as heart rate, skin conductance, etc.) sampling rate is 1 kHz. During the time alignment process, the system adopts the dynamic time warping (DTW) algorithm to align the data of all modalities to a unified time axis, and uses the highest sampling rate (48 kHz for audio in this example) as the benchmark.

[0176] In one embodiment of the present invention, the system further includes a real-time feedback optimization module 8. This module is electrically connected to the performance action execution unit 3, the screen rendering unit 4, and the user interaction unit 5, and is used to monitor in real time the execution effect of the performance action execution unit 3 and the rendering effect of the screen rendering unit 4, receive real-time feedback information from the user interaction unit 5, calculate real-time optimization parameters based on this information, and feedback these parameters to the analysis and processing unit 2 for dynamically adjusting the synchronization optimization strategy.

[0177] The real-time feedback optimization module 8 adopts an innovative online learning algorithm, which can continuously adjust and optimize the system parameters during the performance. Specifically, this module uses the method of incremental learning, and updates the model parameters every time a new feedback sample is received. Preferably, the present invention adopts the Online Stochastic Gradient Descent (OnlineSGD) algorithm, and its update formula is as follows:

[0178]

[0179] where θ t represents the model parameters at time t, η t is the learning rate (adopting a decay strategy, with an initial value of 0.1 and decaying to 0.9 times the original value every 1000 iterations), L is the loss function, x t and y t are the input features and target outputs at time t, respectively.

[0180] Through this real-time feedback optimization mechanism, the system of the present invention can quickly adapt to changes during the performance, such as the fatigue state of the performer, the reaction of the audience, etc., so as to maintain the best synchronization effect.

[0181] Finally, the present invention also provides a method for optimizing the synchronization of sound and action, which is applied to the sound and action synchronization optimization system described above. This method includes the following steps:

[0182] S1. Collect the sound signal and action signal emitted by the performer;

[0183] S2. Convert the sound signal and action signal into digital signals;

[0184] S3. Extract features and perform synchronization analysis on the digitized sound signal and action signal;

[0185] S4. Generate a synchronization adjustment signal executable by the performance action execution unit;

[0186] S5. Execute the synchronization adjustment signal and drive the screen rendering unit;

[0187] S6. Render the screen corresponding to the synchronization adjustment signal in real time and send the rendered screen to the user;

[0188] S7. Receive the operation information input by the user and set the parameters of the sound and action synchronization optimization system;

[0189] S8. Use the artificial intelligence learning module to train the historical synchronization analysis data to generate an optimized synchronization prediction model;

[0190] S9. Through the multi-modal data fusion module, perform time alignment and feature fusion on various modal data to generate multi-modal feature vectors;

[0191] S10. Use the real-time feedback optimization module to dynamically adjust the synchronization optimization strategy to achieve continuous optimized synchronization of sound and action.

[0192] In the specific implementation process, step S3 can use the aforementioned multi-scale cross-correlation analysis algorithm for synchronization analysis, and step S4 can use the adaptive delay compensation algorithm to generate a synchronization adjustment signal. The training process in step S8 can use the aforementioned deep reinforcement learning algorithm, and the feature fusion in step S9 can use a method combining the attention mechanism and the graph neural network. Step S10 can use the online stochastic gradient descent algorithm for real-time optimization.

[0193] Through this method, the present invention can achieve high-precision and adaptive synchronization optimization of sound and action, providing strong technical support for various application scenarios such as performing arts, virtual reality, and human-computer interaction.

[0194] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A sound and action synchronization optimization system, characterized in that, Including: A sound and motion signal acquisition unit, configured to: Acquire the sound signals and motion signals emitted by the performer; Convert the sound signals and motion signals into digital signals; An analysis and processing unit, electrically connected to the sound and motion signal acquisition unit, configured to: Receive the digitized sound signals and motion signals sent by the sound and motion signal acquisition unit; Extract features and perform synchronous analysis on the digitized sound signals and motion signals; Generate a synchronous adjustment signal executable by the performance action execution unit; A performance action execution unit, electrically connected to the analysis and processing unit, configured to: Receive the synchronous adjustment signal sent by the analysis and processing unit; Execute the synchronous adjustment signal; Drive the screen rendering unit to combine the synchronous adjustment signal with the screen; A screen rendering unit, electrically connected to the performance action execution unit, configured to: Receive the drive signal from the performance action execution unit; Render the screen corresponding to the synchronous adjustment signal in real time; Send the rendered screen to the user; A user interaction unit, electrically connected to the analysis and processing unit, the performance action execution unit, and the screen rendering unit, configured to: Receive the operation information input by the user; Perform parameter settings on the sound and motion synchronization optimization system; Feedback the synchronization optimization effect to the user.

2. The sound and action synchronization optimization system according to claim 1, characterized in that The sound and motion signal acquisition unit includes: A sound processing module, configured to: Acquire the sound signals emitted by the performer; Convert the sound signals into digital signals to obtain A / D audio signals; A motion processing module, configured to: Acquire the motion signals emitted by the performer; Convert the motion signals into digital signals to obtain motion signals; Wherein, the sound processing module and the motion processing module are connected through a data bus to achieve synchronous data transmission.

3. The sound and motion synchronization optimization system according to claim 2, wherein The analysis and processing unit includes: A synchronous adjustment module, configured to: Receive the A / D audio signals and motion signals; Process the A / D audio signals and motion signals; Generate an executable synchronous adjustment signal; A motion processing module, configured to: Analyze the motion signals; Generate an executable motion signal; An output signal generation module, electrically connected to the synchronous adjustment module and the motion processing module, configured to: Receive the synchronous adjustment signal and the executable motion signal; Generate a final executable signal.

4. The sound and motion synchronization optimization system according to claim 3, characterized in that The synchronous adjustment module includes: A synchronous analysis unit, configured to: Analyze the synchronization relationship between the A / D audio signals and the motion signals; Calculate the synchronous adjustment coefficient; A synchronous adjustment unit, electrically connected to the synchronous analysis unit, configured to: Generate an A / D audio delay adjustment coefficient according to the calculation result of the synchronous analysis unit; Generate a music enhancement coefficient; Realize the synchronization of the A / D audio signals and the motion signals.

5. The sound and action synchronization optimization system according to claim 3, characterized in that The sound and motion signal acquisition unit further includes a preprocessing module, and the preprocessing module is configured to: Perform noise reduction processing on the A / D audio signals; Perform filtering processing on the motion signals; Wherein, the preprocessing module is arranged between the sound processing module and the synchronous adjustment module, and the preprocessing module is also electrically connected to the synchronous adjustment module and the output signal generation module, and is configured to: Transmit the sound signals after noise reduction processing to the synchronous adjustment module and the output signal generation module; The processed action signal after filtering is transmitted to the synchronization adjustment module and the output signal generation module.

6. The sound and action synchronization optimization system according to claim 5, characterized in that, The signal denoising module, the filtering module, the synchronization analysis unit, the synchronization adjustment unit, and the output signal generation module are electrically connected to the user interaction unit; The user interaction unit includes a tuning signal module, and the tuning signal module is used for: Sending adjustment parameters to the signal denoising module, the filtering module, the synchronization analysis unit, the synchronization adjustment unit, and the output signal generation module; Adjusting their parameters; Among them, the adjustment parameters include denoising parameters, filtering parameters, synchronization evaluation parameters, and output control parameters, and are specifically used for: The preprocessing module performs denoising processing on the A / D audio signal according to the denoising parameters; The preprocessing module performs filtering processing on the action signal according to the filtering parameters; The synchronization analysis unit analyzes the synchronization relationship between the A / D audio signal and the action signal according to the synchronization evaluation parameters to calculate the synchronization adjustment coefficient; The synchronization adjustment unit generates an A / D audio delay adjustment coefficient and a music enhancement coefficient according to the synchronization adjustment coefficient; The output signal generation module splices the A / D audio signal and the action signal according to the output control parameters to generate an executable synchronization adjustment signal.

7. The sound and action synchronization optimization system according to claim 1, characterized in that It further includes an artificial intelligence learning module, and the artificial intelligence learning module is electrically connected to the analysis and processing unit for: Receiving the historical synchronization analysis data of the analysis and processing unit; Training the historical data based on machine learning algorithms; Generating an optimized synchronization prediction model; Feeding back the optimized synchronization prediction model to the analysis and processing unit to improve the accuracy and efficiency of synchronization analysis.

8. The sound and action synchronization optimization system according to claim 1, wherein It further includes a multimodal data fusion module, and the multimodal data fusion module is electrically connected to the sound and action signal acquisition unit and the analysis and processing unit for: Receiving various modal data from the sound and action signal acquisition unit, including audio data, video data, motion capture data, and physiological signal data; Performing time alignment and feature fusion on the various modal data; Generating a fused multimodal feature vector; Transmitting the multimodal feature vector to the analysis and processing unit to improve the comprehensiveness and robustness of synchronization analysis.

9. The sound and action synchronization optimization system according to claim 1, characterized in that, It further includes a real-time feedback optimization module, and the real-time feedback optimization module is electrically connected to the performance action execution unit, the screen rendering unit, and the user interaction unit for: Real-time monitoring the execution effect of the performance action execution unit and the rendering effect of the screen rendering unit; Receiving real-time feedback information from the user interaction unit; Calculating real-time optimization parameters based on the execution effect, rendering effect, and user feedback information; Feeding back the real-time optimization parameters to the analysis and processing unit for dynamically adjusting the synchronization optimization strategy.

10. A method for optimizing the synchronization of sound and action, applied to the sound and action synchronization optimization system according to any one of claims 1-9, characterized in that, It includes the following steps: S1. Collect the sound signal and action signal emitted by the performer; S2. Convert the sound signal and action signal into digital signals; S3. Perform feature extraction and synchronization analysis on the digitized sound signal and action signal; S4. Generate an executable synchronization adjustment signal for the performance action execution unit; S5. Execute the synchronization adjustment signal and drive the screen rendering unit; S6. Render the screen corresponding to the synchronization adjustment signal in real time and send the rendered screen to the user; S7. Receive the operation information input by the user and set the parameters of the sound and action synchronization optimization system; S8. Use the artificial intelligence learning module to train the historical synchronization analysis data to generate an optimized synchronization prediction model; S9. Perform time alignment and feature fusion on multiple modal data through the multi-modal data fusion module to generate multi-modal feature vectors; S10. Use the real-time feedback optimization module to dynamically adjust the synchronization optimization strategy to achieve continuous optimized synchronization of sound and action.