Data alignment method in remote playing, remote playing system and medium
By employing timestamps and score alignment algorithms in remote performances, the problem of audio-visual asynchrony in remote performances is solved, achieving high-precision data alignment, improving user experience and device compatibility, and making it suitable for remote teaching and virtual ensembles.
Patent Information
- Application Number
- CN202511300971.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-12
AI Technical Summary
During remote performances, there are delay differences in various types of data at the receiving end, resulting in asynchrony between sound and picture, and misalignment between movements and sounds, which affects the continuity and realism of the performance and significantly reduces the user experience.
By employing a timestamp-based data alignment method, combining hardware data and video data, and through a preset score alignment algorithm and dynamic adaptive strategy, high-precision alignment of heterogeneous multi-type data is achieved under environments with differences in processing links and fluctuations in device computing power.
It achieves high audio-visual synchronization, expands the device compatibility for high-quality remote performance, and provides a low-threshold, highly immersive remote music experience, suitable for scenarios such as remote teaching and virtual ensembles.
Smart Images

Figure CN120825604A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of remote performance, and in particular to a data alignment method, a remote performance system, and a medium in remote performance. Background Art
[0002] With the rapid development of internet technology, remote performance, as an emerging artistic form, is gradually transforming the way music is created, performed, and disseminated. Remote performance not only transcends geographical limitations, enabling artists to collaborate across space, but also opens up new possibilities for music education, cultural exchange, and the popularization of the arts. To enhance the immersiveness and interactivity of remote performances, various types of data are being introduced, such as real-time control signals from intelligent instruments, multi-angle high-definition video streams, and audio signals.
[0003] For example, the invention patent application with patent application publication number CN116939237A discloses a live teaching method based on IOT piano, including setting up a host end; the first terminal is connected to the host end IOT piano through a MIDI data cable to collect the MIDI data of the IOT piano, and the first terminal is connected to multiple video devices through multiple data cables to collect audio and video data in multiple directions; the video acquisition device combines the multiple audio and video data into live picture data of different scenes, and after encoding and compressing, pushes it to the server through the RTMP protocol with the MIDI data of the IOT piano; setting up the audience end; the second terminal is connected to the audience end IOT piano through a MIDI data cable, pulls the live stream of the server through the RTMP protocol for parsing, and parses it into combined audio and video picture data and MIDI data; sends the MIDI data to the audience end IOT piano through the audience end MIDI data cable, and aligns the combined audio and video picture data and MIDI data through timestamps and plays them synchronously.
[0004] For example, the invention patent application with patent application publication number CN110392276A discloses a live broadcast and recording method for synchronously transmitting MIDI based on the RTMP protocol. First, images, audio and MIDI signals are collected, and the collected images and audio are encoded and compressed respectively; secondly, the collected MIDI signals and the encoded and compressed audio are mixed and pushed to the server together with the encoded and compressed images, and the encoded and compressed images and the mixed audio and MIDI signals are recorded and broadcast at the same time; then, based on the RTMP protocol, the stream is pulled from the server and the mixed audio and MIDI data are separated, and the audio data and image data are decoded at the same time; then, the decoded image data and audio data are integrated into audio and video for playback, and the separated MIDI data is sent to the playing instruments; finally, the synchronous linkage between the playback of audio and video and the playing of musical instruments is realized.
[0005] However, there are delay differences among various types of data at the receiving end, making it difficult to achieve precise alignment. This leads to problems such as asynchrony between sound and picture, and misalignment between action and sound, which seriously affects the continuity and realism of the performance and significantly reduces the user experience. Summary of the Invention
[0006] The main purpose of this application is to provide a data alignment method, remote performance system and medium for remote performance. In order to solve the above-mentioned technical problems, this application specifically adopts the following technical solutions: A first aspect of the present application is to provide a data alignment method for remote performance, the method comprising: S101, collecting hardware data and video data during a performance based on a first device end, and sending the data to a second device end, wherein the hardware data and the video data are associated with a timestamp; the second device end includes a musical instrument end and a first display end and a second display end; S102, controlling the instrument end to parse and present the hardware data; controlling the first display end to parse and present the video data; S103 determines score location data of the current hardware data based on a preset score alignment algorithm and a plurality of hardware data and a preset performance score received within a preset time period of the current hardware data, wherein the score location data is associated with a timestamp of the current hardware data; Depending on the computing power of the device, choose to execute step S1031 or execute steps S1032 to S1034: S1031, when the current hardware data is presented on the musical instrument end, controlling the second display end to synchronously present the music score segment pointed to by the music score positioning data; S1032: Collecting music score advancement data during the performance based on the first device, and sending the data to the second device, wherein the music score advancement data is associated with a timestamp; S1033, based on a preset period, comparing the position deviation of the music score segment pointed to by the music score positioning data and the music score advancement data at the same timestamp; S1034, when the current hardware data is presented on the instrument end, selecting the music score segment pointed to by the music score positioning data or the music score advancement data according to the position deviation, and controlling the second device end to present the corresponding music score segment; S1041, extracting timestamps of hardware data and video data presented at the same time, and calculating a first time offset between the two timestamps; S1042: When the first time deviation is greater than or equal to a first deviation threshold, based on the target timestamp of the hardware data to be presented at the next moment, call the target video data of the target timestamp, and control the first display end to present the target video data at the next moment.
[0007] In some embodiments, the preset time period includes a first time period before and a second time period after the timestamp of the current hardware data, wherein the length of the first time period is a first preset duration, and the length of the second time period is a second preset duration.
[0008] In some embodiments, S1034 includes: when the position deviation is less than a preset difference, presenting the second music score segment pointed to by the music score advancement data on the second device end; or, when the position deviation is greater than or equal to the preset difference, presenting the first music score segment pointed to by the music score positioning data on the second device end.
[0009] In some embodiments, the instrument end is used to drive the sound structure to execute the hardware data; the first display end is used to play the video data; and the second display end is used to display the music score fragment corresponding to the music score advancement data and / or music score positioning data.
[0010] In some embodiments, the method also includes: when the first time deviation is greater than the second deviation threshold and less than the first deviation threshold, analyzing the presentation time difference between the hardware data and the video data at the same timestamp based on the first time deviation within a preset time; adjusting the presentation speed of the video data based on the presentation time difference so that the first time deviation remains less than the first deviation threshold.
[0011] In some embodiments, the method further includes: in response to the user's progress adjustment operation at the first display end, determining the video data to be displayed, and based on the first timestamp of the video data to be displayed, calling the hardware data of the first timestamp for presentation; and / or, in response to the user's music score selection operation at the second display end, determining the music score advancement data or music score positioning data to be displayed, and based on the second timestamp of the music score advancement data or music score positioning data to be displayed, calling the hardware data of the second timestamp for presentation.
[0012] In some embodiments, the method also includes: obtaining first performance data, second performance data, and third performance data respectively based on the hardware data, the video data, and the music score advancement data at the same timestamp; when the first performance data is different from the second performance data and the third performance data, generating correction hardware data based on the second performance data and the third performance data; and presenting it on the instrument end based on the correction hardware data.
[0013] In some embodiments, the method further includes: based on the timestamp, encapsulating the hardware data and the video data with the same timestamp into one data packet; or, based on the timestamp, encapsulating the hardware data and the video data and the music score advancement data with the same timestamp into one data packet; and distributing the encapsulated data packets over the network.
[0014] A second aspect of the present application is to provide a remote performance system, wherein the system is applied to the data alignment method for remote performance provided in any embodiment of the present application; the system comprises: The first device end is used to collect hardware data, video data, and music score advancement data during the performance and send them to the second device end, wherein the hardware data, the video data, and the music score advancement data are associated with a timestamp; The second device end includes a musical instrument end and a first display end and a second display end; the musical instrument end is used to parse and present the hardware data; the first display end is used to parse and present the video data; the second display end is used to determine the score location data of the current hardware data based on a preset score alignment algorithm and a number of hardware data and preset performance scores received within a preset time period of the current hardware data, wherein the score location data is associated with a timestamp of the current hardware data; when the current hardware data is presented on the musical instrument end, the score segment pointed to by the score location data or the score advancement data is presented; The control unit is configured to extract the timestamps of the hardware data and the video data presented at the same moment and calculate a first time offset between the two timestamps; when the first time offset is greater than or equal to a first offset threshold, call the target video data of the target timestamp based on the target timestamp of the hardware data to be presented at the next moment, and control the first display end to present the target video data at the next moment; The control unit is also used to, when it is detected that the computing power of the device is sufficient, control the second display end to synchronously present the music score segment pointed to by the music score positioning data; or, when it is detected that the computing power of the device is insufficient, compare the position deviation of the music score segment pointed to by the music score positioning data and the music score advancement data at the same timestamp based on a preset period; select the music score segment pointed to by the music score positioning data or the music score advancement data according to the position deviation, and control the second device end to present the corresponding music score segment.
[0015] The third aspect of the present application is to provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor performs the steps of the data alignment method in remote performance provided in any embodiment of the present application.
[0016] Beneficial effects: The embodiments of the present application provide a data alignment method, a remote performance system, and a medium for remote performance. Based on stable and reliable hardware data, the system combines video data and music score data for intelligent collaborative presentation, thereby achieving high-precision alignment of heterogeneous and multi-category data in an environment with processing link differences and fluctuations in device computing power. While ensuring high synchronization of sound and picture, the system expands the device compatibility for high-quality remote performances, achieving a low-threshold, highly immersive remote music experience, and is widely applicable to remote performance scenarios such as remote teaching and virtual ensemble.
[0017] During data analysis and processing, priority is given to ensuring real-time processing and accurate restoration of hardware data, allocating limited computing power to the hardware data processing paths directly related to sound generation, ensuring low latency and high responsiveness in driving the instrument's sound generation structure, maintaining a stable auditory experience and a coherent playing rhythm. Video and music scores serve as auxiliary information, and their processing can be dynamically adjusted based on device load: In terms of video display, a dynamic adaptive strategy is adopted: based on the deviation threshold, timely jump to the corresponding timestamp frame to complete alignment, which not only ensures the synchronization of audio and video, but also avoids obvious frame skipping that affects the visual experience. At the same time, regular deviations are analyzed and smooth compensation is performed by pre-adjusting the video presentation speed. The user is unaware of the whole process, thus achieving high-quality immersive presentation even on devices with weaker computing power.
[0018] In terms of music score display, performance and experience are optimized through the collaboration of dual-source data: when computing power is sufficient, high-precision music score positioning is generated based on hardware data within a preset time period to achieve delay-free advancement; when computing power is insufficient, the music score advancement signal pushed remotely is used as the basis, and the deviation is dynamically corrected in combination with a local lightweight correction algorithm to avoid music score lag or jamming due to insufficient local processing power, significantly reducing the terminal processing burden and enabling ordinary devices to smoothly participate in high-quality remote performances.
[0019] Furthermore, for special situations (such as user selection of video progress, music score position or hardware data anomalies), an interactive priority and fault-tolerant mechanism is introduced: when responding to user operations, the hardware data benchmark can be temporarily broken through and directly aligned to the specified position; when hardware data anomalies occur, reverse calibration is performed through video data or music score positioning data to avoid error propagation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the embodiments or the description of the prior art. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the various elements or parts are not necessarily drawn according to the actual scale. Obviously, the drawings described below are some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without paying any creative work.
[0021] Figure 1 is a schematic flow chart of a data alignment method in a remote performance provided by an embodiment of the present application; Figure 2 This is a schematic diagram of a data transmission process provided by an embodiment of the present application; Figure 3 This is a schematic diagram of a process for dynamically adjusting the display of music scores provided in an embodiment of the present application; Figure 4 This is a flow chart of dynamically adjusting video display provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] To make the purpose, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0023] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.
[0024] Herein, suffixes such as "module," "component," or "unit" used to represent elements are only used to facilitate the description of the present application and have no specific meaning. Therefore, "module," "component," or "unit" can be used interchangeably.
[0025] As used herein, terms such as "upper," "lower," "inner," "outer," "front," "back," "one end," and "the other end" indicate positions or locations based on those shown in the accompanying drawings. These terms are intended solely to facilitate the description of this application and simplify the description. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0026] As used herein, unless otherwise expressly specified or limited, the terms "installed," "provided with," "connected," etc., should be understood broadly. For example, "connected" may refer to a fixed connection, a detachable connection, or an integral connection; it may refer to a mechanical connection, a direct connection, an indirect connection through an intermediate medium, or internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this application.
[0027] As used herein, "and / or" includes any and all combinations of one or more of the associated listed items.
[0028] Herein, "plurality" means two or more, ie, it includes two, three, four, five, etc.
[0029] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0030] In this article, the remote performance scenario includes a first device and a second device. The first device, also known as the live broadcast initiator or data acquisition terminal, includes an intelligent musical instrument (such as a smart piano) located on the performer's side, with built-in or external image acquisition devices and sensors. Correspondingly, the second device, also known as the live broadcast receiver or data receiver, includes various receiving devices of remote users, such as tablets, mobile phones, and other intelligent musical instruments, with multiple ports for synchronously restoring sound, playing videos, and displaying music scores, enabling a real-time cross-regional performance experience.
[0031] Based on this, the embodiments of the present application provide a data alignment method, a remote performance system and a medium in remote performance. It uses stable and reliable hardware data as a benchmark, combines video data and music score data for intelligent collaborative presentation, and realizes high-precision alignment of heterogeneous and multi-category data in an environment with processing link differences and device computing power fluctuations. While ensuring high synchronization of sound and picture, it expands the device compatibility of high-quality remote performances, and realizes a low-threshold, highly immersive remote music experience. It can be widely used in remote performance scenarios such as remote teaching and virtual ensemble.
[0032] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features of the embodiments can be combined with each other. Figure 1 , Figure 1 This is a schematic flow chart of a data alignment method for remote performance provided by an embodiment of the present application. Figure 1 As shown, an embodiment of the present application provides a data alignment method in remote performance.
[0033] S101, collecting hardware data and video data during a performance based on a first device end, and sending the data to a second device end, wherein the hardware data and the video data are associated with a timestamp.
[0034] Specifically, the performer can perform through an intelligent musical instrument, which collects hardware data through sensors and video data of the user's performance through image acquisition devices. The hardware data and video data are associated with timestamps and sent to the second device end. Among them, hardware data refers to the instrument control signal collected during the performer's performance, which is used by the remote second device end to restore the performer's performance on the instrument through the local instrument end, such as the ID, displacement, force, speed, etc. of the pressed or released keys, and the ID, stepping force, and stepping speed of the pedals that are stepped on or released. Among them, video data refers to the performance process image recorded synchronously during the performer's performance, which allows remote users to observe the performer's hand shape, posture, expression and other visual information, thereby enhancing the immersive feeling of teaching or performance.
[0035] In some embodiments, the first device can also process real-time hardware data to generate music score advancement data, and the music score advancement data is associated with the timestamp of the real-time hardware data. When the second device needs to use the music score advancement data, the hardware data, video data, and music score advancement data are sent to the second device.
[0036] In some embodiments, the first device end includes a source instrument end, a third display end and / or a fourth display end, the source instrument end is used for performance by a performer, such as a smart piano, the third display end is used to present the video data, and the fourth display end is used to process real-time hardware data to generate music score advancement data, and display a second music score segment corresponding to the music score advancement data.
[0037] In some embodiments, the first device collects various types of data (such as hardware data, video data, and music score advancement data) and associates them with timestamps, which are then encapsulated and synchronously transmitted to the second device via a network protocol.
[0038] In some embodiments, the second device includes a musical instrument and a first and second display terminals working in concert. The musical instrument receives hardware data and drives the sound-generating structure to execute the hardware data to reproduce the performance movements and generate sound. The first display terminal receives and plays corresponding video data, presenting the performer's movements in real time. The second display terminal displays a music score fragment generated by music score progression data or music score positioning data. This achieves distributed, synchronized presentation of sound, images, and music scores, making it suitable for remote learning and multi-screen interactive scenarios.
[0039] In some embodiments, multiple ports of the second device end (such as the instrument end and the first display end and the second display end) can be connected to a specially configured remote data processing unit to process various types of received data through shared computing resources.
[0040] In another embodiment, multiple ports on the second device can independently process received data using the data processing unit of their respective application devices. For example, hardware data is parsed by a dedicated processor on a musical instrument (e.g., a smart piano), resulting in stable computing power. The first and second display terminals can be used on user-defined devices, such as mobile phones, tablets, or PCs, which vary in hardware configuration, system load, and decoding capabilities, making computing power uncontrollable.
[0041] In some embodiments, the sound-producing structure refers to the physical or electronic device on the instrument that produces sound. In acoustic instruments, such as a smart piano, the sound-producing structure includes mechanical components such as strings, soundboard, and hammer mechanism. These components can be automatically controlled through the mechanical structure, enabling playing actions such as the sound produced by the keys striking the strings. In electroacoustic or digital instruments, the sound-producing structure consists of an audio processor and a speaker or headphone output module, responsible for converting received hardware data (such as note and velocity) into audible sound signals.
[0042] In some embodiments, the source instrument on the first device and the instrument on the second device can be the same or different, and can both be intelligent instruments with autonomous sound generation capabilities, such as existing intelligent pianos with automated remote control. The sound generation structure of the source instrument generates sound in response to the performer's playing movements, while the instrument automatically controls the sound generation structure based on hardware data corresponding to the playing movements. Furthermore, the instrument can also be a sound player or an independent automatic sound generation structure.
[0043] In some embodiments, the application devices of the first display end and the second display end can be the same or different, or they can be the same device. After data processing, the display functions of the first display end and the second display end are realized through different interfaces in the device, such as partitioning the display of video screen and music score screen.
[0044] In some embodiments, when the second device does not need to use music score to advance data, the method includes: based on the timestamp, encapsulating the hardware data and the video data with the same timestamp into a data packet, and distributing the encapsulated data packet over the network.
[0045] In some embodiments, when the second device needs to use music score advancement data, the method includes: based on the timestamp, encapsulating the hardware data and the video data and the music score advancement data with the same timestamp into a data packet, and distributing the encapsulated data packet over the network.
[0046] Specifically, at the live broadcast initiator, various data such as hardware data, video data, and music score advancement data are synchronized with timestamps, and various data with the same timestamps are encapsulated into a unified data packet. Figure 2 , Figure 2 This is a flow chart of a data transmission process provided by an embodiment of the present application, such as Figure 2 As shown, when the music score advancement data needs to be transmitted, the hardware data, video data, and music score advancement data with the same timestamp are encapsulated into a unified data packet and transmitted to the second device end.
[0047] In some embodiments, data is uniformly distributed over the network via the UDP protocol to ensure consistent timing at the receiving end.
[0048] S102, controlling the instrument end to parse and present the hardware data; controlling the first display end to parse and present the video data.
[0049] The first display end parses the video data based on the corresponding data processing unit. The video data is high-bandwidth streaming media and needs to undergo multiple processing links such as decapsulation, decoding, and rendering at the receiving end. The processing link is long and complex. On this basis, the allocation priority of the shared computing power of the video data is lower than that of the instrument end, or the computing power of the user-defined device is unstable. Therefore, the video data is easily affected by fluctuations in device performance, which leads to delays in presentation.
[0050] Specifically, the instrument side parses the hardware data through its dedicated or shared data processing unit. This data is a lightweight control signal, such as key depression status, touch force, and pedal position. It directly drives the sound structure through a short processing link, with low computational overhead. Furthermore, during resource scheduling, computing power allocation is prioritized for tasks directly related to sound production, ensuring their highest processing priority. Furthermore, the instrument side typically uses a dedicated processor with a stable operating environment and sufficient available computing power. Therefore, hardware data processing latency is low and highly stable, providing a stable and reliable listening experience for remote performances.
[0051] The first display terminal must parse the video data through its own device or a shared data processing unit. Video data is high-bandwidth streaming media and requires multiple processing steps, including decapsulation, decoding, and rendering. This results in a long and complex link. Furthermore, in scenarios where computing resources are shared, video data has a lower priority than instrument-based computing, resulting in a higher risk of latency. In scenarios where computing power is provided by user-defined devices, performance fluctuations between devices can cause the image to lag or lag.
[0052] It should be understood that even if the second display terminal receives two types of data at the same time, the final presentation moment may still cause audio and video delays due to differences in data processing links and computing power. This audio and video asynchrony phenomenon is particularly obvious on low-performance display devices, which may seriously affect the immersiveness of remote performance and the accuracy of teaching.
[0053] In some embodiments, hardware data is prioritized over video data, score location data, and score advancement data in computing power allocation to prevent the processing of auxiliary information from excessively occupying resources. Furthermore, based on different user needs, the priority of score location data and score advancement data can be set higher than video data to prioritize score tracking, or video data can be set higher than score location data and score advancement data to prioritize video smoothness.
[0054] In some embodiments, the second display end is controlled to generate music score positioning data of the current hardware data according to the received hardware data. When the current hardware data is presented on the instrument end, the music score positioning data is synchronously presented on the second display end.
[0055] Specifically, the second device analyzes continuous playing actions (such as note sequences and rhythm patterns) based on the received hardware data stream, and dynamically matches the current score position corresponding to the current hardware data in the local score to obtain score positioning data, wherein the score positioning data is information about the specific position of the current performance in the preset score, which may include information such as measure number and beat, and is used to drive the automatic page turning and highlighting of the score. Furthermore, the score positioning data can be bound to the timestamp of the current hardware data. Based on the timestamp, when the hardware data is played and presented on the instrument end, the corresponding score positioning data can be called and displayed synchronously on the second display end. Since the generation process of the score positioning data is completed locally at the receiving end, there is no need to rely on remote data, which avoids network transmission delays and realizes delay-free score display synchronized with the hardware data.
[0056] S103, based on a preset score alignment algorithm, according to a number of hardware data and preset performance scores received within a preset time period of the current hardware data, determine the score positioning data of the current hardware data, wherein the score positioning data is associated with a timestamp of the current hardware data.
[0057] Specifically, based on the timestamp of the current hardware data, the continuous hardware data received within a preset time period before and after the current moment, as well as the locally pre-stored or pre-acquired performance score are obtained, and matching analysis is performed based on the preset score alignment algorithm, so as to accurately calculate the score position corresponding to the current hardware data, generate score positioning data with a timestamp, and then determine the score segment that should be displayed currently, that is, the first score fragment, based on the score positioning data.
[0058] In some embodiments, the preset time period includes a first time period before and a second time period after the timestamp of the current hardware data, wherein the length of the first time period is a first preset duration, and the length of the second time period is a second preset duration. In other words, the preset time period is a time window obtained by extending the first preset time or the second preset duration before and after the current moment of the current hardware data as a reference moment, so as to generate accurate music score positioning data based on the context information of the current hardware data. The first preset duration and the second preset duration can be the same or different, and the specific values can be flexibly set and adjusted according to actual needs.
[0059] In some embodiments, the preset performance score is pre-stored locally or downloaded music score data, including note sequence, rhythm, and other information. The preset score alignment algorithm is an algorithm used to match real-time performance data (e.g., pitch and duration corresponding to hardware data) with a standard score. For details, please refer to the relevant art.
[0060] According to the computing power of the device, step S1031 or steps S1032 to S1034 are selected. The computing power of the device refers to the computing power of each component in the second device end (such as the instrument end, the first display end, and the second display end) to process the received data, which can be divided into centralized shared computing power and independent distributed computing power at each end. It should be understood that when the computing power of the device is sufficient, all types of data can be decoded and presented in real time and without delay, with less lag, frame loss or dislocation. When the computing power of the device is insufficient or tight, the audio and video will be out of sync (such as the frequent occurrence of the first time deviation being greater than or equal to the first deviation threshold), the score refresh will be delayed (such as the generation rate of the score positioning data is lower than the preset rate), or data packet loss and response lag will occur due to overload of processing power.
[0061] For example, pre-set criteria for determining sufficient or insufficient computing power are used, and relevant technologies are used to monitor both centralized shared computing power and independent distributed computing power on each end in real time to assess the computing power of the device. For example, real-time monitoring of CPU usage, memory usage, GPU load, and temperature; real-time monitoring of the processing speed of key processing tasks (such as video frame decoding); and monitoring of data packet loss rates are also possible.
[0062] When it is monitored that the computing power of the device is sufficient, step S1031 is executed; when it is monitored that the computing power of the device is insufficient, steps S1032 to S1034 are executed.
[0063] S1031, when the current hardware data is presented on the musical instrument end, controlling the second display end to synchronously present the music score segment pointed to by the music score positioning data.
[0064] Specifically, when the computing power of the device is sufficient, the timestamp of the music score positioning data is the same as the timestamp of the current hardware data. Therefore, the second device end can be controlled based on the timestamp to display the corresponding first music score fragment when the current hardware data is presented.
[0065] S1032: Based on the first device end, music score advancement data during the performance is collected, and the music score advancement data is sent to the second device end, where the music score advancement data is associated with a timestamp.
[0066] S1033: Based on a preset period, compare the position deviation of the music score segment pointed to by the music score positioning data and the music score advancement data at the same timestamp.
[0067] Among them, the preset period is a pre-set time interval for comparing two types of music score data, and the specific value can be flexibly set and adjusted according to actual needs.
[0068] S1034, when the current hardware data is presented on the instrument side, the music score positioning data or the music score segment pointed to by the music score advancement data is selected according to the position deviation, and the second device side is controlled to present the corresponding music score segment.
[0069] Specifically, in special circumstances such as limited device computing power and excessive local computing burden, the second device end needs to use the remote music score advancement data. At this time, the music score advancement data generated by the first device end is collected and sent to the second device end. At the same time, the music score positioning data is generated locally on the second device end based on the hardware data. During the remote performance, the position deviation between the music score positioning data generated by this end at the same timestamp and the music score fragment pointed to by the transmitted music score advancement data is compared based on a preset period.
[0070] It should be understood that score positioning data is based on hardware data from a period around the current moment, analyzed through a preset score alignment algorithm to analyze the preceding and following performance sequences. The contextual information it relies on is complete, resulting in greater stability and accuracy. However, score advancement data is dynamically generated by the first device based on real-time performance, predicting the current score position based solely on the performed portion (i.e., the contextual information). It lacks subsequent information correction and is susceptible to various factors, such as repeated passages, misplays, and improvisation, leading to the risk of recognition bias.
[0071] Therefore, the accuracy of the music score advancement data generated remotely in real time is verified through the music score positioning data. When the accuracy is low, the music score segment pointed to by the music score positioning data is selected for presentation; when the accuracy is high, the music score segment pointed to by the music score advancement data is selected for presentation.
[0072] In some embodiments, S1034 includes: when the position deviation is less than a preset difference, presenting the second music score segment pointed to by the music score advancement data on the second device end; or, when the position deviation is greater than or equal to the preset difference, presenting the first music score segment pointed to by the music score positioning data on the second device end.
[0073] The preset difference is a preset maximum threshold of the position deviation allowed between the two types of score data. The specific value can be flexibly set and adjusted according to actual needs.
[0074] If the positional deviation is less than the preset difference, the accuracy of the score advancement data at this point is determined to be high, and the second score segment corresponding to the score advancement data is maintained on display to conserve local computing power. If the positional deviation is greater than or equal to the preset difference, the score advancement data at this point is determined to have a recognition error due to relying solely on the above information, and the display is switched back to the first score segment corresponding to the score positioning data, thereby ensuring display accuracy under different network and device performance conditions. Furthermore, after displaying the first score segment corresponding to the score positioning data for a certain period of time, the accuracy of the score advancement data is re-verified. If the score advancement data has completed self-correction at this point, the display of the second score segment corresponding to the score advancement data is restored.
[0075] Among them, the preset period is a pre-set time interval for comparing two types of music score data, and the preset difference is a pre-set maximum threshold value of the allowable position deviation of the two types of music score data. The specific value can be flexibly set and adjusted according to actual needs.
[0076] It should be understood that score positioning data is based on hardware data from a period around the current moment, analyzed through a preset score alignment algorithm to analyze the preceding and following performance sequences. The contextual information it relies on is complete, resulting in greater stability and accuracy. However, score advancement data is dynamically generated by the first device based on real-time performance, predicting the current score position based solely on the performed portion (i.e., the contextual information). It lacks subsequent information correction and is susceptible to various factors, such as repeated passages, misplays, and improvisation, leading to the risk of recognition bias.
[0077] See also Figure 3 , Figure 3 This is a flow chart of a music score display provided by an embodiment of the present application. Figure 3 As shown, the first device collects hardware data and music score advancement data during the performance and transmits them to the second device.
[0078] When computing power resources are sufficient, the locally generated music score positioning data is used to drive the music score presentation on the second display terminal to ensure display accuracy. Figure 3 As shown, the second device end can determine the score positioning data of the current hardware data based on a preset score alignment algorithm, according to a number of hardware data and preset performance scores received within a preset time period of the current hardware data, and present the first score fragment pointed to by the score positioning data on the second device end.
[0079] Furthermore, in the case of limited computing power of the device, excessive system load and other conditions where computing power resources are tight, in order to reduce the local computing burden, the music score advancement data transmitted remotely can be switched as an alternative to maintain the continuity of music score advancement. Figure 3As shown, based on a preset period, the position deviation of the music score positioning data and the music score advancement data is compared; when the position deviation is less than the preset difference, the second music score segment corresponding to the music score advancement data is presented on the second device end; or, when the position deviation is greater than or equal to the preset difference, the corresponding first music score segment is presented on the second device end according to the music score positioning data.
[0080] It should be understood that in order to control the display error, a periodic comparison mechanism is set to verify the consistency of the two types of data, limiting the number of times the second device generates music score positioning data, and significantly reducing the computing resources it occupies. Once it is found that the deviation exceeds the threshold, it is immediately switched back to the high-precision music score positioning data for correction, and maintained for use within the second preset time. As the performance progresses, the music score advancement data gradually corrects itself, and the verification is triggered again after the second preset time. If the position deviation is less than the preset difference, the music score advancement data is restored. Among them, the second preset time can be flexibly determined based on the computing power of the second device. For example, when the computing power is relatively tight, the second preset time can be set shorter to actively verify whether the music score advancement data has completed self-correction, and to switch back to the application of the music score advancement data in a timely manner. In this way, while ensuring accuracy, the device compatibility of high-quality remote performances is expanded.
[0081] S1041 , extracting the timestamps of the hardware data and the video data presented at the same time, and calculating a first time offset between the two timestamps.
[0082] Specifically, the second device extracts the timestamps associated with the hardware data and video data presented at the same time and calculates the time difference between the two, known as the first time offset. This first time offset reflects the degree of synchronization between the sound and video during actual output, and is used to determine whether there is any audio-visual asynchrony, thereby dynamically triggering subsequent correction mechanisms to improve the consistency of sound and video during remote performances.
[0083] S1042: When the first time deviation is greater than or equal to a first deviation threshold, based on the target timestamp of the hardware data to be presented at the next moment, call the target video data of the target timestamp, and control the first display end to present the target video data at the next moment.
[0084] Specifically, when it is detected that the first time deviation exceeds a preset first deviation threshold, the target timestamp of the hardware data to be presented at the next moment is used as a reference, and the target video data with the same timestamp is searched and called from the received video data stream. If the target video data has been parsed and processed at this time, it can be directly called for rendering and display at the next moment. If the target video data has not been parsed and processed at this time, the processing and presentation steps of the delayed video data can be skipped, and the first display end can be directly instructed to process and present the target video data at the next moment.
[0085] It should be understood that by achieving rapid frame alignment of the video screen through precise timestamp matching, the presentation of some video data can be omitted, saving computing power resources while maintaining a high degree of consistency between the hardware data and the video data presented to the user at the next moment. Even if the processing and presentation of the target video data may introduce new delays, it can, to a certain extent, correct the impending significant and perceptible audio and video delays, avoiding the disconnection between audio and video caused by the accumulation of decoding or rendering delays on the display end. This is particularly suitable for scenarios with insufficient device performance, effectively restoring the synchronization between hearing and vision, and ensuring the continuity and realism of remote performances.
[0086] The first deviation threshold is a preset maximum allowable time deviation between audio and video presentation, used to trigger video frame skipping. The specific value can be set based on the human eye's sensitivity to audio-video asynchrony and image jumps. For example, if the first time deviation is equal to or slightly greater than the first deviation threshold, frame skipping will be less noticeable or have a minimal impact, thus minimizing the impact of frame skipping on viewing experience.
[0087] In some embodiments, the method further includes: when the first time deviation is greater than a second deviation threshold and less than a first deviation threshold, accelerating or slowing down the playback speed of the video data based on a preset playback speed.
[0088] Specifically, when the first time deviation is greater than the second deviation threshold and less than the first deviation threshold, it is determined that there is a slight asynchrony between the sound and the video. Although it has not yet reached the level of obvious perception by the human eye, it may continue to accumulate and eventually affect the experience if no intervention is made. In this case, no frame skipping operation is performed. Instead, the playback rate of the video data is dynamically fine-tuned based on the preset playback speed to achieve smooth synchronization. This smooth adjustment method is particularly suitable for scenarios where the performance of the first display terminal is good and the video data processing speed is fast. It can support slightly speeding up or slowing down the playback rhythm of the video, allowing the video picture to gradually catch up or delay the processing and presentation performance and changes of the hardware data, avoiding visual interruptions caused by sudden picture changes or frame skipping, and maintaining playback continuity.
[0089] The second deviation threshold is smaller than the first, used to distinguish between minor and noticeable frame skipping or audio / video mismatches. When the deviation exceeds the second but falls below the first, speed adjustment can be initiated instead of frame skipping. The preset playback speed is the unit of adjustment used for speed acceleration or deceleration. The value should be set to ensure a natural and imperceptible adjustment process.
[0090] In some embodiments, the method also includes: when the first time deviation is greater than the second deviation threshold and less than the first deviation threshold, analyzing the presentation time difference between the hardware data and the video data at the same timestamp based on the first time deviation within a preset time; adjusting the presentation speed of the video data based on the presentation time difference so that the first time deviation remains less than the first deviation threshold.
[0091] Specifically, the number of frame skips within the preset time period is counted based on the first time deviation within the preset time period. The number of frame skips is the number of times step S105 is executed, and is used to quantify the difference in data processing and presentation capabilities between the second device and the instrument. For example, during a remote piano lesson, the student device may trigger four frame skips within one minute due to limited computing resources or insufficient decoding performance.
[0092] In this way, the actual presentation time difference between hardware data and video data can be analyzed. For example, when the first time deviation is 80ms, each frame skip is because the time deviation accumulates to 80ms. After each frame skip, the deviation returns to zero and re-accumulates. The total accumulated delay is 320ms, which is the presentation time difference. Thus, a continuous delay trend is identified. Based on the presentation time difference, the computing power resources occupied by the first display end are increased or decreased based on the presentation time difference, and then the presentation speed of the video data is increased or decreased, so that the first time deviation is continuously controlled within the first deviation threshold, achieving pre-smooth synchronization, reducing the frame skipping trigger frequency, and improving playback continuity. Among them, the preset time refers to a fixed time window used to count the first time deviation, such as 1 minute, which is used to evaluate the time difference performance of video data and hardware data in application presentation in the near future.
[0093] In some embodiments, a mapping relationship table between the presentation time difference and the computing power resources of the first display terminal is preset based on prior data. Based on this mapping relationship table, the computing power resources that need to be adjusted can be quickly determined based on the presentation time difference. Furthermore, the computing power resources occupied by the first display terminal have an upper limit, which is determined based on the total computing power resources of the remote data processing unit, the computing power resources occupied by the instrument terminal, and other necessary computing power resources. This avoids allocating too much computing power resources to the first display terminal, which may affect the presentation of hardware data.
[0094] See also Figure 4 , Figure 4 This is a flow chart of a method for dynamically adjusting video display provided by an embodiment of the present application. Figure 4 As shown, the timestamps of the hardware data and video data presented at the same time are extracted, the first time deviation between the two timestamps is calculated, and the degree of audio-visual mismatch is distinguished based on the first deviation threshold and the second deviation threshold.
[0095] When the first time deviation is greater than the first deviation threshold, the audio and video delay will be perceived by the user, and synchronization is quickly achieved by frame skipping: based on the target timestamp of the hardware data, the target video data of the target timestamp is called, and the first display end is controlled to present the target video data.
[0096] Furthermore, when the first time deviation is greater than the second deviation threshold and less than the first deviation threshold, a slight audio and video delay occurs, and speed adjustment can be activated instead of frame skipping. Figure 4 As shown, the playback speed of the video data can be accelerated or slowed down based on the preset playback speed; or, after running for a period of time, the presentation time difference of the application presentation between the hardware data and the video data can be analyzed based on the first time deviation within the preset time; based on the presentation time difference, the presentation speed of the video data can be accelerated or slowed down to keep the first time deviation less than the first deviation threshold, thereby avoiding frame skipping operations.
[0097] It should be understood that humans are extremely sensitive to audio latency. The direct connection between hardware data and sound control, with hardware data as the core benchmark, can ensure a better auditory experience. Delays or errors in video and music scores, which serve as auxiliary information, are relatively tolerable. By using methods such as screen-jumping video displays and switching to lightweight music score displays, overhead is reduced, avoiding resource competition with hardware data. While ensuring high synchronization between audio and video, this significantly reduces dependence on terminal device performance, effectively improving cross-platform compatibility and providing a low-threshold, highly immersive remote music experience for scenarios such as remote teaching and virtual performances.
[0098] In some embodiments, the method further includes: in response to the user's progress adjustment operation at the first display end, determining the video data to be displayed, and based on the first timestamp of the video data to be displayed, calling the hardware data of the first timestamp for presentation; and / or, in response to the user's music score selection operation at the second display end, determining the music score advancement data or music score positioning data to be displayed, and based on the second timestamp of the music score advancement data or music score positioning data to be displayed, calling the hardware data of the second timestamp for presentation.
[0099] Specifically, when the user performs a progress adjustment operation on the first display end (such as dragging the video playback progress bar), the video data to be displayed corresponding to the target position is determined, and its associated first timestamp is extracted. Subsequently, the hardware data-driven sound structure with the same timestamp is searched in the hardware data stream, and the sound output at that moment is restored first, thereby achieving synchronized jump between audio and video. Similarly, when the user performs a score selection operation on the second display end (such as clicking on a certain measure of the score), the score advancement data or score positioning data at the corresponding position and its second timestamp are determined, and the hardware data-driven sound structure with the same timestamp is called, and the sound output at that moment is restored first, thereby achieving consistency between the sound and the selected score position. In this way, the original playback sequence is broken and interactive jump is achieved, triggered by user operation, so that heterogeneous multi-category data remains time-aligned after manual adjustment.
[0100] The progress adjustment operation refers to the user's operation of selecting the video playback progress on the first display terminal, such as manually sliding or clicking the progress bar or voice control to jump to a specified time point. The score selection operation refers to the user's operation of selecting a music score segment to be performed on the second display terminal, such as clicking a phrase or measure to locate the performance position.
[0101] In some embodiments, the method also includes: obtaining first performance data, second performance data, and third performance data respectively based on the hardware data, the video data, and the music score advancement data at the same timestamp; when the first performance data is different from the second performance data and the third performance data, generating correction hardware data based on the second performance data and the third performance data; and presenting it on the instrument end based on the correction hardware data.
[0102] Specifically, after receiving the hardware data, video data, and music score advancement data at the same time stamp, the corresponding first performance data, second performance data, and third performance data are parsed respectively. The performance data may be data such as specific pitch and duration.
[0103] For example, the first performance data is directly parsed from hardware data, for example, by mapping the ID, displacement, force, and velocity of the pressed or released key to directly capture the performance data. The second performance data can be obtained through visual recognition analysis of video data, using image processing techniques to detect the finger press position and key movement trajectory, and then identifying the performance data fed back from the piano keyboard layout. The third performance data is based on the music score segment pointed to by the music score progression data, determining the performance data expected to be played at that moment. The second and third performance data can represent ranges of data values rather than single, fixed values.
[0104] If the first performance data differs from the second and third performance data, and the second and third performance data are sufficiently similar, there may be an acquisition or transmission error in the hardware data. The video and music score data, acting as auxiliary information, provide reference for the performance content. In this case, correction hardware data is generated based on the common data between the second and third performance data. If there are slight differences between the second and third performance data, the second performance data determined by the video data is prioritized, and correction hardware data is generated for correction and sent to the instrument for sound presentation.
[0105] This mechanism is enabled when hardware anomalies occur, and uses data cross-validation to achieve fault-tolerant recovery, ensuring the accuracy and continuity of performance content and improving system robustness.
[0106] In some embodiments, an embodiment of the present application provides a schematic flowchart of another data alignment method in remote performance, the method including S201 to S204: S201, based on the first device end, collecting hardware data and video data during the performance, the hardware data and the video data are associated with a timestamp; and sending the hardware data and the video data to the second device end; wherein the second device end includes a musical instrument end and a first display end; S202, controlling the musical instrument end to parse and present the hardware data; controlling the first display end to parse and present the video data; S203, extracting the timestamps of the hardware data and video data presented at the same moment, and calculating the first time deviation between the two timestamps; S204, when the first time deviation is greater than or equal to the first deviation threshold, based on the target timestamp of the hardware data to be presented at the next moment, calling the target video data of the target timestamp, and controlling the first display end to present the target video data at the next moment.
[0107] The present application also provides a remote performance system, which is applied to the data alignment method for remote performance provided in any embodiment of the present application; the system includes: The first device end is used to collect hardware data, video data, and music score advancement data during the performance and send them to the second device end, wherein the hardware data, the video data, and the music score advancement data are associated with a timestamp; The second device end includes a musical instrument end and a first display end and a second display end; the musical instrument end is used to parse and present the hardware data; the first display end is used to parse and present the video data; the second display end is used to determine the score location data of the current hardware data based on a preset score alignment algorithm and a number of hardware data and preset performance scores received within a preset time period of the current hardware data, wherein the score location data is associated with a timestamp of the current hardware data; when the current hardware data is presented on the musical instrument end, the score segment pointed to by the score location data or the score advancement data is presented; The control unit is configured to extract the timestamps of the hardware data and the video data presented at the same moment and calculate a first time offset between the two timestamps; when the first time offset is greater than or equal to a first offset threshold, call the target video data of the target timestamp based on the target timestamp of the hardware data to be presented at the next moment, and control the first display end to present the target video data at the next moment; The control unit is also used to, when it is detected that the computing power of the device is sufficient, control the second display end to synchronously present the music score segment pointed to by the music score positioning data; or, when it is detected that the computing power of the device is insufficient, compare the position deviation of the music score segment pointed to by the music score positioning data and the music score advancement data at the same timestamp based on a preset period; select the music score segment pointed to by the music score positioning data or the music score advancement data according to the position deviation, and control the second device end to present the corresponding music score segment.
[0108] The embodiments of the present application provide a remote performance system that aims to achieve cross-device, high-precision collaborative presentation of audio, video and music scores, and enhance the immersion and interactivity in scenarios such as remote teaching and virtual ensemble.
[0109] The system includes a first device end, a second device end, and a control unit. The first device end is used to collect hardware data (such as key status and strength) and video data (performance action images) during the performance, generate score advancement data, and stamp the three with a high-precision unified timestamp, and then send the data to the second device end. The second device end includes a musical instrument end, a first display end, and a second display end: the musical instrument end parses the hardware data and drives the sound structure to restore the sound; the first display end is responsible for playing the video data and presenting the performance picture; the second display end generates the score positioning data of the current performance position based on the continuous hardware data received through the score alignment algorithm. The score positioning data uses the unified timestamp of the hardware data synchronously, and displays the corresponding score fragment synchronously when the hardware data sounds, realizing the precise linkage between notes and scores.
[0110] The control unit continuously monitors the state of audio and video synchronization, extracts the timestamps of the hardware and video data presented at the same time, and calculates the first time deviation. When the deviation exceeds the first deviation threshold, it uses the target timestamp of the hardware data as the basis, calls the target video frame with the corresponding timestamp, and instructs the first display terminal to directly present it, achieving frame skipping alignment. This effectively addresses audio and video synchronization issues caused by insufficient terminal computing power or network fluctuations, ensuring the real-time and audio-visual consistency of remote performances. Furthermore, the control unit also compares and selects the score advancement data and score positioning data when the device computing power is insufficient.
[0111] Exemplarily, the control unit is also used to implement the steps of the data alignment method in remote performance provided in any embodiment of the present application, which will not be repeated here.
[0112] An embodiment of the present application provides a computer device, which may be a terminal device or a server. Exemplarily, the above method may be implemented in the form of a computer program, which may be run on the computer device.
[0113] The computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0114] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any data alignment method in remote performance.
[0115] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0116] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any data alignment method in the remote performance.
[0117] This network interface is used for network communication, such as sending assigned tasks.
[0118] It should be understood that the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0119] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps: S101, collecting hardware data and video data during a performance based on a first device end, and sending the data to a second device end, wherein the hardware data and the video data are associated with a timestamp; the second device end includes a musical instrument end and a first display end and a second display end; S102, controlling the instrument end to parse and present the hardware data; controlling the first display end to parse and present the video data; S103 determines score location data of the current hardware data based on a preset score alignment algorithm and a plurality of hardware data and a preset performance score received within a preset time period of the current hardware data, wherein the score location data is associated with a timestamp of the current hardware data; Depending on the computing power of the device, choose to execute step S1031 or execute steps S1032 to S1034: S1031, when the current hardware data is presented on the musical instrument end, controlling the second display end to synchronously present the music score segment pointed to by the music score positioning data; S1032: Collecting music score advancement data during the performance based on the first device, and sending the data to the second device, wherein the music score advancement data is associated with a timestamp; S1033, based on a preset period, comparing the position deviation of the music score segment pointed to by the music score positioning data and the music score advancement data at the same timestamp; S1034, when the current hardware data is presented on the instrument end, selecting the music score segment pointed to by the music score positioning data or the music score advancement data according to the position deviation, and controlling the second device end to present the corresponding music score segment; S1041, extracting timestamps of hardware data and video data presented at the same time, and calculating a first time offset between the two timestamps; S1042: When the first time deviation is greater than or equal to a first deviation threshold, based on the target timestamp of the hardware data to be presented at the next moment, call the target video data of the target timestamp, and control the first display end to present the target video data at the next moment.
[0120] In one embodiment, the processor is configured to execute a computer program stored in the memory to implement the following steps: S201, collecting hardware data and video data during a performance based on a first device, wherein the hardware data and the video data are associated with a timestamp; and sending the hardware data and the video data to a second device; wherein the second device includes a musical instrument and a first display. S202, controlling the instrument end to parse and present the hardware data; controlling the first display end to parse and present the video data; S203, extracting the timestamps of the hardware data and the video data presented at the same time, and calculating a first time offset between the two timestamps; S204 , when the first time deviation is greater than or equal to a first deviation threshold, calling target video data with a target timestamp based on the target timestamp of the hardware data to be presented at the next moment, and controlling the first display end to present the target video data at the next moment.
[0121] Exemplarily, the processor is used to run a computer program stored in the memory, and is also used to implement the steps of the data alignment method in remote performance provided in any embodiment of the present application, which will not be repeated here.
[0122] A computer-readable storage medium is also provided in an embodiment of the present application. The computer-readable storage medium stores a computer program, which includes program instructions. The processor executes the program instructions to implement the steps of any one of the data alignment methods in remote performance provided in the embodiments of the present application. The computer-readable storage medium can be a product program.
[0123] The computer-readable storage medium may be an internal storage unit of the computer device described in the aforementioned embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a SmartMedia Card (SMC), a Secure Digital (SD) card, a flash memory card, etc., equipped on the computer device.
[0124] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A data alignment method for remote performance, characterized in that: The method comprises: S101, collecting hardware data and video data during a performance based on a first device end, and sending the data to a second device end, wherein the hardware data and the video data are associated with a timestamp; the second device end includes a musical instrument end and a first display end and a second display end; S102, controlling the instrument end to parse and present the hardware data; controlling the first display end to parse and present the video data; S103 determines score location data of the current hardware data based on a preset score alignment algorithm and a plurality of hardware data and a preset performance score received within a preset time period of the current hardware data, wherein the score location data is associated with a timestamp of the current hardware data; Depending on the computing power of the device, choose to execute step S1031 or execute steps S1032 to S1034: S1031, when the current hardware data is presented on the musical instrument end, controlling the second display end to synchronously present the music score segment pointed to by the music score positioning data; S1032: Collecting music score advancement data during the performance based on the first device, and sending the data to the second device, wherein the music score advancement data is associated with a timestamp; S1033, based on a preset period, comparing the position deviation of the music score segment pointed to by the music score positioning data and the music score advancement data at the same timestamp; S1034, when the current hardware data is presented on the instrument end, selecting the music score segment pointed to by the music score positioning data or the music score advancement data according to the position deviation, and controlling the second device end to present the corresponding music score segment; S1041, extracting timestamps of hardware data and video data presented at the same time, and calculating a first time offset between the two timestamps; S1042: When the first time deviation is greater than or equal to a first deviation threshold, based on the target timestamp of the hardware data to be presented at the next moment, call the target video data of the target timestamp, and control the first display end to present the target video data at the next moment.
2. The method according to claim 1, wherein The preset time period includes a first time period before and a second time period after the timestamp of the current hardware data, wherein the length of the first time period is a first preset duration, and the length of the second time period is a second preset duration.
3. The method according to claim 1 or 2, wherein: The S1034 includes: When the position deviation is less than a preset difference, the second music score segment pointed to by the music score advancement data is presented on the second device end; or When the position deviation is greater than or equal to the preset difference, the first music score segment pointed to by the music score positioning data is presented on the second device end.
4. The method according to claim 3, wherein The instrument end is used to drive the sound structure to execute the hardware data; the first display end is used to play the video data; and the second display end is used to display the music score fragment corresponding to the music score advancement data and / or music score positioning data.
5. The method according to claim 1, wherein The method further comprises: When the first time deviation is greater than the second deviation threshold and less than the first deviation threshold, analyzing the presentation time difference between the hardware data and the video data at the same timestamp according to the first time deviation within the preset time; The presentation speed of the video data is adjusted based on the presentation time difference so that the first time deviation remains smaller than the first deviation threshold.
6. The method according to claim 1, wherein The method further comprises: In response to a progress adjustment operation by a user on the first display terminal, determining video data to be displayed, and based on a first timestamp of the video data to be displayed, calling hardware data of the first timestamp for presentation; and / or, In response to the user's music score selection operation on the second display end, the music score advancement data or music score positioning data to be displayed is determined, and based on the second timestamp of the music score advancement data or music score positioning data to be displayed, the hardware data of the second timestamp is called for presentation.
7. The method according to claim 3, wherein The method further comprises: obtaining first performance data, second performance data, and third performance data respectively according to the hardware data, the video data, and the music score advancement data at the same time stamp; generating correction hardware data based on the second performance data and the third performance data when the first performance data is different from both the second performance data and the third performance data; Presentation is performed on the instrument end based on the correction hardware data.
8. The method according to claim 1, wherein The method further comprises: Based on the timestamp, the hardware data and the video data with the same timestamp are encapsulated into one data packet; or based on the timestamp, the hardware data, the video data, and the music score advancement data with the same timestamp are encapsulated into one data packet; Distribute the encapsulated data packets over the network.
9. A remote performance system, characterized in that: The system is applied to the data alignment method in remote performance according to any one of claims 1 to 8; the system comprises: The first device end is used to collect hardware data, video data, and music score advancement data during the performance and send them to the second device end, wherein the hardware data, the video data, and the music score advancement data are associated with a timestamp; The second device end includes a musical instrument end and a first display end and a second display end; the musical instrument end is used to parse and present the hardware data; the first display end is used to parse and present the video data; the second display end is used to determine the score location data of the current hardware data based on a preset score alignment algorithm and a number of hardware data and preset performance scores received within a preset time period of the current hardware data, wherein the score location data is associated with a timestamp of the current hardware data; when the current hardware data is presented on the musical instrument end, the score segment pointed to by the score location data or the score advancement data is presented; The control unit is configured to extract the timestamps of the hardware data and the video data presented at the same moment and calculate a first time offset between the two timestamps; when the first time offset is greater than or equal to a first offset threshold, call the target video data of the target timestamp based on the target timestamp of the hardware data to be presented at the next moment, and control the first display end to present the target video data at the next moment; The control unit is also used to, when it is detected that the computing power of the device is sufficient, control the second display end to synchronously present the music score segment pointed to by the music score positioning data; or, when it is detected that the computing power of the device is insufficient, compare the position deviation of the music score segment pointed to by the music score positioning data and the music score advancement data at the same timestamp based on a preset period; select the music score segment pointed to by the music score positioning data or the music score advancement data according to the position deviation, and control the second device end to present the corresponding music score segment.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the data alignment method in remote performance according to any one of claims 1 to 8.
Citation Information
Patent Citations
Live broadcast recording and broadcasting method for synchronously transmitting MIDI based on RTMP protocol
CN110392276A
Live broadcast teaching method based on IOT piano
CN116939237A
Multifunctional synchronous interaction system and method of music instruments
CN103729062A
Musical instrument playing key position prompting method and device, electronic equipment and storage medium
CN112818981A
Playing process and music score synchronous display method, device and equipment and storage medium
CN113286183A
Cited By
Fingerprint identification method, fingering identification system and fingering identification equipment applied to musical instrument system, and medium
CN121459429A