A multi-modal cooperative control method and system

By generating beat pulse signals and synchronizing them with an electronic stopwatch, and dynamically routing audio signals for mixing, the problem of coordinated control of data signals in music teaching is solved, achieving precise matching and automated data processing, thereby improving training effectiveness and user experience.

CN122120905APending Publication Date: 2026-05-29SHENZHEN HONGSHENGDA ELECTRONIC TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN HONGSHENGDA ELECTRONIC TECH CO LTD
Filing Date
2026-03-03
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

In rhythm training scenarios in music teaching, the beat pulse, timing start, and various data signals of audio teaching are not effectively coordinated and controlled, resulting in difficulties in timing matching and data association storage, which affects the quantitative evaluation of training effects and user experience.

Method used

By receiving encrypted instructions, parsing the beat rhythm type and timing duration information, generating beat pulse signals, aligning the electronic stopwatch timing start point with the beat pulses, dynamically routing Bluetooth audio and teaching voice signals for mixing, recording the time correspondence, generating synchronized timeline data, and outputting training reports.

Benefits of technology

It achieves precise matching of audio signals and beat pulses, automatically collects and records training data, improves the automation efficiency of data processing and user experience, and adapts to the actual needs of music teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122120905A_ABST
    Figure CN122120905A_ABST
Patent Text Reader

Abstract

The application provides a multimodal collaborative control method and system, and relates to the technical field of data processing.The method comprises the following steps: starting an electronic stopwatch according to timing duration information, taking a beat pulse signal at the starting moment as a synchronization reference, aligning the timing starting point of the electronic stopwatch with the beat pulse, and obtaining synchronized stopwatch timing data; during the timing of the electronic stopwatch, dynamically routing and mixing a Bluetooth audio signal and a teaching voice signal based on the beat pulse signal and the stopwatch timing data, generating a mixed audio signal, synchronously recording the time correspondence between the stopwatch timing data and the beat pulse signal, and obtaining synchronization time axis data; after the timing of the electronic stopwatch ends, generating a training report according to the synchronization time axis data, and outputting the training report through voice.The application improves the automation degree, synchronization accuracy and user experience of multimodal collaborative control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a multimodal collaborative control method and system. Background Technology

[0002] In rhythm training scenarios in music education, beat pulses, timing start signals, and various audio teaching data signals are processed independently. Some lack an effective timing coordination control architecture, leading to limitations in timing matching, linked output, and data association storage for multiple signal types. Specifically, the generation of beat pulse signals and the triggering of timing start signals lack hardware-level synchronous triggering logic; their start commands must be issued independently, potentially failing to achieve microsecond-level precise alignment of signal starting points and easily resulting in uncontrollable timing deviations. The output of audio teaching signals lacks a linked routing configuration mechanism with beat pulse signals, making it impossible to precisely match the playback rhythm of audio signals with the beat pulse signals. Furthermore, during training, the timestamp data generated by timing and the pulse signal data output by beat cannot be automatically collected by the system and a one-to-one correspondence cannot be established; data processing and integration must rely on external manual methods. This discrete signal processing mode prevents the formation of a precise and complete time-rhythm correspondence data chain during training, making it difficult to objectively and systematically quantify the rhythm training effect based on this data. This, to some extent, limits the application experience and practical effectiveness of music rhythm training technologies. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a multimodal cooperative control method and system, which improves the automation level, synchronization accuracy and user experience of multimodal cooperative control.

[0004] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0005] Firstly, a multimodal cooperative control method, the method comprising:

[0006] Step 1: Receive the encrypted command sent by the host and obtain the original command data;

[0007] Step 2: Parse the raw instruction data to obtain the beat rhythm type and timing duration information;

[0008] Step 3: Generate the corresponding beat pulse signal according to the beat rhythm type;

[0009] Step 4: Based on the timing duration information, start the electronic stopwatch and use the beat pulse signal at the moment of start as the synchronization reference to align the starting point of the electronic stopwatch with the beat pulse, thereby obtaining the synchronized stopwatch timing data.

[0010] Step 5: During the electronic stopwatch timing, based on the beat pulse signal and stopwatch timing data, the Bluetooth audio signal and the teaching voice signal are dynamically routed and mixed to generate a mixed audio signal, and the time correspondence between the stopwatch timing data and the beat pulse signal is recorded synchronously to obtain synchronous time axis data.

[0011] Step 6: After the electronic stopwatch finishes timing, generate a training report based on the synchronized timeline data, and output the training report via voice.

[0012] Secondly, a multimodal cooperative control system includes:

[0013] The acquisition module is used to receive encrypted commands sent by the host and obtain the raw command data;

[0014] The parsing module is used to parse the raw instruction data to obtain the beat rhythm type and timing duration information;

[0015] The processing module is used to generate corresponding beat pulse signals according to the beat rhythm type;

[0016] The synchronization module is used to start the electronic stopwatch based on the timing duration information and use the beat pulse signal at the moment of start as the synchronization reference to align the starting point of the electronic stopwatch with the beat pulse, so as to obtain the synchronized stopwatch timing data.

[0017] The compensation module is used to dynamically route and mix the Bluetooth audio signal and the teaching voice signal based on the beat pulse signal and stopwatch timing data during the electronic stopwatch timing period, generate a mixed audio signal, and synchronously record the time correspondence between the stopwatch timing data and the beat pulse signal to obtain synchronous time axis data.

[0018] The output module is used to generate a training report based on the synchronized timeline data after the electronic stopwatch has finished timing, and output the training report via voice.

[0019] Thirdly, a computing device, comprising:

[0020] One or more processors;

[0021] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0022] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0023] The above-described solution of the present invention has at least the following beneficial effects:

[0024] A linkage processing mechanism for beat pulse signals and audio signals was established. Based on the beat pulse signals and stopwatch timing data, the Bluetooth audio signals and teaching voice signals were dynamically routed and mixed to ensure that the audio output rhythm was precisely matched with the beat pulses, and that the audio teaching signals were highly consistent with the beat rhythm, thus meeting the actual needs of music teaching rhythm training. The mechanism also enabled the automatic collection, association, and recording of data during the training process, synchronously storing the time correspondence between stopwatch timing data and beat pulse signals and forming standardized synchronous timeline data. This eliminated the need for manual post-processing and integration, improving the automation efficiency of data processing. Attached Figure Description

[0025] Figure 1 This is a schematic flowchart of a multimodal cooperative control method provided by an embodiment of the present invention.

[0026] Figure 2 This is a schematic diagram of a multimodal cooperative control system provided by an embodiment of the present invention. Detailed Implementation

[0027] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0028] like Figure 1 As shown, an embodiment of the present invention proposes a multimodal cooperative control method, the method comprising the following steps:

[0029] Step 1: Receive the encrypted command sent by the host and obtain the original command data;

[0030] Step 2: Parse the raw instruction data to obtain the beat rhythm type and timing duration information;

[0031] Step 3: Generate the corresponding beat pulse signal according to the beat rhythm type;

[0032] Step 4: Based on the timing duration information, start the electronic stopwatch and use the beat pulse signal at the moment of start as the synchronization reference to align the starting point of the electronic stopwatch with the beat pulse, thereby obtaining the synchronized stopwatch timing data.

[0033] Step 5: During the electronic stopwatch timing, based on the beat pulse signal and stopwatch timing data, the Bluetooth audio signal and the teaching voice signal are dynamically routed and mixed to generate a mixed audio signal, and the time correspondence between the stopwatch timing data and the beat pulse signal is recorded synchronously to obtain synchronous time axis data.

[0034] Step 6: After the electronic stopwatch finishes timing, generate a training report based on the synchronized timeline data, and output the training report via voice.

[0035] In this embodiment of the invention, a linkage processing mechanism between the beat pulse signal and the audio signal is established. Based on the beat pulse signal and stopwatch timing data, the Bluetooth audio signal and the teaching voice signal are dynamically routed and mixed, so that the audio output rhythm is precisely matched with the beat pulse, and the audio teaching signal is highly consistent with the beat rhythm, which is suitable for the actual needs of music teaching rhythm training. It realizes the automatic collection, association and recording of data during the training process, and synchronously retains the time correspondence between the stopwatch timing data and the beat pulse signal to form a standardized synchronous time axis data, which eliminates the need for manual post-processing and integration, and improves the automation efficiency of data processing.

[0036] In a preferred embodiment of the present invention, step 1, receiving the encryption instruction sent by the host to obtain the original instruction data, may include:

[0037] Step 101: Receive encrypted commands sent by the host via wireless communication to obtain an encrypted command data stream; decrypt the encrypted command data stream to obtain the decrypted command text. Specifically, this includes: establishing a fixed-frequency wireless communication link between the device and the teaching host using a quad-mode 76 or 868MHz wireless communication unit connected to the device's dedicated main control. This communication unit continuously listens for encrypted command transmission signals sent by the host. Upon detecting an encrypted command transmission request from the host, a stable fixed-frequency wireless data transmission channel is established. Following the transmission rules defined by the device based on the fixed-frequency communication specification, which specify the baud rate, data frame format, and data verification method, the device receives continuously sent encrypted command data from the host. The received discrete encrypted instruction data is integrated according to the chronological order of transmission to form a complete and continuous encrypted instruction data stream, ensuring that there is no packet loss or error during data transmission. The device's multi-core coprocessor calls the decryption logic, which is burned into the chip in the form of hardware logic circuits, and uses the unique hardware identification code configured at the device's factory as the decryption key. This hardware identification code is a unique serial code written during the device's production stage. The integrated encrypted instruction data stream is decrypted byte by byte. The decryption process strictly follows the original segmentation order of the encrypted instruction data stream, and restores the plaintext instruction content corresponding to each segment of encrypted data in turn. Then, all the decrypted plaintext instruction content is spliced ​​and integrated according to the original transmission order to form a complete decrypted instruction text.

[0038] Step 102 involves performing an integrity check on the decrypted instruction text. Once the check passes, it is used as the original instruction data. Specifically, this includes: the device's multi-core coprocessor extracting data features from the decrypted instruction text. The extracted features include the total number of characters in the instruction text, the fixed identifier characters for teaching and timing instructions, and their positions. The fixed identifier characters for teaching instructions are unique character combinations representing rhythm teaching instructions, and the fixed identifier characters for timing instructions are unique character combinations representing timing instructions. The character positions are the specific byte positions of the identifier characters within the instruction text. Simultaneously, a feature information verification request is initiated to the host via the four-mode radio frequency wireless communication unit, and the corresponding feature information of the original plaintext instruction returned by the host is received. The multi-core coprocessor then processes the data from the device side... The decrypted instruction text feature information is compared item by item with the original plaintext instruction feature information fed back by the host. First, the total number of characters is compared to see if they are exactly the same. Then, the existence and specific position of the fixed identifier characters for teaching and timing instructions are compared to see if they are consistent. If all the comparison items are completely matched, the integrity verification of the decrypted instruction text is determined to be passed, and the decrypted instruction text is directly used as the original instruction data for subsequent instruction parsing and processing. If any comparison item does not match, the integrity verification is determined to be failed. The four-mode radio frequency wireless communication unit immediately sends an instruction retransmission request to the host and re-executes the encryption instruction reception and decryption operation in step 101 until the received encryption instruction decrypted instruction text passes the integrity verification.

[0039] In this embodiment, the device's four-mode radio frequency wireless communication unit enables stable reception of encrypted commands at a fixed frequency. Clear transmission rules ensure the stability of command data transmission in teaching scenarios, and the decryption method using the device's unique hardware identification code as the key enhances the security of command transmission.

[0040] In a preferred embodiment of the present invention, step 2, parsing the original instruction data to obtain the beat rhythm type and timing duration information, may include:

[0041] Step 201 involves identifying the teaching instruction identifier and timing instruction identifier in the original instruction data, and segmenting the original instruction data into a teaching instruction part and a timing instruction part. Specifically, the device's multi-core coprocessor scans and identifies the original instruction data that has passed integrity verification character by character. According to the instruction encoding rules formulated based on the transmission requirements of teaching instructions, which clearly define the specific character combinations of the teaching instruction identifier and the timing instruction identifier and their arrangement specifications in the instruction data, the device searches for the teaching instruction identifier used to mark beat-related instructions and the timing instruction identifier used to mark timing-related instructions in the original instruction data. After scanning and identifying the two exclusive identifiers, the device accurately locates the specific character positions of the teaching instruction identifier and the timing instruction identifier in the original instruction data. Using the specific character positions of the two identifiers as the clear segmentation basis, the multi-core coprocessor divides all content from the teaching instruction identifier to the timing instruction identifier in the original instruction data into the teaching instruction part, and divides all content after the timing instruction identifier in the original instruction data into the timing instruction part. This completes the accurate segmentation processing of the original instruction data, ensuring that the segmented teaching instruction part only contains beat-related instruction information, and the timing instruction part only contains instruction information related to training time duration.

[0042] Step 202: Parse the beat speed parameter and beat type parameter from the teaching instruction section, and combine the beat speed parameter and beat type parameter into a beat rhythm type; parse the timing duration parameter from the timing instruction section, and convert the timing duration parameter into timing duration information in a preset format. Specifically, this includes: the device's multi-core coprocessor extracting core parameters from the divided teaching instruction section, and according to the teaching instruction parameter encoding rules formulated based on the transmission requirements of beat teaching parameters, the beat speed parameter in the teaching instruction is a fixed-byte numeric parameter with a character length of three digits, represented as beats per minute, with an accuracy of ±0.1 BPM; the beat type parameter in the teaching instruction is a character parameter following the beat speed parameter with a character length of three digits, represented as a fractional beat pattern. The teaching instruction section accurately identifies and extracts the beat speed parameter (beats per minute) used to represent the tempo, and the beat type parameter (3 / 4, 6 / 8) used to represent the composition of beat measures. After extracting the beat pattern from the composite beat, the beat speed parameters and beat type parameters are combined in an orderly manner. The overall information after combination is the beat rhythm type, which can completely represent the beat pattern required for teaching. The multi-core coprocessor extracts the core parameters of the divided timing instruction part. According to the timing instruction parameter encoding rules based on the timing duration parameter transmission requirements, the timing duration parameter in the timing instruction is a fixed byte number combination parameter with a character length of six digits, representing the total training time in the order of hours, minutes, and seconds. The timing duration parameter used to represent the total training time is accurately identified and extracted in the timing instruction part. After extraction, according to the timing operation rules of the temperature-compensated RTC stopwatch built into the device, the timing data input format of the stopwatch is a six-digit number combination of hours, minutes, and seconds, with a timing accuracy of ±2ppm. The extracted timing duration parameter is converted into timing duration information in the preset format of hours, minutes, and seconds that can be directly recognized and executed by the temperature-compensated RTC stopwatch.

[0043] This embodiment, relying on clear parameter encoding rules, achieves accurate extraction of core parameters in teaching and timing instructions. The orderly combination of beat rhythm types provides an accurate basis for the generation of subsequent beat pulse signals. According to the clear timing operation rules of the temperature-compensated RTC stopwatch, the format conversion of the timing duration parameter is completed, allowing the stopwatch to directly recognize and execute timing operations, eliminating additional format conversion steps and improving the instruction execution efficiency of the device.

[0044] In a preferred embodiment of the present invention, step 3, generating a corresponding beat pulse signal according to the beat rhythm type, may include:

[0045] Step 301: Based on the beat type parameter, read the corresponding rhythm pattern data from the pre-stored non-volatile memory to obtain the strong and weak beat distribution pattern within each measure. Specifically, this includes: the device's multi-core coprocessor receiving the beat type parameter parsed in step 202, using this beat type parameter as the sole data retrieval basis, and initiating a rhythm pattern data read request to the device's built-in non-volatile memory; this non-volatile memory is a supporting storage unit for the device's digital beat generator, and its construction process is as follows: before the device leaves the factory, according to the frequency of beat usage in music teaching scenarios, the composite rhythm pattern data such as 3 / 4 beats and 6 / 8 beats are standardized and encoded. The encoding format is that the dynamic feature of each beat corresponds to an 8-bit binary value, where the strong beat is encoded as 0xFF, the secondary strong beat as 0x80, and the weak beat as 0x40. Then, the encoded rhythm pattern data of each beat type are stored in the designated address range of the memory according to a fixed data block size. The 3 / 4 beat rhythm pattern data is stored in a dedicated fixed address segment, and the 6 / 8 beat rhythm pattern data is stored in another independent fixed address segment. At the same time, the beat type parameter and the storage address of the corresponding rhythm pattern data are... The non-volatile memory is pre-built by binding the rhythm data into a relational data set and storing it in a dedicated address segment of the memory. This relational data records the one-to-one correspondence between the character values ​​of each beat type parameter and the corresponding storage address, thus completing the pre-storage construction of the non-volatile memory. All pre-stored rhythm data are classified and stored according to the beat type parameters, and a one-to-one correspondence retrieval relationship is established between the beat type parameters and the rhythm data. This retrieval relationship uses the beat type parameters as a retrieval index and binds them to the fixed storage address of the corresponding rhythm data in the memory. After receiving a data read request, the non-volatile memory first retrieves the storage address that matches the received beat type parameter from the relational data in the dedicated address segment of the memory. Then, it quickly reads the corresponding rhythm data based on that address. After a successful search, the rhythm data is transmitted to the multi-core coprocessor via the hardware bus. The multi-core coprocessor performs segment-by-segment parsing of the received rhythm data, restoring the 8-bit binary value to the corresponding dynamic feature. It extracts the core information representing the dynamic feature of each beat within a measure from the rhythm data. This information is the strong and weak beat distribution pattern within each measure.

[0046] Step 302: Calculate the time interval of each beat pulse according to the beat speed parameter to obtain the basic pulse period; generate a continuous beat pulse signal according to the strong-weak beat distribution pattern and the basic pulse period. The timing of each pulse in the beat pulse signal is determined by the basic pulse period, and the amplitude characteristic of each pulse is determined by the strong-weak attribute at the corresponding position in the strong-weak beat distribution pattern. Specifically, the multi-core coprocessor of the device receives the beat speed parameter parsed in Step 202, i.e., the number of beats per minute. Taking sixty seconds as the basic time value for calculation, divide the basic time value of sixty seconds by the beat speed parameter, i.e., the number of beats per minute. The result obtained through this division is the time interval of each beat pulse, and this time interval is the basic pulse period for generating the beat pulse signal. For example, if the beat speed parameter is one hundred and twenty beats per minute, that is, sixty seconds divided by one hundred and twenty, the calculated basic pulse period is 0.5 seconds, and this period can accurately represent the time difference between two adjacent beat pulses. The multi-core coprocessor receives the strong-weak beat distribution pattern obtained in Step 301. Taking the calculated basic pulse period as the only timing basis, set the hardware trigger moment of each beat pulse at a fixed time interval according to this period in sequence to determine the sequential trigger timing of all beat pulses. At the same time, taking the strong-weak beat distribution pattern as the only amplitude basis, match the strong-weak attribute at each position in the strong-weak beat distribution pattern with the beat pulses set in sequence one by one, and set the electrical signal amplitude characteristic of each beat pulse according to the strong-weak attribute at the corresponding position. The beat pulse corresponding to the strong beat is set with a high-amplitude electrical signal characteristic, the beat pulse corresponding to the weak beat is set with a low-amplitude electrical signal characteristic, and the beat pulse corresponding to the sub-strong beat is set with a medium-amplitude electrical signal characteristic. The digital beat generator of the device generates a continuous beat pulse electrical signal in sequence according to the trigger timing and electrical signal amplitude characteristic set by the multi-core coprocessor.

[0047] In this embodiment, relying on the classified storage of the non-volatile memory supporting the digital beat generator and the retrieval relationship bound with the beat type parameter, it realizes the fast and accurate reading of the rhythm data matching the beat type parameter. The segmented parsing of the rhythm data can extract the accurate strong-weak beat distribution pattern, ensuring that the generated beat pulse signal can conform to the strong-weak rules of different beat types.

[0048] In a preferred embodiment of the present invention, Step 4: According to the timing duration information, start an electronic stopwatch, and use the beat pulse signal at the moment of startup as the synchronization reference to align the timing start point of the electronic stopwatch with the beat pulse to obtain the synchronized stopwatch timing data, which may include:

[0049] Step 401: Set the timing endpoint of the electronic stopwatch based on the timing duration information to obtain the timing endpoint parameter; monitor the beat pulse signal based on the timing endpoint parameter, and when the first pulse edge is detected, use the trigger time of the detected pulse edge as the timing start point of the electronic stopwatch, and record the beat number corresponding to the detected pulse edge to obtain the synchronization start time and the starting beat mark. Specifically, the device's multi-core coprocessor receives the converted hour, minute, and second format timing duration information, extracts the hour, minute, and second values ​​from the information, adds these three sets of values ​​to the initial timing value of zero of the device's built-in temperature-compensated RTC stopwatch, and obtains the timing endpoint parameter of the electronic stopwatch. This parameter is the target value for the temperature-compensated RTC stopwatch to complete the timing operation. The multi-core coprocessor writes the timing endpoint parameter into the timing of the temperature-compensated RTC stopwatch. The timing configuration register is used to set the timing endpoint of the electronic stopwatch. The timing accuracy of this temperature-compensated RTC stopwatch is ±2ppm. Based on the written timing endpoint parameters, the multi-core coprocessor sends a beat pulse signal monitoring command to the digital beat generator, starting continuous monitoring of the beat pulse signal output by the digital beat generator. The beat accuracy of the digital beat generator is ±0.1BPM. During the monitoring process, the rising edge of the beat pulse signal is identified by hardware-level detection to ensure that the time difference between the timing start point and the beat pulse is less than 1 millisecond. When the rising edge of the first beat pulse signal is detected for the first time, the hardware trigger time of the rising edge is used as the timing start point of the electronic stopwatch. At the same time, the beat number corresponding to the first pulse edge is set as the initial sequence number one. The detected trigger time and the initial sequence number one are associated and bound to form a synchronous start time and the starting beat mark.

[0050] Step 402: The electronic stopwatch is started based on the synchronization start time and the initial beat marker. During the timing process, the current beat pulse number corresponding to each timing moment is determined based on the continuous pulse sequence of the beat pulse signal. The time value of each timing moment is associated with and stored with the corresponding current beat pulse number to obtain the synchronized stopwatch timing data. Specifically, the device's multi-core coprocessor sends the synchronization start time and the initial beat marker to the temperature-compensated RTC stopwatch, triggering the temperature-compensated RTC stopwatch to perform timing operations with a timing accuracy of ±2ppm starting from the synchronization start time. Throughout the entire timing process of the electronic stopwatch, the multi-core coprocessor continuously receives the continuous beat pulse sequence output by the digital beat generator, which supports three or four beats. The system outputs composite rhythm patterns such as 6 / 8 beats. Each time a new beat pulse edge is detected, the initial sequence number in the starting beat marker is incremented by one to obtain the current beat pulse sequence number corresponding to each beat pulse edge. Simultaneously, the real-time timing value of the electronic stopwatch at each beat pulse edge trigger moment is extracted, ensuring the synchronization of the extracted time value with the pulse edge trigger moment. The multi-core coprocessor maps the real-time timing value of each timing moment to the corresponding current beat pulse sequence number, storing this correspondence information in chronological order to the device's non-volatile memory. This memory is a supporting storage unit for the digital beat generator, ensuring that each time value has a unique corresponding beat pulse sequence number during storage, ultimately forming synchronized stopwatch timing data.

[0051] This embodiment achieves hardware-level synchronous timing between the electronic stopwatch and the beat pulse signal. By successively calculating the beat pulse sequence number, each timing moment has a corresponding beat pulse sequence number. The associated storage method ensures the correspondence between timing data and beat data. The resulting synchronized stopwatch timing data completely records the synchronization relationship between timing and beat.

[0052] In a preferred embodiment of the present invention, step 5, during the electronic stopwatch timing, dynamically routes and mixes the Bluetooth audio signal and the teaching voice signal based on the beat pulse signal and stopwatch timing data to generate a mixed audio signal, and synchronously records the time correspondence between the stopwatch timing data and the beat pulse signal to obtain synchronized timeline data, may include:

[0053] Step 501: Based on the correspondence between time values ​​and beat pulse numbers stored in the synchronized stopwatch timing data, extract the mapping relationship between time values ​​and beat pulse numbers to obtain beat time mapping data; determine the beat pulse number of the real-time time based on the beat time mapping data, and query the preset audio routing rules based on the beat pulse number of the real-time time to obtain audio mixing ratio parameters. Specifically, this includes: the device's multi-core coprocessor reading the synchronized stopwatch timing data in the non-volatile memory, extracting all corresponding time values ​​and beat pulse numbers from the data, standardizing and organizing the correspondence, removing redundant data to form beat time mapping data, which fully represents the correspondence between time and beat during the electronic stopwatch timing process; the multi-core coprocessor extracting the current real-time time value of the temperature-compensated RTC stopwatch in real time, accurately matching the real-time time value with the time value in the beat time mapping data, and determining the beat pulse number of the real-time time value. The multi-core coprocessor retrieves the audio routing rules and Bluetooth audio priority rules pre-stored in the device's non-volatile memory. The audio routing rules are defined by the real-time audio routing engine of the multi-core coprocessor, specifying that the strong beat pulse sequence number corresponds to a 70% proportion of teaching voice signal and a 30% proportion of Bluetooth audio signal, while the weak beat and secondary strong beat pulse sequence numbers correspond to a 40% proportion of teaching voice signal and a 60% proportion of Bluetooth audio signal. The Bluetooth audio priority rules specify that when Bluetooth call audio input is detected, the proportion of Bluetooth audio signal should be immediately adjusted to 100% and the beat pulse output of the digital beat generator should be paused. The multi-core coprocessor performs a precise query in the audio routing rules based on the beat pulse sequence number of the real-time time obtained by matching, and obtains the corresponding proportion value of Bluetooth audio signal and teaching voice signal. This proportion value is the audio mixing ratio parameter.

[0054] Step 502: Obtain water depth data detected by the water pressure sensor; perform frequency response compensation processing on the Bluetooth audio signal based on the water depth data to obtain a compensated Bluetooth audio signal; perform frequency response compensation processing on the teaching voice signal based on the water depth data to obtain a compensated teaching voice signal; mix and synthesize the compensated Bluetooth audio signal and the compensated teaching voice signal according to the audio mixing ratio parameters to generate a mixed audio signal. Specifically, the device's multi-core coprocessor obtains the water depth data detected by the water pressure sensor in real time through the hardware interface. This data represents the actual water depth of the device's current environment and can accurately detect water depth changes exceeding 0.5 meters. Based on this water depth data, the multi-core coprocessor performs gain adjustment processing on the Bluetooth audio signal transmitted by the Bluetooth 5.3 communication unit according to a preset underwater frequency response compensation rule. This underwater frequency response compensation rule is specifically for the entire audio frequency band from 20 Hz to 20 kHz. When the water depth does not exceed 0.5 meters, the gain of each frequency band remains at 0d. B. When the water depth exceeds 0.5 meters, the gain of the 1000 Hz center frequency band is increased according to a parabolic gain curve. The gain of the other frequency bands is gradually adjusted as the frequency deviates from 1000 Hz to complete the frequency response compensation of the Bluetooth audio signal, thus obtaining the compensated Bluetooth audio signal. At the same time, according to the underwater frequency response compensation rule, the gain of the teaching voice signal transmitted by the four-mode RF 76 or 868 MHz communication unit is adjusted band by band to complete the frequency response compensation of the teaching voice signal, thus obtaining the compensated teaching voice signal. This ensures that the output distortion of the two audio signals in the underwater environment meets the usage requirements. The multi-core coprocessor extracts the ratio of the Bluetooth audio signal to the teaching voice signal from the audio mixing ratio parameter. The signal amplitude of the compensated Bluetooth audio signal is multiplied by its ratio value, and then the signal amplitude of the compensated teaching voice signal is multiplied by its ratio value. The results of the two multiplication operations are added together, and the amplitude of the added signal is normalized to finally generate the mixed audio signal.

[0055] Step 503: During the generation of the mixed audio signal, the generation time corresponding to each frame of the mixed audio signal is associated and stored with the time value and beat pulse number of the same moment in the synchronized stopwatch timing data to obtain synchronized timeline data. Specifically, this includes: the device's multi-core coprocessor, in the process of generating the mixed audio signal frame by frame according to the audio mixing ratio parameters, relies on the real-time audio routing engine to extract the hardware generation time of each frame of the mixed audio signal in real time, and at the same time retrieves the electronic stopwatch time value that perfectly matches the generation time from the synchronized stopwatch timing data, as well as the beat pulse number corresponding to the time value, to ensure the time synchronization of the three; the multi-core coprocessor associates and binds the generation time of each frame of the mixed audio signal, the stopwatch time value at the same moment, and the corresponding beat pulse number, and stores all the associated and bound information continuously in the device's non-volatile memory according to the generation order of the mixed audio signal. The memory pre-stores at least three rhythmic data, and during the storage process, it ensures that the relevant information of each frame of audio signal is complete and unique, ultimately forming synchronized timeline data.

[0056] This embodiment utilizes water depth data from a water pressure sensor to achieve targeted frequency response compensation for Bluetooth audio signals and teaching voice signals, effectively improving the attenuation and distortion problem of audio signals in underwater environments. Through precise proportional calculation and signal synthesis, it achieves reasonable mixing of the two audio signals, and the generated mixed audio signal meets the output requirements of both Bluetooth audio and teaching voice.

[0057] In a preferred embodiment of the present invention, step 6, after the electronic stopwatch has finished timing, generating a training report based on the synchronized timeline data and outputting the training report via voice, may include:

[0058] Step 601: Extract all stored timestamps, beat pulse numbers, and timing values ​​from the synchronization timeline data to obtain the training process timing data. Specifically, this includes: the device's multi-core coprocessor reading the synchronization timeline data stored in the non-volatile memory, parsing the synchronization timeline data segment by segment according to the data storage order, and extracting from the data all the timestamps of the mixed audio signal generation time, the beat pulse numbers corresponding to the timestamps, and the timing values ​​of the electronic stopwatch at the same time. During the extraction process, the integrity and accuracy of the data are ensured, and no relevant data corresponding to any frame of audio signal is omitted. The multi-core coprocessor then organizes all the extracted timestamps, beat pulse numbers, and timing values ​​in an orderly manner, arranging them continuously according to the chronological order of the timestamps to form complete training process timing data.

[0059] Step 602: Calculate the total training time based on the timing values, statistically analyze the beat following deviation based on the changes in the beat pulse sequence number, and extract the difference between the actual trigger time and the theoretical trigger time for each beat point based on the correspondence between the timestamp and the timing values ​​to obtain beat synchronization accuracy data. The total training time, beat following deviation values, and beat synchronization accuracy data are then summarized into training statistics. Specifically, this includes: the device's multi-core coprocessor extracting the final and initial timing values ​​of the electronic stopwatch from the training process timing data. The timekeeping rules for temperature-compensated RTC stopwatches are as follows: subtract the hours, minutes, and seconds sequentially. First, subtract the seconds value. If the minuend is less than the subtrahend, borrow one from the minutes place, counting as 60 seconds, and then complete the seconds subtraction. The minutes value is then decremented by one. Next, subtract the minutes value. If the minuend is less than the subtrahend, borrow one from the hours place, counting as 60 minutes, and then complete the minutes subtraction. Finally, subtract the hours value. The resulting combination of hour, minute, and second values ​​is the training result. Total duration; The multi-core coprocessor extracts the maximum value of the actually detected beat pulse sequence from the timing data of the training process, subtracts this maximum value from the maximum value of the beat pulse sequence output by the digital beat generator, and takes the absolute value of the result, which is the beat following deviation value. The beat output accuracy of the digital beat generator is ±0.1 BPM; The multi-core coprocessor determines the theoretical trigger time of each beat point based on the beat pulse sequence in the timing data of the training process and the output beat speed of the digital beat generator, and then extracts the actual timestamp corresponding to each beat point as the actual trigger time. The actual trigger time of each beat point is subtracted from its theoretical trigger time to obtain the time difference value of each beat point. All differences are sorted in an orderly manner to form beat synchronization accuracy data. This data can intuitively reflect the synchronization accuracy of timing and beat, ensuring that the time difference is controlled within 1 millisecond; The multi-core coprocessor integrates and collects the calculated total training duration, the statistically obtained beat following deviation value, and the extracted and sorted beat synchronization accuracy data, and stores them in non-volatile memory according to a preset data format to form training statistics data.

[0060] Step 603: Match the training statistics data to a preset report template, fill the corresponding fields of the report template with the training statistics data, and generate a training report text containing text descriptions. Specifically, the device's multi-core coprocessor retrieves the report template pre-stored in non-volatile memory. This template is designed according to the actual needs of music teaching and includes fixed fields such as total training time, beat following deviation value, and beat synchronization accuracy data. Each field has a corresponding standardized text description format. The text description format for the total training time field is "The total training time for this teaching session is X hours X minutes X seconds," and the text description format for the beat following deviation value field is "The beat following deviation value for this teaching session is X, and the beat synchronization accuracy data is X." The text description format for the precision data field is that the synchronization time difference of each beat point in this teaching and training is within 1 millisecond, and the overall beat synchronization precision meets the teaching requirements. The multi-core coprocessor accurately matches the various data in the training statistics with the fixed fields in the report template. The total training duration is filled into the corresponding duration field, the beat following deviation value is filled into the corresponding deviation field, and the beat synchronization precision data is filled into the corresponding precision field. According to the exclusive text description format of each field preset in the template, the specific data is substituted into the corresponding description, and a complete text description of each data is provided. In the process of description, the data accuracy and description are ensured to be clear. Finally, a training report text containing a complete text description is generated.

[0061] Step 604: The training report text is input into the speech synthesizer for speech conversion processing to generate a training report speech signal. The training report speech signal is then output through a bone conduction speaker. Specifically, the device's multi-core coprocessor transmits the generated training report text to the device's built-in speech synthesizer. The speech synthesizer converts the text content in the training report text sentence by sentence into analog speech electrical signals according to preset speech synthesis rules. During the conversion process, the clarity and fluency of the speech are ensured, and the output characteristics of the bone conduction speaker are adapted to generate the training report speech signal. The multi-core coprocessor transmits the training report speech signal to the device's bone conduction speaker. This speaker is equipped with a hydrophobic acoustic diaphragm and can achieve underwater frequency response compensation from 20 Hz to 20 kHz. After receiving the speech electrical signal, the bone conduction speaker converts the electrical signal into a mechanical vibration signal and outputs the speech content of the training report through bone conduction. Even in an underwater environment, the recognizability of the speech output is ensured, and the speech broadcast of the training report is completed.

[0062] This embodiment achieves accurate conversion of training report text into speech signals, and completes the speech output of the training report by relying on a bone conduction speaker. The bone conduction output method is compatible with the waterproof characteristics of the device and can ensure the clarity of the speech output, allowing users to quickly obtain the training report content.

[0063] like Figure 2As shown, embodiments of the present invention also provide a multimodal cooperative control system, comprising:

[0064] The acquisition module is used to receive encrypted commands sent by the host and obtain the raw command data;

[0065] The parsing module is used to parse the raw instruction data to obtain the beat rhythm type and timing duration information;

[0066] The processing module is used to generate corresponding beat pulse signals according to the beat rhythm type;

[0067] The synchronization module is used to start the electronic stopwatch based on the timing duration information and use the beat pulse signal at the moment of start as the synchronization reference to align the starting point of the electronic stopwatch with the beat pulse, so as to obtain the synchronized stopwatch timing data.

[0068] The compensation module is used to dynamically route and mix the Bluetooth audio signal and the teaching voice signal based on the beat pulse signal and stopwatch timing data during the electronic stopwatch timing period, generate a mixed audio signal, and synchronously record the time correspondence between the stopwatch timing data and the beat pulse signal to obtain synchronous time axis data.

[0069] The output module is used to generate a training report based on the synchronized timeline data after the electronic stopwatch has finished timing, and output the training report via voice.

[0070] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.

[0071] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0072] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0073] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multimodal cooperative control method, characterized in that, The method includes: Step 1: Receive the encrypted command sent by the host and obtain the original command data; Step 2: Parse the raw instruction data to obtain the beat rhythm type and timing duration information; Step 3: Generate the corresponding beat pulse signal according to the beat rhythm type; Step 4: Based on the timing duration information, start the electronic stopwatch and use the beat pulse signal at the moment of start as the synchronization reference to align the starting point of the electronic stopwatch with the beat pulse, thereby obtaining the synchronized stopwatch timing data. Step 5: During the electronic stopwatch timing, based on the beat pulse signal and stopwatch timing data, the Bluetooth audio signal and the teaching voice signal are dynamically routed and mixed to generate a mixed audio signal, and the time correspondence between the stopwatch timing data and the beat pulse signal is recorded synchronously to obtain synchronous time axis data. Step 6: After the electronic stopwatch finishes timing, generate a training report based on the synchronized timeline data, and output the training report via voice.

2. The multimodal cooperative control method according to claim 1, characterized in that, Receive encrypted instructions sent by the host and obtain the original instruction data, including: The encrypted instruction data stream is obtained by receiving the encrypted instruction sent by the host via wireless communication; the encrypted instruction data stream is then decrypted to obtain the decrypted instruction text. The decrypted instruction text is subjected to integrity verification. Once the verification is successful, it is used as the original instruction data.

3. The multimodal cooperative control method according to claim 2, characterized in that, Parsing the raw instruction data yields the beat rhythm type and timing duration information, including: Identify the teaching instruction identifier and timing instruction identifier in the original instruction data, and segment the original instruction data into teaching instruction part and timing instruction part; The beat speed and beat type parameters are extracted from the teaching instruction section and combined into a beat rhythm type. The timing duration parameter is extracted from the timing instruction section and converted into a preset format timing duration information.

4. The multimodal cooperative control method according to claim 3, characterized in that, Based on the rhythm type, generate corresponding beat pulse signals, including: Based on the beat type parameter, the corresponding rhythm pattern data is read from the pre-stored non-volatile memory to obtain the strong and weak beat distribution pattern in each measure; Based on the beat speed parameters, the time interval of each beat pulse is calculated to obtain the basic pulse period; based on the strong and weak beat distribution pattern and the basic pulse period, a continuous beat pulse signal is generated. The timing of each pulse in the beat pulse signal is determined by the basic pulse period, and the amplitude characteristics of each pulse are determined by the strength attribute of the corresponding position in the strong and weak beat distribution pattern.

5. The multimodal cooperative control method according to claim 4, characterized in that, Based on the timing duration information, the electronic stopwatch is started, and the beat pulse signal at the moment of start is used as the synchronization reference to align the starting point of the electronic stopwatch with the beat pulse, thus obtaining synchronized stopwatch timing data, including: The timing endpoint of the electronic stopwatch is set according to the timing duration information to obtain the timing endpoint parameter; the beat pulse signal is monitored according to the timing endpoint parameter, and when the first pulse edge is detected, the trigger time of the detected pulse edge is taken as the timing start of the electronic stopwatch, and the beat number corresponding to the detected pulse edge is recorded to obtain the synchronization start time and the starting beat mark. The electronic stopwatch is started based on the synchronization start time and the starting beat mark. During the timing process, the current beat pulse number corresponding to each timing moment is determined based on the continuous pulse sequence of the beat pulse signal. The time value of each timing moment is associated with and stored with the corresponding current beat pulse number to obtain the synchronized stopwatch timing data.

6. The multimodal cooperative control method according to claim 5, characterized in that, During the electronic stopwatch timing, based on the beat pulse signal and stopwatch timing data, the Bluetooth audio signal and the teaching voice signal are dynamically routed and mixed to generate a mixed audio signal. Simultaneously, the time correspondence between the stopwatch timing data and the beat pulse signal is recorded to obtain synchronized timeline data, including: Based on the correspondence between the time value and the beat pulse number stored in the synchronized stopwatch timing data, the mapping relationship between the time value and the beat pulse number is extracted to obtain the beat time mapping data; the beat pulse number of the real time is determined based on the beat time mapping data, and the preset audio routing rules are queried based on the beat pulse number of the real time to obtain the audio mixing ratio parameters. The system acquires water depth data detected by a water pressure sensor, performs frequency response compensation processing on the Bluetooth audio signal based on the water depth data to obtain a compensated Bluetooth audio signal, and performs frequency response compensation processing on the teaching voice signal based on the water depth data to obtain a compensated teaching voice signal; and mixes and synthesizes the compensated Bluetooth audio signal and the compensated teaching voice signal according to the audio mixing ratio parameter to generate a mixed audio signal. During the generation of the mixed audio signal, the generation time corresponding to each frame of the mixed audio signal is associated with and stored in the time value and beat pulse sequence number of the same moment in the synchronized stopwatch timing data to obtain the synchronized time axis data.

7. The multimodal cooperative control method according to claim 6, characterized in that, After the electronic stopwatch finishes timing, a training report is generated based on the synchronized timeline data and output via voice. The training report includes: Extract all stored timestamps, beat pulse numbers, and timing values ​​from the synchronous timeline data to obtain the training process time series data; The total training time is calculated based on the timing value, the beat following deviation value is statistically analyzed based on the change of the beat pulse sequence number, and the difference between the actual trigger time and the theoretical trigger time of each beat point is extracted based on the correspondence between the timestamp and the timing value to obtain the beat synchronization accuracy data. The total training time, beat following deviation value and beat synchronization accuracy data are summarized into training statistics data. Match the training statistics data with the preset report template, fill the corresponding fields of the report template with the training statistics data, and generate a training report text containing text descriptions; The training report text is input into a speech synthesizer for speech conversion processing to generate the training report speech signal, which is then output through a bone conduction speaker.

8. A multimodal cooperative control system, wherein the system implements the method as described in any one of claims 1 to 7, characterized in that, include: The acquisition module is used to receive encrypted commands sent by the host and obtain the raw command data; The parsing module is used to parse the raw instruction data to obtain the beat rhythm type and timing duration information; The processing module is used to generate corresponding beat pulse signals according to the beat rhythm type; The synchronization module is used to start the electronic stopwatch based on the timing duration information and use the beat pulse signal at the moment of start as the synchronization reference to align the starting point of the electronic stopwatch with the beat pulse, so as to obtain the synchronized stopwatch timing data. The compensation module is used to dynamically route and mix the Bluetooth audio signal and the teaching voice signal based on the beat pulse signal and stopwatch timing data during the electronic stopwatch timing period, generate a mixed audio signal, and synchronously record the time correspondence between the stopwatch timing data and the beat pulse signal to obtain synchronous time axis data. The output module is used to generate a training report based on the synchronized timeline data after the electronic stopwatch has finished timing, and output the training report via voice.

9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.