Modularized music creation and playing device, control method and device thereof and medium
The modular music creation and playback device achieves harmonious sound production from multiple units and high degree of creative freedom through the detachable connection of the main control module and functional sub-modules. It solves the problems of poor sound coordination and insufficient connection stability of existing musical toys, and enhances the user's creative experience and the device's expandability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANSONG NANJING TECH LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing musical toys suffer from poor sound coordination, significant noise interference, limited creative modes, insufficient connection stability, and poor expandability, making it difficult to meet users' needs for developing musical awareness and creating music independently.
The design incorporates a modular music creation and playback device, featuring detachable connections between the main control module and functional sub-modules. The central processing unit identifies the modules' identities to enable track synthesis and playback. It supports harmonious sound generation from multiple units, offering a high degree of creative freedom. Furthermore, the magnetic and Pogo Pin connection method ensures circuit stability and convenience.
It achieves harmonious sound generation between multiple modules, enhances the flexibility and fun of creation, strengthens users' creative confidence and aesthetic experience, supports multi-level cascading expansion and physical creation, and solves the problems of noise interference and connection stability.
Smart Images

Figure CN121963674A_ABST
Abstract
Description
Modular music creation and playback devices, control methods, devices and media Technical Field
[0001] This specification relates to the field of playback device technology, and in particular to a modular music creation and playback device and its control method, apparatus and medium. Background Technology
[0002] Currently, most musical toys on the market suffer from poor sound coordination. Many are designed with independent sound units, and using multiple such toys simultaneously can easily generate noise, hindering the development of a good musical aesthetic in users and potentially discouraging their creative enthusiasm. Furthermore, existing music education toys often focus on unidirectional music playback, lacking the ability to translate abstract musical concepts into tangible, interactive forms that users can perceive. Users can only perform simple playback operations, making it difficult to create music through intuitive physical manipulation.
[0003] In addition, the creation mode of traditional user music toys is limited, and they cannot collect and integrate the user's own voice, which restricts the creative autonomy and personalized expression. Their circuit connection structure also has problems such as insufficient connection stability or inconvenient plug-and-play operation, and most of them do not have the ability to expand through multi-level cascading, resulting in poor flexibility of use and making it difficult to meet the actual needs of users' music cognition cultivation and independent creation.
[0004] Therefore, there is an urgent need for a modular music creation and playback device and its control method that can achieve harmonious sound production from multiple units and also has a high degree of creative freedom. Summary of the Invention
[0005] This specification provides one or more embodiments of a modular music creation and playback device, the device including a main control module and functional sub-modules; the functional sub-modules are detachably connected to the main control module, the functional sub-modules include a first type interface and a second type interface, and the functional sub-modules store identity information corresponding to preset audio tracks; the main control module includes a central processing unit, an audio playback unit, and the first type interface; the central processing unit is configured to: in response to the occurrence of a target event, read the identity information of the target module; the target event includes the second type interface of the target module being connected to the first type interface of the main control module; or the second type interface of the target module being connected to the first type interface of an already connected sub-module; determine the audio track to be processed based on the identity information of the target module; synthesize the audio track to be processed with the current audio stream to obtain a target audio stream; and control the audio playback unit to play the target audio stream.
[0006] This specification provides one or more embodiments of a control method for a modular music creation and playback device. The method includes: in response to the occurrence of a target event, reading the identity information of a target module; the target event includes the target module's second type interface being connected to the main control module's first type interface; or the target module's second type interface being connected to the first type interface of an already connected sub-module; determining the audio track to be processed based on the target module's identity information; synthesizing the audio track to be processed with the current audio stream to obtain a target audio stream; and controlling an audio playback unit to play the target audio stream.
[0007] This specification provides one or more embodiments of a control device for a modular music creation and playback apparatus, the apparatus including a processing device for executing a control method for the modular music creation and playback apparatus.
[0008] This specification provides one or more embodiments of a computer-readable storage medium that stores computer instructions that, when executed by a processor, implement a control method for the modular music creation and playback device. Attached Figure Description
[0009] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same reference numerals denote the same structures, wherein: FIG1 is a schematic diagram of an application scenario of a modular music creation and playback device according to some embodiments of this specification; FIG2 is a schematic diagram of the structure of a modular music creation and playback device according to some embodiments of this specification; FIG3 is an exemplary flowchart of a control method for a modular music creation and playback device according to some embodiments of this specification.
[0010] Figure 4 is a schematic diagram of the control flow of different functional sub-modules according to some embodiments of this specification; Figure 5 is an exemplary schematic diagram of determining the target audio stream according to some embodiments of this specification. Detailed Implementation
[0011] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.
[0012] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0013] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.
[0014] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0015] Figure 1 is a schematic diagram of an application scenario of a modular music creation and playback device according to some embodiments of this specification.
[0016] In some embodiments, the modular music creation and playback device can realize the free splicing of multifunctional sub-modules and automatic synthesis of audio tracks, ensuring the harmony of the melody played after splicing. At the same time, the device is applicable to a wide range of scenarios, such as family games, music enlightenment teaching, and parent-child interactive entertainment.
[0017] In some embodiments, as shown in FIG1, the application scenario of the modular music creation and playback device (hereinafter referred to as application scenario 100) may include a modular music creation and playback device 110, a user terminal 120, a network 130, a processor 140, and a database 150.
[0018] The modular music creation and playback device 110 can realize the functions of access recognition of functional sub-modules, track synthesis, and audio playback. In some embodiments, the modular music creation and playback device 110 may include a main control module and functional sub-modules. For more information about the modular music creation and playback device 110 and its main control module and functional sub-modules, please refer to Figure 2 and its related description.
[0019] User terminal 120 is a terminal device that interacts with a user. For example, a user terminal may include a mobile phone 120-1, a tablet 120-2, a computer 120-3, etc. A user refers to one or more users who use the modular music creation and playback device 110, such as a child using the modular music creation and playback device 110 or a child's parent, teacher, etc.
[0020] In some embodiments, the main control module is connected to the user terminal 120 via a built-in main Bluetooth or WiFi communication module. The user can update the built-in sound library of the main control module using an application (App) on the user terminal 120. More information about the main control module can be found in Figure 2 and its related description.
[0021] The built-in tone library refers to a standardized tone data set integrated within the main control module. For example, the built-in tone library may include standardized tone data corresponding to various preset tracks. See Figures 2 and 3 for a description of preset tracks.
[0022] Network 130 may include any suitable network capable of facilitating the exchange of information and / or data. In some embodiments, at least one component of application scenario 100 (e.g., modular music creation and playback device 110, user terminal 120, processor 140, database 150, etc.) may exchange information and / or data with at least one other component of application scenario 100 via network 130.
[0023] For example, processor 140 can obtain data related to the operating status of modular music creation and playback device 110 via network 130. As another example, user terminal 120 can update or set preset tracks for each functional sub-module via network 130. Furthermore, database 150 can receive and store identity information and matched preset tracks uploaded by modular music creation and playback device 110 via network 130. More information about functional sub-modules, preset tracks, and identity information can be found in Figure 2 and its related description.
[0024] In some embodiments, network 130 can be any one or more of wired or wireless networks. For example, network 130 may include cable networks, fiber optic networks, telecommunications networks, cable connections, etc. The network connections between the various parts can be achieved using one or more of the above methods. In some embodiments, the network can be a point-to-point, shared, centralized, or other topologies, or a combination of multiple topologies. In some embodiments, network 130 may include one or more network access points.
[0025] The processor 140 is used to process data and / or information related to the application scenario 100 of the modular music creation and playback device. In some embodiments, the processor 140 may process data, information and / or processing results obtained from other devices or system components, and execute program instructions based on such data, information and / or processing results to perform one or more functions described in this specification.
[0026] For example, the processor 140 can obtain relevant user information from the user terminal 120 via the network 130, obtain the identity information of functional sub-modules from the modular music creation and playback device 110, and access the correspondence data between functional sub-modules and preset audio tracks stored in the database 150. Furthermore, the processor can execute related instructions based on this data, such as determining the audio track to be processed, synthesizing the target audio stream, and controlling the audio playback unit to play the target audio stream.
[0027] For more information on identity information, preset audio tracks, audio tracks to be processed, and target audio streams, please refer to Figures 2 and 3 and their related descriptions. For more information on the audio playback unit, please refer to Figure 2 and its related descriptions.
[0028] In some embodiments, processor 140 may include one or more sub-processing devices (e.g., a single-core processing device or a multi-core multi-chip processing device). By way of example only, processor 140 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), a microprocessor, or any combination thereof. In some embodiments, processor 140 may be part of a modular music creation and playback device 110.
[0029] Database 150 may store data, instructions, and / or any other information related to the modular music creation and playback device 110. For example, database 150 may store data and / or information acquired by the modular music creation and playback device 110, user terminal 120, and processor 140 (e.g., the identity information of functional sub-modules of the modular music creation and playback device 110, the correspondence between functional sub-modules and preset tracks, etc.). As another example, database 150 may also store the target audio stream obtained by the modular music creation and playback device 110, relevant information of user terminal 120, etc.
[0030] In some embodiments, database 150 may store data and / or instructions used by processor 140 to execute or use in order to perform the exemplary methods described herein. For example, database 150 may store instruction data generated by processor 140, such as pre-release of an audio track to be processed.
[0031] In some embodiments, database 150 may include one or more storage units, each of which may be a separate device or part of another device. In some embodiments, database 150 may be implemented on a cloud platform. In some embodiments, database 150 may be part of user terminal 120, processor 140, and / or modular music creation and playback device 110. Processor 140 may also be part of modular music creation and playback device 110.
[0032] Figure 2 is a structural schematic diagram of a modular music creation and playback device according to some embodiments of this specification.
[0033] In some embodiments, as shown in FIG2, the modular music creation and playback device includes a main control module 1 and a functional sub-module 2; the functional sub-module 2 is detachably connected to the main control module 1, and the functional sub-module 2 includes a first type interface 104 (not shown in the figure) and a second type interface 202, and the functional sub-module 2 stores identity information corresponding to a preset audio track; the main control module includes a central processing unit 102, an audio playback unit (not shown in the figure) and the first type interface 104.
[0034] The central processing unit 102 is configured to: in response to a target event, read the identity information of the target module; the target event includes the target module's second type interface 202 being connected to the main control module 1's first type interface 104; or the target module's second type interface 202 being connected to the first type interface 104 of an already connected sub-module; determine the audio track to be processed based on the target module's identity information; synthesize the audio track to be processed with the current audio stream to obtain the target audio stream; and control the audio playback unit to play the target audio stream. For further explanation of the above, please refer to the corresponding content in Figures 3 to 5.
[0035] The main control module is the core module of the modular music creation and playback device, used to manage functional sub-modules, process audio tracks, and control audio playback. The management of functional sub-modules includes, but is not limited to: the main control module enabling detachable connections with functional sub-modules through a first-type interface, identifying the cascaded connection status between functional sub-modules, and reading the identity information of each functional sub-module.
[0036] In some embodiments, the main control module may be designed in the shape of a cube, or a base-like structure. More information about the first type of interface and functional submodules can be found in the following descriptions.
[0037] The central processing unit (CPU) is the core computing component of the main control module, configured to execute logical instructions such as identity information reading, determining the audio track to be processed, audio stream synthesis, and playback control. For example, the CPU can be a microcontroller unit (MCU), a digital signal processor (DSP), or the like. In some embodiments, the CPU can be implemented based on processor 140.
[0038] For further details on what the central processing unit is configured to execute, see Figure 3 and its related description.
[0039] The audio playback unit is used to convert the digital audio signal output by the central processing unit into an audible audio signal and output it. In some embodiments, the audio playback unit may include a digital-to-analog converter circuit (not shown), a power amplifier (not shown), a speaker 101, etc., connected in sequence.
[0040] A digital-to-analog converter circuit is used to decode digital audio signals and convert them into analog audio signals.
[0041] Power amplifiers are used to amplify both the current and voltage of analog audio signals.
[0042] A loudspeaker is used to convert a power-amplified analog audio signal into an audible audio signal for playback. In some embodiments, the loudspeaker is electrically connected to the output of the power amplifier.
[0043] The first type of interface refers to the adapter connection interface set on the main control module and functional sub-modules, which is used for the connection between the main control module and the functional sub-modules or the connection between multiple functional sub-modules.
[0044] In some embodiments, the main control module further includes a battery 103 for providing power.
[0045] In some embodiments, the main control module may also include communication components such as Bluetooth and WIFI communication modules. The communication components can be used to realize the communication connection between the main control module and the app of the user terminal 120, so that the user can control the main control module through the user terminal 120, for example, to update the preset audio tracks of the function sub-module.
[0046] A functional submodule refers to a functional unit that can be flexibly combined with the main control module. In some embodiments, there can be one or more functional submodules. Furthermore, different functional submodules may correspond to the same or different functions.
[0047] In some embodiments, the appearance of a functional submodule can be designed in any form, such as a cube, a sphere, etc.
[0048] In some embodiments, the functional submodule includes a first type of interface and a second type of interface, and the functional submodule stores identity information corresponding to a preset audio track.
[0049] In some embodiments, the first type interface 104 and the second type interface 202 are adapter interfaces; wherein, one of the first type interface 104 and the second type interface 202 is provided with a connector, and the other is provided with metal contacts; both the first type interface 104 and the second type interface 202 are provided with a ring magnet, a power pin, and a data communication pin. In some embodiments, the connector can be a Pogo Pin spring pin connector, and the metal contacts can be metal contacts that are compatible with the connector.
[0050] In some embodiments, the power supply pin can be the voltage common collector (VCC), ground (GND), etc. The data communication pin can be the data transmit pin (TX), the data receive pin (RX), the inter-integrated circuit (I2C), etc.
[0051] For more information on the first type of interface, please refer to the preceding text and its related descriptions.
[0052] The second type of interface is an adapter connection interface set on a functional submodule. In some embodiments, the second type of interface is detachably connected to the first type of interface.
[0053] In some embodiments, based on the first type of interface and the second type of interface, multiple functional sub-modules and the main control module can achieve quick connection and quick data transfer in various ways. For example, based on the positional design of the first type of interface on the main control module and the positional design of the first type of interface and the second type of interface on the functional sub-modules, multiple functional sub-modules and the main control module can be connected end-to-end like train carriages, or stacked like a pyramid, etc.
[0054] In some embodiments, when a user attaches a functional submodule to the main control module or another functional submodule that is already connected to the main control module, the circuits between the modules are connected through connectors, and the main control module can detect the newly connected functional submodule by polling via I2C or UART bus based on the central processing unit.
[0055] In some embodiments, the Pogo Pin spring pin can work together with the ring magnet to achieve module adsorption, thus completing the cascading expansion of functional sub-modules.
[0056] Some embodiments in this specification employ an adapter interface design, enabling quick module docking via a ring magnet. The connector and metal contacts work together to ensure stable circuit connection. Independent power supply pins and data communication pins enable reliable transmission of power and signals, balancing ease of connection with transmission stability.
[0057] In some embodiments, in addition to the power supply pins (VCC, GND) and data communication pins (TX, RX), the first and second type interfaces may also have a dedicated detection pin (DET). The detection pin is connected to a preset port of the central processing unit (such as an interrupt input port) via a pull-up resistor. When a functional submodule is connected, the ring magnet of the interface ensures stable physical connection. After the connector contacts the metal contacts, the level of the detection pin changes from high to low, triggering a hardware interrupt. Upon responding to the interrupt, the central processing unit immediately initiates an enumeration query through the data communication pin to read the target module's identity information, achieving millisecond-level plug-and-play identification. Simultaneously, the first and second type interfaces may also have built-in anti-bounce circuitry to filter out false interrupt signals caused by contact jitter.
[0058] In some embodiments, the main control module can also integrate a power management chip (PMIC) to support multiple outputs and dynamic current distribution. The power supply pin (VCC) of the first type interface is independently controlled by the PMIC and has overcurrent and short-circuit protection functions. The functional submodules have internal low-dropout regulators (LDOs) or DC-DC converters to convert the input voltage to the module operating voltage (e.g., 3.3V). When multiple functional submodules are cascaded, the PMIC dynamically adjusts the output current according to the power consumption requirements reported by each module, and the cascade current limiting threshold (e.g., total current ≤ 2A) can be set through the user terminal to prevent system overload.
[0059] A preset track is a pre-recorded audio data file of a fixed length. Different preset tracks can correspond to different sound effects. For example, a preset track may include recordings of piano sounds, guitar sounds, embellishments, rhythm changes, etc. Embellishments refer to short, decorative sound effect clips.
[0060] In some embodiments, different functional sub-modules correspond to the same or different preset audio tracks. Users can update or set the preset audio tracks of each functional sub-module based on their user terminal.
[0061] In some embodiments, all preset audio tracks are stored in the functional submodules in a standard format (such as 16-bit PCM, 44.1kHz sampling rate, mono). The central processing unit can upsample the input audio to 48kHz and convert it to a 32-bit floating-point internal representation via a built-in sample rate converter (SRC) for easy subsequent digital processing.
[0062] Identity information is a unique identifier stored within a functional submodule, used by the central processing unit to identify the type of the functional submodule and its corresponding preset track. For example, identity information can be an ID number plus a track number. Different track numbers correspond to different preset tracks, such as 01 representing piano sound effects, 02 representing guitar sound effects, 03 representing drum sound effects, 04 representing recording, and 05 representing preset sound effects, etc.
[0063] In some embodiments, a functional submodule corresponds to a unique identity information, and the identity information of different functional submodules corresponds to different ID numbers.
[0064] In some embodiments, the central processing unit can read the identity information of the functional sub-modules through the connection between the first type interface of the main control module and the second type interface of the functional sub-modules. The central processing unit can have a built-in intelligent music arrangement algorithm to dynamically adjust the beats per minute (BPM) and key on the bus according to the number and type of the connected functional sub-modules. It can also call a preset audio algorithm to synchronize the beat and match the harmony of the corresponding instrument sound effects track of the connected functional sub-modules with the current audio stream, ensuring that the sounds emitted by all functional sub-modules are in the same key and rhythm, and then synthesize and play them through the speakers.
[0065] Intelligent music arrangement algorithms refer to algorithms that can automatically identify the type and number of instruments in the connected functional sub-modules, and then dynamically adjust the overall beat (BPM) and key. For example, when connecting two instrument modules corresponding to guitar and drums, the intelligent music arrangement algorithm sets the BPM to 120 and the key to C major; after adding an instrument module corresponding to bass, the intelligent music arrangement algorithm automatically lowers the BPM to 90 to match the rhythm and style of bass.
[0066] The preset audio algorithm refers to an algorithm that can perform processing logic to calibrate the beat and unify the tonality of multiple instrument tracks. For example, if the guitar track has a beat deviation or the bass track's tonic is inconsistent, the preset audio algorithm can calibrate the guitar beat to the reference BPM set by the intelligent arrangement algorithm and convert the bass tonality to a unified tonic. Tonality refers to the core pitch reference and modal attribute of an audio track.
[0067] The intelligent music arrangement algorithm and audio algorithm can be integrated into the central processing unit of the main control module as instruction programs, so that the central processing unit can directly read and execute them. For further explanation of the implementation methods of the intelligent music arrangement algorithm and audio algorithm, please refer to the corresponding content in Figures 3 to 5.
[0068] In some embodiments, the main control module may also incorporate a high-precision crystal oscillator (e.g., ±10ppm) as a global clock source. The central processing unit periodically sends synchronization packets to all connected sub-modules via a data communication pin of a first-type interface (e.g., I2C). Each packet contains the current beat count (BPM), measure count, and beat position. Upon receiving the synchronization packet, each functional sub-module adjusts its internal audio playback timing accordingly, achieving beat alignment between multiple modules. To address transmission delays, the central processing unit employs a timestamp compensation mechanism, dynamically adjusting the synchronization offset based on the number of cascaded module jumps to ensure that all modules output corresponding audio data at the same musical moment.
[0069] In some embodiments, the functional submodule further includes an indicator device, which can be used to indicate the current operating state of the functional submodule or to issue corresponding instructions based on received control commands. For example, the indicator device can be a light-emitting device such as RGB LED beads. RGB LED beads can serve as an audio-visual synchronization feedback component integrated within each functional submodule, used to achieve rhythmic flashing in sync with the audio rhythm.
[0070] In some embodiments, the functional submodule 2 includes at least one of the following: instrument module 20, recording module 30, and special effects module (not shown in the figure). For example, the preset track corresponding to the instrument module can be a pre-recorded segment such as piano sound effect, guitar sound effect, and drum kit sound effect; the preset track corresponding to the recording module is blank, and when the recording module is activated or recognized by the central processing unit, the recording module can generate a dynamic recording track in real time through its built-in acquisition component; the preset track corresponding to the special effects module can be a specific sound effect track (such as embellished sound effects, rhythm change segments, etc.).
[0071] The instrument module refers to a functional submodule that stores preset tracks corresponding to specific instrument sound effects and specific instrument IDs (i.e., identity information).
[0072] In some embodiments, each functional submodule may include an identity recognition circuit 201, etc.
[0073] The identity recognition circuit 201 is integrated within each functional submodule and is used to store the identity information of each functional submodule. In some embodiments, the identity recognition circuit can establish communication with the first type interface of the main control module through the second type interface of the functional submodule, or establish communication with the first type interface of the connected submodule through the second type interface of the functional submodule, thereby transmitting the identity information to the central processing unit, so that the central processing unit can identify the type of the connected functional submodule and the corresponding preset audio track.
[0074] An already connected submodule refers to a functional submodule that is directly or indirectly connected to the main control module. The second type interface of a functional submodule that is directly connected to the main control module is connected to the first type interface of the main control module; the second type interface of a functional submodule that is indirectly connected to the main control module is cascaded with the first type interface of other already connected submodules.
[0075] A recording module refers to a functional submodule that has audio acquisition and storage capabilities. In some embodiments, a recording module may include a recording button 301, a microphone 302, an analog-to-digital converter circuit 303, etc.
[0076] The recording button is located on the surface of the recording module. In some embodiments, the user can press the recording button to trigger the audio acquisition function of the recording module, record short audio content such as "Mom, I love you", and press it again to stop the acquisition. After the acquisition is completed, the central processing unit of the main control module will automatically quantize the short audio segment, decode the digital audio data obtained after quantization through an audio decoding algorithm to obtain an audio stream, and use the audio stream as a dynamic recording track.
[0077] Quantization refers to the process of converting the acquired analog audio signal into digital audio data that can be recognized and processed by the central processing unit. The processed digital audio data can accurately match the tempo and key of the current audio track.
[0078] An audio decoding algorithm is a signal processing algorithm that converts encoded digital audio data back into an audio stream that can be directly edited, synthesized, or played. For example, audio decoding algorithms may include PCM decoding algorithms, MP3 decoding algorithms, etc. In some embodiments, the audio decoding algorithm is built into the central processing unit.
[0079] In some embodiments, to achieve real-time pitch matching of audio captured by the recording module, the central processing unit (CPU) incorporates lightweight digital signal processing (DSP) algorithms, including resampling-based pitch shifting algorithms and waveform similarity-based time stretching algorithms. To balance real-time performance and sound quality, the CPU can employ a segmented processing strategy, such as dividing the recorded audio stream into several frames (e.g., 512 samples / frame), performing pitch shifting and alignment at the frame level, and using overlap-add technology to ensure smooth transitions between frames. Simultaneously, the CPU can pre-store fundamental frequency mapping tables for commonly used keys (e.g., C, G, F major), and combine this with Fast Fourier Transform (FFT) to achieve low-latency pitch detection and adjustment. The overall processing latency is controlled within a preset threshold (e.g., 50 milliseconds) to meet real-time interactive requirements.
[0080] The microphone is integrated inside the recording module and serves as an audio acquisition component to receive sound signals (such as human voices) emitted by the user.
[0081] The analog-to-digital converter circuit is integrated inside the recording module and is used to convert the analog audio signals captured by the microphone into digital audio signals.
[0082] The special effects module refers to a functional sub-module that stores preset audio tracks and specific sound effect IDs (i.e., identity information) for specific sound effects (such as decorative sound effects, rhythm change segments, etc.).
[0083] Some embodiments in this specification enrich the types of materials available for music creation and enhance the diversity and flexibility of creation through the design of multiple types of functional sub-modules.
[0084] The main control module and functional sub-modules are detachably connected through the adaptation and cooperation of the first type of interface and the second type of interface.
[0085] In some embodiments, the central processing unit (CPU) may also incorporate a multi-channel digital mixer, supporting real-time mixing of up to 16 audio streams. The mixer applies dynamic gain control (DGC) to each input audio stream, automatically adjusting volume weights based on track priority and user settings to prevent signal overload after synthesis. Furthermore, the mixer integrates a soft limiter to smoothly limit the synthesized audio, avoiding clipping distortion caused by instantaneous peaks. Users can select preset mixing modes (such as "Equalized Mix" or "Emphasis on Melody") via a user terminal application, and the CPU adjusts the equalizer (EQ) parameters and panning distribution of each track accordingly.
[0086] In some embodiments, when the second type interface of a functional submodule is connected to the first type interface of the main control module, or when the second type interface of a functional submodule is connected to the first type interface of an already connected submodule, the Pogo Pin conduction circuit at the interface is activated. The central processing unit polls and detects via I2C or UART bus to obtain the ID of the functional submodule and complete the connection identification. After the connection identification is completed, the central processing unit in the main control module waits for the start point of the next measure or beat based on its own maintained global clock, and then controls the audio playback unit to synchronously play the preset audio track corresponding to the connected functional submodule to ensure that all audio tracks have the same rhythm.
[0087] During playback, the central processing unit (CPU) intelligently mixes audio based on the type of the connected functional sub-modules. For example, if only the instrument module corresponding to the piano sound effect is connected, a piano solo is played. If an instrument module corresponding to the drum sound effect is subsequently added, drum beats are automatically overlaid and the piano volume is fine-tuned. If a recording module is then connected, the digital signal processor (DSP) modulates the digital audio signal output from the analog-to-digital converter in the recording module to match the pitch of the current background music. Simultaneously, the CPU outputs a synchronization signal, driving the RGB LEDs within each functional sub-module to flash rhythmically with the music, achieving an integrated audiovisual interactive experience.
[0088] For example, after the user activates the main control module, the central processing unit controls the audio playback unit to play the basic background noise rhythm, laying the foundation for music creation. The user selects the instrument module corresponding to the bass sound effect, attaches its second type interface to the first type interface of the main control module, and after the central processing unit recognizes the module ID, it waits for the start of the next measure or beat to play the bass melody synchronously, realizing melody looping. Then, the user continues to attach the recording module, presses the record button on the module to record a "hey!", and the recording module performs pitch shifting processing on the recording through the digital signal processor (DSP) to make its pitch match the current bass background music key, and automatically plays a rhythmic "hey!" sound at the end of every two beats. Finally, the user attaches the instrument module corresponding to the violin sound effect, and after the central processing unit recognizes it, it automatically cuts the melodious chorus melody into the current audio stream, completing the intelligent mixing of multiple tracks. The user can easily complete a complex electronic music composition with just simple module attachment and recording operations.
[0089] In some embodiments described in this manual, the modular music creation and playback device solves the problem of noise caused by multiple devices independently producing sound in existing musical toys through unified clock synchronization and tonality alignment of the main control module. This ensures that regardless of how the modules are assembled, the generated melodies are musically harmonious and pleasant to listen to, effectively enhancing the creative confidence and aesthetic experience of children and other users. By transforming the abstract concept of sound tracks into modules with a realistic block-like effect, and using magnetic and Pogo Pin connections, the device not only adapts to the cognitive development patterns of children and other users, realizing the materialization and visualization of music creation, but also ensures stable circuit connections and is easy for children and other users to plug and unplug, supporting multi-level cascading expansion. The addition of a recording module allows users to incorporate personalized sounds such as human voices and clapping sounds into the music. Combined with the free combination characteristics of the modules, it gives users a highly free creative interactive experience, further enhancing the fun and playability of the device.
[0090] Figure 3 is an exemplary flowchart of a control method for a modular music creation and playback device according to some embodiments of this specification. The process 300 includes steps 310-340, and the process 300 can be executed by a central processing unit.
[0091] Step 310: In response to the occurrence of a target event, read the identity information of the target module.
[0092] A target event refers to a specific connection event that triggers the central processing unit to perform an operation to read identity information.
[0093] In some embodiments, the target event includes the second type of interface of the target module being connected to the first type of interface of the main control module; or the second type of interface of the target module being connected to the first type of interface of the connected sub-module.
[0094] The target module refers to a module that is newly connected to the main control module or an already connected sub-module during the operation of the modular music creation and playback device (hereinafter referred to as the device). An already connected sub-module refers to a functional sub-module that has been identified by the main control module and participated in the synthesis of the current audio stream before the target event occurs.
[0095] For example, if a user connects a recording module to a connected instrument module when the main control module is already connected to it, then the instrument module is considered an accessed sub-module, and the recording module is the target module. For more information on the main control module and functional sub-modules, please refer to Figure 2 and its related content.
[0096] A connection (hereinafter also referred to as a link) can be a physical connection or an electrical connection. For example, a physical connection can be a physical attraction. An electrical connection can be achieved by connecting power and signal pins through a connector (such as a first interface or a second interface). For more information about the first interface, the second interface, and the connector, please refer to Figure 2 and related content.
[0097] In some embodiments, the target module can be identified and obtained by the central processing unit. For example, when a user connects a functional submodule to the main control module or an already connected submodule, the functional submodule returns identity information to the main control module based on the interface. The central processing unit automatically identifies the newly connected functional submodule as the target module by reading the identity information.
[0098] For more information on identity information, please refer to Figure 2 and its related description.
[0099] In some embodiments, the central processing unit may also send an identity query instruction to the target module through the main control module connected to the target module or the first type of interface of the sub-module that has been connected, thereby reading the identity information of the target module.
[0100] Step 320: Determine the audio track to be processed based on the identity information of the target module.
[0101] The audio track to be processed refers to the audio data corresponding to the identity information of the target module, which is to be synthesized with the current audio stream. The audio data can be preset or generated in real time. For example, the audio data can be preset drum beats or user-input recording data.
[0102] In some embodiments, the central processing unit determines whether the target module is a recording module based on the target module's identity information. When the target module is not a recording module, the central processing unit can determine the preset audio track corresponding to the identity information through a first preset table based on the target module's identity information, and determine the preset audio track corresponding to the identity information as the audio track to be processed. When the target module is a recording module, it activates an audio input device such as a microphone according to user instructions, and acquires and obtains the audio track to be processed through the audio input device such as a microphone.
[0103] Among them, preset audio tracks refer to pre-set audio segments, such as preset drum sound effect segments, piano sound effect segments, etc. For more explanation about preset audio tracks, please refer to the corresponding content in Figure 2.
[0104] The first preset table includes the identity information of each functional submodule and its corresponding preset audio tracks. The first preset table can be preset according to user needs or updated based on user commands obtained from the user terminal. User commands refer to instructions given by the user to input or update audio tracks. The central processing unit can obtain user commands through methods such as capturing user voice or obtaining user input from the user terminal.
[0105] Step 330: Combine the audio track to be processed with the current audio stream to obtain the target audio stream.
[0106] The current audio stream refers to the audio data stream that the central processing unit is playing before the target module is connected.
[0107] For example, before the target module is connected, when the main control module is not connected to other functional sub-modules, the current audio stream is the reference audio stream of the main control module. The reference audio stream refers to a pre-set audio track played on the main control module. For example, a track with a common eighth-note rhythm.
[0108] For example, before the target module is connected, the main control module is connected to a musical instrument module. The current audio stream is the audio stream synthesized from the preset track corresponding to the identity information of the musical instrument module and the reference audio stream of the main control module.
[0109] Synthesis refers to the process of processing a preset audio track and audio stream data to generate a segment of audio data. In some embodiments, the central processing unit can first decode and then preprocess the preset audio track to form streaming audio data. The data value of the streaming audio data corresponding to each time point (such as multiple timestamps) is then superimposed with the data value of the audio stream data corresponding to each time point (e.g., the audio data corresponding to the currently playing audio stream) to obtain the synthesized audio data for that time point, thus achieving synthesis. Superposition can be a process of weighted summation of the data values of the streaming audio track data and the audio stream data at the same time point. Preprocessing may include format resampling, format normalization, and beat alignment.
[0110] In some embodiments, the central processing unit may obtain the currently playing audio information from the audio playback unit and use it as the current audio stream.
[0111] The target audio stream refers to the audio data stream that the central processing unit needs to play after the target module is connected.
[0112] In some embodiments, the central processing unit can combine the audio track to be processed and the current audio stream to generate the target audio stream. That is, the central processing unit can first decode the audio track to be processed and then preprocess it to form a streaming audio track to be processed. The streaming audio track to be processed is then superimposed on the current audio stream to generate the target audio stream.
[0113] For example, suppose the current audio stream is audio data with a sampling rate of 48kHz, stereo, and internal operating format of Float32, and the audio track to be processed is audio data with a sampling rate of 44.1kHz, mono, and storage format of 16-bit PCM (Pulse Code Modulation). The central processing system first resamples, aligns the format, and aligns the beat of the audio track to be processed, thereby increasing the sampling rate of the audio track to be processed from 44.1kHz to 48kHz, converting the storage format from 16-bit to Float32, expanding the mono to stereo, and aligning the beat of the audio track to be processed with that of the current audio stream.
[0114] As an example only, the preprocessed audio track to be processed has audio data values of [data value 11, data value 12, data value 13, data value 14] at four timestamps, and the current audio stream has audio data values of [data value 21, data value 22, data value 23, data value 24] at four timestamps. The weight of the audio track to be processed is w1, and the weight of the current audio stream is w2. At this time, the corresponding audio data values of the synthesized target audio stream at four timestamps are [w1... Data value 11+ w2 Data value 21, w1 Data value 12+w2 Data value 22, w1 Data value 13+w2 Data value 23, w1 Data value 14+w2 Data value 24].
[0115] In some embodiments, the central processing unit is further configured to detect a connectivity graph, determine audio rendering parameters based on the connectivity graph, determine an updated audio track to be processed based on the audio rendering parameters and the audio track to be processed, and synthesize the updated audio track to be processed with the current audio stream to obtain a target audio stream.
[0116] A connection graph is data that describes the mapping relationship of the physical topology between connected functional sub-modules and the main control module.
[0117] In some embodiments, the connection diagram includes the connection relationships between multiple connected functional sub-modules and the main control module. The connection relationships include positional orientation and relative cascading position. Positional orientation refers to the relative position of the functional sub-module with respect to the main control module. Relative cascading position refers to the number of module levels separating the functional sub-module and the main control module.
[0118] For example, if the main control module and the connected functional submodule A are directly connected, and the connected functional submodule A is connected to the left interface of the main control module, and the connected functional submodule B is connected to the connected functional submodule A and then indirectly connected to the main control module, then the positional orientation between the main control module and the connected functional submodule A is: functional submodule A is located to the left of the main control module, with a relative cascading position of 1. The positional orientation between the main control module and the connected functional submodule B is: functional submodule B is located to the left of the main control module, with a relative cascading position of 2.
[0119] In some embodiments, when a new functional submodule is connected, the central processing unit can determine the location direction of the newly connected functional submodule by detecting the signal source direction of the newly connected functional submodule. The main control module can periodically (e.g., based on a preset period of manual experience) or when a new functional submodule is connected, initiate a topology query command. After receiving the topology query command, each functional submodule determines its relative cascade position according to the hop count parameter received by the functional submodule, and sends a return packet containing identity information and relative cascade position back to the main control module. The main control module determines the relative cascade position between the main control module and each functional submodule based on the return packet.
[0120] A topology query command is used to trigger and collect topology information of physical connection links. The topology query command includes a hop count parameter. The hop count parameter is a numerical value used to record the number of functional submodules the topology query command traverses after it is issued from the main control module. For example, when the main control module issues a topology query command, the initial value of this hop count parameter is set to 0. The first functional submodule A directly connected to the main control module receives the command, increments the hop count parameter by 1 to 1, records it as the relative cascade position of functional submodule A, updates the hop count parameter in the command to 1, and then forwards it. When functional submodule B, cascaded with functional submodule A, receives the hop count parameter, reads that it is 1, increments it by 1 to 2, records it as the relative cascade position of functional submodule B, and then continues forwarding until each functional submodule has updated its hop count parameter. The hop count parameter increments once for each functional submodule it traverses, allowing each functional submodule on the link to determine its relative cascade position with the main control module by reading the hop count parameter.
[0121] Audio rendering parameters are control parameters used to adjust the audio track being processed during the audio synthesis stage.
[0122] In some embodiments, audio rendering parameters may include panning position (horizontal orientation between the left and right channels), volume, etc.
[0123] In some embodiments, the central processing unit (CPU) can determine audio rendering parameters based on a connectivity graph using various methods. For example, the CPU can query a second preset table to determine the audio rendering parameters based on the connectivity relationships between the target module and the main module corresponding to the audio track to be processed in the connectivity graph. The second preset table records the correspondence between connectivity relationships and audio rendering parameters.
[0124] The second preset table can be preset based on human experience. For example, for a target module to the left of the main control module, the pan position of its rendering parameters can be -50% (leaning to the left). The larger the relative cascading position of the target module (i.e., the larger the jump number parameter), the smaller the relative volume of its rendering parameters.
[0125] The updated audio track to be processed refers to the audio data obtained after the audio track to be processed has been processed by the audio rendering parameters.
[0126] In some embodiments, the central processing unit applies audio rendering parameters to the audio track to be processed through a digital signal processing algorithm, thereby generating an updated audio track to be processed.
[0127] In some embodiments, the central processing unit can synthesize the updated audio track to be processed with the current audio stream to obtain the target audio stream. More information on synthesis can be found in the related content above.
[0128] In some embodiments of the specification, by detecting the connection relationship graph, the central processing unit can perceive the connection relationship of each functional sub-module, determine the audio rendering parameters based on the connection relationship, thereby creating collaborative music with a sense of three-dimensional space and depth, significantly improving the immersiveness and fun of creation and playback.
[0129] Step 340: Control the audio playback unit to play the target audio stream.
[0130] In some embodiments, the central processing unit outputs the target audio stream to the audio playback unit. The digital-to-analog converter (DAC) circuit within the audio playback unit receives the target audio stream and performs DAC conversion, converting the target audio stream into a continuous analog voltage signal. A power amplifier amplifies the continuous analog voltage signal, and the amplified analog voltage signal is transmitted to the speaker. The speaker, under the influence of the amplified analog voltage signal, converts the analog voltage signal into audible sound to complete the playback. More information about the DAC circuit, power amplifier, and speaker can be found in Figure 2 and related documentation.
[0131] In some embodiments, the central processing unit employs a streaming architecture. Once the target audio stream has been partially synthesized in the memory buffer, this completed audio data stream is transmitted to the audio playback unit in real time for playback, without waiting for the entire audio data to be synthesized. This approach effectively reduces audio playback latency and enables real-time processing of synthesis and playback simultaneously.
[0132] In some embodiments of the instruction manual, the abstract concept of an audio track is transformed into physical modules, allowing users such as children to intuitively complete music composition through physical assembly actions, thus realizing the materialization and visualization of music creation. The recording module allows users to incorporate their own voices as instruments into their creations, and the high degree of freedom in combining these physical modules provides limitless possibilities for interaction and creation.
[0133] In some embodiments, the central processing unit is further configured to perform a pre-release of the audio track to be processed in response to detecting a disconnection between the target module and the main control module or the connected sub-module.
[0134] In some embodiments, disconnection means that the connection between the target module and the main control module or the connected sub-module is removed.
[0135] In some embodiments, the central processing unit or user can separate the second type interface (or first type interface) of the target module from the first type interface of the main control module or the second type interface (or first type interface) of the connected sub-module, thereby achieving connection disconnection.
[0136] Pre-release refers to the process of smoothly reducing the volume of the audio track to be processed to mute within a preset time period. The preset time is based on human experience.
[0137] In some embodiments, when the connection between the target module and the main control module is disconnected or the connection with the connected sub-module is disconnected, the central processing unit can gradually reduce the volume of the audio track to be processed from the current value to zero in a linear or exponential manner, thereby achieving pre-release.
[0138] In some embodiments, to achieve a smooth audio transition when the target module is disconnected, the central processing unit (CPU) can continuously monitor the connection status of each interface. When the detection pin level changes from low to high (indicating impending disconnection), the CPU immediately initiates a pre-release. For example, within 200 milliseconds before disconnection, the gain of the corresponding audio track is gradually reduced to mute using an exponential decay curve. During the pre-release process, the CPU synchronously updates the mixer parameters to maintain overall volume balance.
[0139] In some embodiments of the specification, the central processing unit can effectively prevent the sound from stopping abruptly due to the sudden removal of the module by pre-releasing when the connection between the target module and the main control module is disconnected or when the connection with the connected sub-module is disconnected, making the transition of music smoother and more natural, and improving the continuity and artistry of the overall listening experience.
[0140] Figure 4 is a schematic diagram of the control flow of different functional sub-modules according to some embodiments of this specification.
[0141] In some embodiments, the special effects module is provided with an acceleration sensor, and the central processing unit 102 is further configured to: in response to the target module being a special effects module, acquire acceleration data monitored by the acceleration sensor; determine a special effects event based on the acceleration data; and in response to the special effects event satisfying a first preset condition, use a preset audio track of the target module's identity information as an audio track to be processed.
[0142] For more information on the target module, effects module, and audio track to be processed, please refer to Figures 2 and 3 and their related descriptions.
[0143] In some embodiments, the central processing unit can identify the identity information of the target module and determine the type of the target module based on the identity information. In response to the target module being a special effects module, the central processing unit acquires acceleration data through the accelerometer in the special effects module.
[0144] In some embodiments, acceleration is generated by a user shaking effect module. When the user shakes the effect module, the acceleration sensor detects the attitude change signal, measures the acceleration data, and uploads it to the main control module.
[0145] Special effects events refer to the trigger events that activate the preset audio tracks of the special effects module.
[0146] In some embodiments, the central processing unit can monitor changes in acceleration data of the special effects module in real time using an accelerometer. When the value or magnitude of the acceleration data exceeds a preset intensity threshold, a special effects event is determined to have occurred.
[0147] The first preset condition refers to the time interval between two adjacent special effects events triggered by the same special effects module being less than the preset duration (e.g., 10 seconds).
[0148] In some embodiments, in response to two adjacent special effects events triggered by the same special effects module satisfying a first preset condition, the central processing unit uses the preset audio track corresponding to the identity information of the special effects module as the audio track to be processed.
[0149] In some embodiments, the central processing unit can also perform dual filtering on the accelerometer data, including: firstly, smoothing the raw data using a moving average filter, and then setting a dynamic intensity threshold to determine whether a special effect event is triggered. To prevent false triggering, the central processing unit can also introduce a debounce mechanism, such as: only when the value or amplitude of the acceleration data sampled three consecutive times exceeds a preset intensity threshold is it recorded as a valid event, and a cooling timer (e.g., 300 milliseconds) is started, ignoring the trigger signal of the same special effect module during the cooling period. Users can adjust the intensity threshold and cooling time through the user terminal to adapt to different usage habits.
[0150] In some embodiments of this specification, an acceleration sensor is integrated into the special effects module, and special effects events are determined based on acceleration data. When a special effects event meets a first preset condition, the corresponding special effects track is automatically invoked, which enriches the interactive forms and sound effects layers of music creation, enhances the creative fun, and reduces the risk of accidental triggering.
[0151] In some embodiments, when the target module is the instrument module 20 or the recording module 30, the audio track to be processed is configured to loop playback mode; the central processing unit 102 is further configured to: monitor the playback progress of the audio track to be processed; when the playback progress of the audio track to be processed meets the second preset condition, at a preset time, reset the playback progress of the audio track to be processed, and synthesize the audio track to be processed with the current audio stream to obtain the updated target audio stream; and control the audio playback unit to play the updated target audio stream.
[0152] For more information on the instrument module, recording module, current audio stream, and target audio stream, please refer to the previous text and its related descriptions.
[0153] The loop playback mode refers to the mode in which the audio track to be processed is played repeatedly according to the complete audio cycle. That is, the playback will automatically restart after the audio track finishes playing, so as to achieve continuous audio output.
[0154] Playback progress is used to indicate the degree to which the audio track being processed has been played.
[0155] In some embodiments, the central processing unit can obtain the current playback position information of the audio track to be processed by reading the playback timing data of the audio track to be processed in real time, and determine the playback progress of the audio track to be processed.
[0156] The second preset condition can be related to the playback progress of the audio track to be processed, such as when the playback progress of the audio track to be processed reaches a preset threshold. For example, the preset threshold can be 90% or 100% of the total duration of the audio track to be processed.
[0157] In some embodiments, if the playback progress of the audio track to be processed meets the second preset condition, it indicates that the audio track to be processed is about to finish or has finished a full playback, and a playback progress reset operation needs to be triggered to ensure the continuity of loop playback.
[0158] The preset time refers to the moment when the playback progress of the audio track to be processed meets the second preset condition and is synchronized with the beat cycle of the current audio stream. For example, the preset time can be the start point of the beat of the current audio stream.
[0159] In some embodiments, the central processing unit can send control commands to restore the current playback position of the audio track to the beginning position of the audio track, thereby resetting the playback progress.
[0160] In some embodiments, at a preset time, the central processing unit can perform a synthesis operation between the audio track to be processed after resetting the playback progress and the current audio stream to generate an updated target audio stream.
[0161] In some embodiments, the central processing unit may send a playback control command to the audio playback unit to control the audio playback unit to stop playing the original target audio stream and instead play the updated target audio stream.
[0162] In some embodiments of this specification, by configuring the unprocessed audio tracks of the instrument module and recording module into a loop playback mode and synchronously resetting the progress and synthesized audio at a preset time, the continuity of loop playback is ensured, while ensuring the rhythm synchronization and harmony of multi-track synthesis, thereby improving the auditory effect of music creation.
[0163] Figure 5 is an exemplary schematic diagram illustrating the determination of a target audio stream according to some embodiments of this specification.
[0164] In some embodiments, the central processing unit may determine the target beat count 510 and the target key 520, and determine the processed audio track 550 based on the initial beat count 530 and the initial key 540 of the audio track to be processed, as well as the target beat count 510 and the target key 520; at a preset time, the processed audio track 550 is synthesized with the current audio stream 560 to obtain the target audio stream 570.
[0165] For more information about target audio stream 570, please see Figure 3 and its related description.
[0166] The initial beat count of 530 is a parameter describing the original playing tempo of the track to be processed, in beats per minute (BPM). For example, an initial beat count of 120 BPM means that the track has 120 beats per minute.
[0167] The initial key 540 is a parameter that describes the original performance pitch reference of the track to be processed, such as C major, A minor, etc.
[0168] The target beat count of 510 refers to the global tempo of all tracks when the central processing unit is used to unify the tempo of all tracks during synthesis.
[0169] The target key 520 refers to the global pitch reference used by the central processing unit to unify all tracks during synthesis.
[0170] In some embodiments, the central processing unit may determine the target beat number 510 and the target key 520 in a variety of ways.
[0171] In some embodiments, the central processing unit may use the number of beats of the reference audio stream as the target number of beats 510 and the key of the reference audio stream as the target key 520.
[0172] In some embodiments, the central processing unit may use the beat count of the preset audio track corresponding to the first functional sub-module connected to the main control module as the target beat count 510, and the tonic of the preset audio track as the target tonic 520. For more information on preset audio tracks and reference audio streams, please refer to Figure 3 and its related description.
[0173] In some embodiments, the central processing unit may, based on preset weights, weight the initial number of beats of the reference audio stream and the initial number of beats of the connected sub-modules to determine the target number of beats; and, based on preset weights, weight the initial key of the reference audio stream and the initial key of the connected sub-modules to determine the target key.
[0174] Preset weights refer to the influence coefficients assigned to preset tracks of different modules when calculating the target beat count or target key.
[0175] In some embodiments, the central processing unit can interact with the user terminal through a communication component to obtain the preset weight of the user input. Further description of the communication component can be found in Figure 2 and its related description.
[0176] In some embodiments, the central processing unit may weight the initial beat count of the reference audio stream and the initial beat count of the connected sub-modules according to a preset weight, and use the weighted result as the target beat count.
[0177] In some embodiments, the central processing unit may weight the initial key of the reference audio stream and the initial key of the connected sub-modules according to preset weights, and use the weighted result as the target key 520.
[0178] In some embodiments of the specification, the central processing unit introduces user-adjustable preset weights, which users can flexibly adjust according to their creative intentions, thus achieving personalized collaborative creation.
[0179] The processed audio track is the audio track to be synthesized with the current audio stream.
[0180] In some embodiments, the central processing unit can sequentially perform time scaling and modulation processing on the audio track to be processed according to the target beat count and target key, to obtain the processed audio track.
[0181] The preset time is the point in time when the central processing unit synthesizes the processed audio track with the current audio stream.
[0182] In some embodiments, the central processing unit can use the moment when the global clock signal arrives at the nearest next beat point or the start point of the next measure in the current audio stream as a preset time. The global clock signal is a hardware-level high-precision timer signal generated by the central processing unit of the main control module and serves as the time reference of the device. For example, the beat point can be a quarter note position, and the start point of a measure can be the first beat in 4 / 4 time.
[0183] In some embodiments, the central processing unit may synthesize the processed audio track with the current audio stream at a preset time to obtain the target audio stream. For more details on how this synthesis is performed, please refer to Figure 3 and its related content.
[0184] In some embodiments of the specification, the central processing unit (CPU) solves the problems of rhythmic inconsistency and tonality conflicts that easily occur in multi-module collaboration by unifying the audio tracks to be processed to the target beat count and target key. By introducing a preset timing method for synthesis, when a new functional sub-module is connected, the CPU does not immediately output sound, but enters a "standby state," ensuring that the new audio track always precisely enters on the beat of the music, resulting in a synthesized effect with a strong sense of rhythm and natural transitions.
[0185] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.
[0186] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.
[0187] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.
[0188] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.
[0189] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values are set as precisely as feasible.
[0190] For each patent, patent application, patent application publication, and other material such as articles, books, specifications, publications, and documents referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.
[0191] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.
Claims
1. A modular music creation and playback device, characterized in that, The system includes a main control module and functional sub-modules. The functional sub-modules are detachably connected to the main control module. Each functional sub-module includes a first type of interface and a second type of interface, and stores identity information corresponding to a preset audio track. The main control module includes a central processing unit, an audio playback unit, and the first type of interface. The central processing unit is configured to: in response to a target event, read the identity information of the target module; the target event includes the second type of interface of the target module being connected to the first type of interface of the main control module; or the second type of interface of the target module being connected to the first type of interface of an already connected sub-module; determine the audio track to be processed based on the identity information of the target module; synthesize the audio track to be processed with the current audio stream to obtain a target audio stream; and control the audio playback unit to play the target audio stream.
2. The apparatus according to claim 1, characterized in that, The functional sub-modules include at least one of the following: musical instrument module, recording module, and special effects module.
3. The apparatus according to claim 2, characterized in that, When the target module is the instrument module or the recording module, the audio track to be processed is configured to loop playback mode; the central processing unit is further configured to: monitor the playback progress of the audio track to be processed; when the playback progress of the audio track to be processed meets the second preset condition, at a preset time, reset the playback progress of the audio track to be processed, and synthesize the audio track to be processed with the current audio stream to obtain an updated target audio stream; and control the audio playback unit to play the updated target audio stream.
4. The apparatus according to claim 1, characterized in that, The audio track to be processed includes an initial number of beats and an initial key; the central processing unit is further configured to: determine a target number of beats and a target key; and determine a processed audio track based on the initial number of beats and the initial key, as well as the target number of beats and the target key. At a preset time, the processed audio track is combined with the current audio stream to obtain the target audio stream.
5. The apparatus according to claim 1, characterized in that, The first type of interface and the second type of interface are adapter interfaces; wherein, one of the first type of interface and the second type of interface is provided with a connector and the other is provided with a metal contact; both the first type of interface and the second type of interface are provided with a ring magnet, a power supply pin and a data communication pin.
6. The control method for the modular music creation and playback device as described in claim 1, characterized in that, include: In response to the occurrence of a target event, the identity information of the target module is read; the target event includes the second type interface of the target module being connected to the first type interface of the main control module; or the second type interface of the target module being connected to the first type interface of the connected sub-module; based on the identity information of the target module, the audio track to be processed is determined; the audio track to be processed is combined with the current audio stream to obtain the target audio stream; and the audio playback unit is controlled to play the target audio stream.
7. The method according to claim 6, characterized in that, When the target module is a musical instrument module or a recording module, the audio track to be processed is configured in a loop playback mode. The method further includes: monitoring the playback progress of the audio track to be processed; when the playback progress of the audio track to be processed meets a second preset condition, resetting the playback progress of the audio track to be processed at a preset time, and merging the audio track to be processed with the current audio stream to obtain an updated target audio stream; and controlling the audio playback unit to play the updated target audio stream.
8. The method according to claim 6, characterized in that, The audio track to be processed includes the initial beat count and the initial tonic key; The step of combining the audio track to be processed with the current audio stream to obtain the target audio stream includes: determining the target number of beats and the target key; determining the processed audio track based on the initial number of beats and the initial key of the audio track to be processed, as well as the target number of beats and the target key; and combining the processed audio track with the current audio stream at a preset time to obtain the target audio stream.
9. A control device for a modular music creation and playback apparatus, comprising a processing unit, characterized in that, The processing device is used to execute the control method of the modular music creation and playback device as described in any one of claims 6-8.
10. A computer-readable storage medium storing computer instructions that, when executed by a processor, implement the control method of the modular music creation and playback apparatus as described in any one of claims 6-8.