Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

59 results about "MIDI" patented technology

MIDI (/ˈmɪdi/; short for Musical Instrument Digital Interface) is a technical standard that describes a communications protocol, digital interface, and electrical connectors that connect a wide variety of electronic musical instruments, computers, and related audio devices for playing, editing and recording music. A single MIDI link through a MIDI cable can carry up to sixteen channels of information, each of which can be routed to a separate device or instrument. This could be sixteen different digital instruments, for example.

Construction method of vem-token vocal emotion multimodal magic modification model

The construction method of the VEM-Token vocal emotion multi-modal magic modification model is different from the natural language processing model NLP-Token which interprets music through text. Instead, the VEM-Token sound-to-text innovation model comes with vocal emotion multi-modal information. The model captures and aligns the music tempo of sample songs and user learning songs, identifies various modalities of vocal emotion, divides the file into word units based on tempo, obtains VEM parameters through supervised learning and reinforcement learning, and decomposes songs into vocals, accompaniment, and emotion. The magic modification model provides a multi-modal magic modification method for imitating sample songs, including vocals, accompaniment, emotion overtones, emotion fluctuations, learning to sing, voice cloning, lyrics, pitch calibration, ornaments, tempo length, rhythm speed, tempo strength, freedom, and multiple sample magic modification. It also provides member management, mobile and PC application systems, dedicated support hardware, and communication protocols including MIDI, making it easy to access popular AI large models, reducing model hallucinations, and forming AI vocal agents and AI karaoke.
Owner:GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD

Music performance accuracy evaluation system based on multi-modal signal fusion

The invention relates to the technical field of artificial intelligence and music information processing, in particular to a multi-modal signal fusion music playing accuracy evaluation system. The system comprises a multi-modal signal synchronous acquisition module, a signal preprocessing and feature extraction module, a time domain alignment and event association engine, a hierarchical collaborative evaluation module and a comprehensive evaluation report generation module. Through high-precision synchronous acquisition of an MIDI instruction, an acoustic signal and a structural vibration signal, a nonlinear time domain alignment mechanism taking dynamic time warping as a core is constructed, accurate matching of a symbol event and a physical response is realized, and an alignment multi-modal event frame is formed; and collaborative quantitative evaluation of basic accuracy, technical skill and music expressive force is carried out through the hierarchical model. By adopting the technical scheme, full-link high-fidelity analysis from playing intention to physical presentation can be realized, and the accuracy, robustness and interpretability of evaluation are remarkably improved.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Singing detection method and singing detection system using the same

A singing detection method and a singing detection system using the same are provided in the embodiments of the present invention. The singing detection method includes the following steps: performing packet trigger judgment to activate the detection function; conducting rhythm detection on the sound signal; performing pitch detection on the sound signal; quantizing the pitch detection data; converting the quantized data and rhythm detection data into MIDI note data; using the MIDI note data to judge rhythm and pitch changes; assigning weighted scores and comparing with a threshold score. Finally, it determines whether the sound signal is singing based on the total score. This method combines multiple musical feature analyses to identify singing voices through comprehensive scoring.
Owner:GENERALPLUS TECH INC

Interactive system for music performance

The invention describes an advanced musical control system based on the identification and tracking of physical objects on an infrared touch surface. Traditional musical controllers offer limited control possibilities. Existing controllers based on object tracking systems, on the other hand, are often bulky (when based on visual detection) or restricted to the use of conductive materials (in the case of capacitive controllers). In the proposed solution, object tracking is achieved by identifying groups of points (protrusions on the objects themselves) through an infrared detection system. The geometric configurations of these points uniquely identify and localize the objects, making it possible to detect data that, once processed through suitable protocols such as MIDI, TUIO, or equivalents, allow the user to control up to three sound parameters with each hand. The invention is intended for interactive musical performance and enables intuitive and compact control of sound parameters through physical objects.
Owner:PAVANELLI ISACCO

Method and system for converting image into music score

The invention relates to a method for converting an image into a music score, which comprises the following steps of: performing note identification on a note image to obtain an MIDI (musical instrument digital interface) file, and performing music score conversion on the MIDI file to obtain music score data. According to the method for converting the image into the music score, the recognized note data can be directly converted into the music score data in a standardized data format used by a music teaching system, the structural consistency of a score recognition result and an original music score is effectively improved, and the generated music score is closer to a manual input standard in visual and logic aspects; and the obtained music score data can also be used for conventional functions of a music teaching system.
Owner:BEIJING JINSANHUI TECH CO LTD

Piano playing audio and video joint generation model based on multi-modal MIDI guidance

The invention belongs to the technical field of multi-modal generation and audio and video synthesis, and relates to a piano playing audio and video joint generation model based on multi-modal MIDI guidance, and the overall structure of the piano playing audio and video joint generation model comprises three core modules: an MIDI to gesture (MIDI-to-Pose) module; a gesture-to-video (Pose-to-Video) generation module is used for generating a gesture (Pose-to-Video); the invention relates to a symbol-to-audio (Symbolic-to-Audio) generation module. The model provided by the invention can be applied to music education, virtual playing, music generation explaining, performance analysis and other scenes. According to the invention, the generated hand action, key knocking and piano sound are strictly aligned in time and semantics, and the real, synchronous and physically consistent piano playing video generation is realized.
Owner:GIANT MOBILE TECH CO LTD

Harmonic texture and sequence reasoning fused audio-to-MIDI method and system

PendingCN121687100ASpeech analysisFrequency spectrumNote value
The invention discloses a harmonic texture and sequence reasoning fused audio-to-MIDI method and system, and relates to the technical field of audio processing. Aiming at the problem of insufficient transcription accuracy of the existing polyphonic music, the method comprises the following steps: carrying out time-frequency transformation on the polyphonic music to generate a spectrogram; constructing a time-frequency target detection network, detecting notes based on harmonic comb texture features, and outputting a time-frequency bounding box; constructing a sequence inference network to extract time sequence context features; carrying out local alignment and weighted fusion on the time sequence context features by using a time-frequency bounding box by adopting an ROI-guided gating fusion mechanism; and inputting the fusion features into a decoder to generate an MIDI sequence, and correcting the note time value by using a time-frequency bounding box. According to the method, the visual texture features and the music sequence logic are effectively fused, and the time precision and robustness of polyphonic music transcription are remarkably improved.
Owner:XINJIANG UNIVERSITY

MIDI and DMX interconversion device

The utility model discloses an MIDI and DMX interconversion device, which comprises an MCU module, a control panel and a plurality of interface modules, wherein the control panel and the interface modules are connected with the MCU module; the interface module is provided with at least one DMX interface used for DMX information communication and at least one MIDI interface used for MIDI information communication, and the MCU module can edit the conversion relation between DMX information and MIDI information through the control panel; when the MCU module receives any one of the DMX information or the MIDI information, the DMX information or the MIDI information can be converted into the corresponding MIDI information or the DMX information according to a preset conversion relation. Compared with the prior art, the interconversion of MIDI and DMX can be realized, so that a music player can quickly operate stage equipment such as lighting equipment, and a stage lamp player can quickly play MIDI musical instruments.
Owner:HUASHI (SHENZHEN) TECH CO LTD

Method and system for extracting duration of singing voice phoneme using midi

There are provided a method and a system for extracting singing voice phoneme duration. A singing voice phoneme duration extraction system using a MIDI according to an embodiment may receive phonemes converted from a text as input, and may output a prior probability distribution, may receive acoustic features as input and may output a posterior probability distribution, may convert the probability distribution, may perform monotonic alignment search by using information on MIDI duration, and may output a waveform which is a voice digital signal, based on input reflecting a result of extracting the phoneme duration.
Owner:KOREA ELECTRONICS TECH INST

Enhanced earplug device with advanced acoustic filtering and smart lighting system

An enhanced earplug device for noise suppression and entertainment lighting comprises a tubular housing with a removable tip portion for ear canal insertion and a visible light element. The device incorporates a quartz acoustic filter for high-fidelity sound preservation while reducing volume, and a vibration sensor detecting musical beats in the 20-250 Hz range for music-reactive lighting patterns. Bluetooth connectivity enables inter-device synchronization and mobile app control. A tap sensor integrated into the housing exterior provides intuitive gesture-based mode switching between static color, music-reactive, and synchronized operational modes. The system includes RF capabilities for professional venue integration via centralized controllers with DMX / MIDI connectivity, enabling coordinated lighting displays across multiple users in entertainment environments.
Owner:DENNIS JOHN MAXWELL NORMAN

Singing detection method and singing detection system using same

The invention provides a singing detection method and a singing detection system using the same. The singing detection method comprises the following steps: performing packet triggering judgment to start a detection function; performing rhythm detection on the sound signal; carrying out fundamental tone detection on the sound signal; carrying out quantization operation on the fundamental tone detection data; converting the quantized data and the rhythm detection data into MIDI note data; performing rhythm and pitch change judgment by using MIDI note data; weight scores are given and compared with threshold scores. And finally, judging whether the sound signal is singing according to the total score. According to the method, analysis of multiple music features is combined, singing sounds are recognized through comprehensive scores, singing sounds and non-singing sounds can be effectively distinguished without a high-performance processor or a large number of memories, and the application range of the singing detection technology in low-power-consumption and small-size equipment is greatly expanded.
Owner:GENERALPLUS TECH INC

Electronic musical instrument keyboard

The utility model provides an electronic musical instrument keyboard, including the face shell and bottom shell who butt joint into the casing, still include gyro wheel support structure, be located in the bottom shell inside, functional potentiometer, rotatable installation in gyro wheel support structure, cross section is semicircular functional gyro wheel, functional gyro wheel bottom direct surface sinks in the bottom shell inside, functional gyro wheel top stretches out the face shell upwards, functional gyro wheel center installation in functional potentiometer, gyro wheel reset spare, connect in gyro wheel support structure and functional gyro wheel between, reset buffer spare, be located in gyro wheel reset spare and gyro wheel support structure junction. The utility model provides electronic musical instrument keyboard under the diameter and rotation angle stroke limit optimization functional gyro wheel layout, realize MIDI keyboard thinning and portable, eliminate the larger rebound noise of shorter elastic arm in the limited space, satisfy individualized color selection while reducing the cost, realize assembly tolerance control, avoid partial deformation, warping in the assembly process, realize circuit board device integration module stacking and guarantee the plug assembly in place.
Owner:SHENZHEN ARTFAST TECH CO LTD

Multifunctional time code synchronization device

The utility model discloses a multifunctional time code synchronizer, which comprises an MCU (Microprogrammed Control Unit) module, a control panel connected with the MCU module, a power supply module connected with the MCU module, a plurality of interface modules and an interface circuit connected between the interface modules and the MCU module, the interface module at least comprises an MIDI IN interface and an MIDI OUT interface which can be used for transmitting an MIDI time code, an LTC IN interface and an LTC OUT interface which can be used for transmitting an SMPTE LTC time code, and an Ethernet port which can be used for transmitting an RTP MIDI time code and an ArtTimecode time code, and the MCU module can transmit and synchronize corresponding time code information through an interface circuit connected with the MIDI IN interface, the MIDI OUT interface, the LTC IN interface, the LTC OUT interface and the Ethernet port. Compared with the prior art, the device has the synchronization of the MIDI time code and the SMPTE LTC time code, and also has the network time codes of the ArtTimecode and the RTP-MIDI, so that the LTC time code can be transmitted more flexibly.
Owner:HUASHI (SHENZHEN) TECH CO LTD

Empathy-Based Wearable Systems for Perceptual and Communicative Assistance

PendingUS20260179644A1Speech recognitionHeadphonesEmpathy
The invention provides a unified assistive technology system, the ADA Empathy Wearable Bundle, enhancing perception, communication, and self-expression for individuals with visual, auditory, or speech impairments.It integrates Empathy Glasses, Empathy Headphones, Word Articulation Guidance (WAGS), Word Articulation Response Monitoring (W.A.R.M.), and Conversational MIDI (C-MIDI), using AI-driven narrative interpretation, augmented reality, and blockchain-based security.The system delivers vivid narrative audio for blind users, multimodal AR overlays for hearing-impaired users, and real-time visual and auditory articulation guidance for speech-impaired or non-verbal users. A tone-based trust protocol ensures operational integrity, while blockchain storage preserves privacy and consent.Grounded in ethical AI, empathy, and human dignity, the system enables users to perceive, understand, and engage with the world fully, offering a holistic, compassionate, and transformative solution in assistive technology.
Owner:BOWES DEAN

Symbol music generation method based on texture perception and adaptive representation alignment

The invention discloses a symbol music generation method based on texture perception and adaptive representation alignment. The method is characterized by comprising the steps of conditional decoupling and extraction of multi-scale music element features, conditional coding, texture perception FLUX music model construction, adaptive representation alignment, MIDI file generation and the like. Compared with the prior art, the method has the advantages that the structural advantage and multi-scale information of the symbol music are fully utilized, the texture consistency, chord accuracy and style controllability of the generated music are remarkably improved, the training and reasoning cost is reduced, and the generalization ability and practicability of the model are further improved. Through the three-in-one design of'conditional decoupling, texture perception diffusion and self-adaptive expression alignment ', the problems of semantic loss, inconsistent texture, uncontrollable chord and the like caused by neglecting of internal association of music in a traditional method are effectively solved, and the method is efficient, reliable and easy to implement and has good application prospects in the fields of automatic composition, interactive music creation and the like.
Owner:EAST CHINA NORMAL UNIV

AI light robot (piano version)

1. The name of the design product: AI light robot (piano version). 2. The use of the design product: can be connected to the keyboard through Bluetooth or midi for piano use, can also be played automatically according to the music score through AI, the corresponding piano keys will be lighted when playing, and can also be used in various entertainment places, birthday parties and other places to create light atmosphere and play music. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view. 5. Other circumstances that need to be explained: A is transparent material.
Owner:江青

No-reference sound correction method based on music score big language model

The invention discloses a no-reference sound correction method based on a music score big language model, and relates to a mismatch song correction method in the technical field of intelligent audio processing. The objective of the invention is to solve the problems of poor pitch and rhythm collaborative correction effect, low note feature extraction precision and easy loss of emotional expression in a non-reference scene in the existing sound correction technology. The method comprises the following steps: S1, acquiring and preprocessing off-tune singing audio data; s2, carrying out note accurate segmentation on the preprocessed audio; s3, note level features are extracted, and Octuple-MIDI symbolization input is generated; s4, realizing non-reference pitch-rhythm joint error correction through a music score big language model, and outputting a standard MIDI sequence; s5, based on a standard MIDI sequence, completing cooperative accurate correction of the sound pitch and the rhythm of the song; and S6, outputting the natural corrected singing sound through a high-fidelity synthesis technology. The method is used in the field of off-tune singing processing such as karaoke entertainment, AI music creation, singing sound correction and the like.
Owner:HARBIN INST OF TECH

Pendant (MIDI)

1. Name of the designed product: pendant (MIDI). 2. Use of the designed product: for wearing decoration. 3. Design points of the designed product: in shape. 4. Picture or photo best indicating the design points: perspective view 1.
Owner:SHENGGE (SHANGHAI) TECHNOLOGY CO LTD

Audio recognition method and device, computer device and computer readable storage medium

The application discloses an audio recognition method and device, computer equipment and a computer readable storage medium, and applies to the technical field of computers. The method comprises the following steps: after obtaining a cappella audio data, extracting a fundamental frequency sequence of the cappella audio data; obtaining a musical instrument digital interface (MIDI) template of a song corresponding to the cappella audio data, wherein the MIDI template is used for representing standardized music parameters of the song; adjusting a sound area corresponding to the fundamental frequency sequence and a sound production speed corresponding to the fundamental frequency sequence based on the MIDI template; calculating a matching degree between the adjusted fundamental frequency sequence and the MIDI template; and determining a recognition result of the cappella audio data according to the matching degree, wherein the recognition result is used for indicating whether the cappella audio data is qualified. Through the application, the accuracy of audio evaluation can be improved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

A music generation method based on a large pre-training language model

The present application relates to the technical field of language processing, and more particularly to a music generation method based on a large pre-training language model. The GPT-3 is fine-tuned by using thousands of melody MIDI files, and then the fine-tuned model is used for melody generation. The main advantages of the method are as follows: (1) the algorithm can learn the long-term dependency structure of the melody and generate music with long-term structure and musicality; (2) the algorithm can simulate different melody generation methods by adjusting the fine-tuning data format; (3) the algorithm allows only a small amount of data to be used to generate similar style melodies.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Electronic device with graphical user interface for editing a polyphonic panflute panel of a stringed instrument

1. The name of the design product: the graphical user interface of the electronic device with the riffer MIDI editor panel of stringed musical instrument. 2. The use of the design product: an electronic device. 3. The design points of the design product: the graphical user interface. 4. The picture or photo that best shows the design points: the front view. 5. The use of the graphical user interface: to display the relevant information of the riffer MIDI editor panel of stringed musical instrument. 6. The human-computer interaction mode of the graphical user interface: the main view is the main interface view, and clicking the first blue "switch mode" button from left to right in the upper right corner of the main view page can enter the interface change state diagram.
Owner:BEIJING BOSHENGYUAN TECH CO LTD

Graphical user interface [computer screen layout]

This image is mainly used as an operation and display image for the music editing view of music composition software. At the bottom, the song track displays the song representation and provides song editing tools, and the groove (MIDI file and chords) is represented by blocks. The buttons and tools in the arrangement section allow the user to arrange the song by groove editing. The middle section allows a preview, and allows the user to replace the selected song track MIDI groove with the groove selected within the song track, while maintaining the song track chords. Figure 9.2 is a reference diagram showing the names of the parts. 9) Image; 9.2) Reference diagram
Owner:TOONTRACK MUSIC

A MIDI music generation method, device and terminal equipment

The application is suitable for the technical field of computers, and provides a MIDI music generation method, device and terminal equipment, which comprises the following steps: preprocessing original MIDI music data to obtain a note track and a chord track; obtaining MCST elements and chord elements based on the note track and the chord track; obtaining a first prediction model and a second prediction model based on the MCST elements and the chord elements; inputting a random MIDI music segment into the first prediction model to generate an intermediate MCST segment; inputting the generated intermediate MCST segment into the second prediction model to perform chord replacement and generate new MIDI music. The application can improve the fluency and harmony of music and also enrich the diversity of music.
Owner:HEBEI UNIV OF SCI & TECH +1

Cross-platform mobile power supply data interaction and state monitoring system and method

The invention relates to the technical field of mobile power supply monitoring and data communication, in particular to a cross-platform mobile power supply data interaction and state monitoring system and method, and aims to solve the problems of unreliability, high delay and poor compatibility of cross-platform communication caused by limitation of a safety mechanism of an operating system in the prior art. According to the system, a triple redundancy communication architecture which is based on a standard MIDI protocol and fuses an MTP protocol and gamepad equipment simulation is constructed, sensor data is embedded into an MIDI SysEx message through a displacement coding algorithm, and millisecond-level safety warning is achieved through a kernel-level audio path; and meanwhile, dynamic switching of communication protocols is supported, and drive-free, low-delay and high-reliability data interaction and active safety intervention of the whole platform are ensured.
Owner:SHENZHEN I4SEASON HONGSHENG TECH CO LTD

Musical instrument playing evaluation method, device, equipment, medium and product

The invention discloses a musical instrument playing evaluation method, device and equipment, a medium and a product. The method comprises the following steps: determining an acoustic feature sequence corresponding to an audio signal played by a user; extracting an audio semantic feature sequence of the acoustic feature sequence through an audio semantic extraction model; generating a first MIDI vector sequence corresponding to the audio semantic feature sequence through an MIDI generation network based on a context dependency relationship of the audio semantic feature sequence; encoding reference MIDI data corresponding to the standard music score into a reference MIDI vector sequence through an MIDI encoding network; and based on the first MIDI vector sequence and the reference MIDI vector sequence, generating an evaluation result of the playing of the user. According to the embodiment of the invention, intelligent evaluation of music playing can be realized.
Owner:XIAN JIAOTONG LIVERPOOL UNIV

Effects unit (EC-2)

1. The name of the design product: effecter (EC-2). 2. The use of the design product: the design product is used for effecter which can send MIDI control information. 3. The design points of the design product: in shape. 4. The picture or photo which can best show the design points: perspective view.
Owner:CHANGSHA HOTONE AUDIO

Cylindrical Digital Audio Workstation (C-DAW)

The present invention relates to a digital audio workstation (DAW) system and method for sequencing, synchronizing, and visualizing musical, conversational, and multimedia content on a cylindrical timeline rather than a traditional linear interface. Each track is mapped to a rotating cylinder, wherein bar length defines the circumference and tempo determines angular velocity. Multiple tracks of arbitrary lengths and independent tempos can run concurrently in perfect synchronization using a shared global time reference. The system inherently supports polymeter and polyrhythm, with events such as notes, gestures, or conversational turns encoded as angular positions and arc lengths. Conversational MIDI (C-MIDI) is integrated to encode emotion, tone, tempo, and phrasing for multi-participant dialogues. The invention further supports distributed orchestration, supercycle detection, AI-driven media generation, and applications extending to music composition, conversation playback, narrative sequencing, and planetary or galactic cyclic modeling.
Owner:BOWES DEAN

MIDI synthesizer

The utility model provides an MIDI synthesizer comprising a first circuit board and a silica gel button assembly, and the front surface of the first circuit board is provided with a plurality of touch switches; the silica gel key assembly comprises a key supporting plate and a plurality of silica gel keys; wherein the key supporting plate is provided with a plurality of key installation positions, the silica gel key comprises a silica gel key cap and at least one auxiliary supporting column, the cap body part of the silica gel key cap protrudes out of the key supporting plate, the cap edge part of the silica gel key cap extends into the corresponding key installation position, and a containing space is defined by the silica gel key cap and the key supporting plate; the touch switch and the auxiliary supporting column are located in the storage space, one end of the auxiliary supporting column is connected with the cap body part of the silica gel key cap, and a distance is kept between the other end of the auxiliary supporting column and the first circuit board; the silica gel button is provided with the auxiliary supporting column, the auxiliary supporting column can avoid the problem that the silica gel button sinks due to the fact that a user presses the edge corner of the silica gel button by mistake, and the use experience of the user is improved.
Owner:SHENZHEN ARTFAST TECH CO LTD

Music source extraction method, device and product based on reference audio and MIDI guidance

The invention provides a music source extraction method, device and product based on reference audio and MIDI guidance, and the method comprises the steps: processing an obtained mixed audio signal and a reference audio signal through a preset music source extraction model, and outputting a separated audio signal which corresponds to the reference audio signal; the music source extraction model processing step comprises the steps of splicing features extracted from the mixed audio signal and the reference audio signal from MIDI features extracted from the mixed audio signal to obtain multi-scale fusion features; carrying out multiple times of feature extraction on the multi-scale fusion features with different time resolutions to obtain cross-resolution representation; the cross-resolution representation is processed by a multi-resolution mask decoder to obtain target masks of different time resolutions, and required separated audio signals are obtained based on the target masks. According to the method, the fuzzy problem caused by spectrum overlapping is effectively relieved by using the MIDI features, and the fidelity of the separated audio signal is improved by combining the reference audio signal as a timbre condition.
Owner:TRUE SPACE (ZHUHAI) TECH CO LTD

Wind instrument adapter method based on reinforcement learning under multiple agents

The invention discloses a wind instrument matching method based on reinforcement learning under multiple agents. The method comprises the following steps: extracting music features according to an original MIDI project file, and constructing a music feature model; constructing an expert data set according to the music feature model, and constructing a pre-training model according to the expert data set; constructing a wind instrument dictionary according to a wind instrument method, and constructing a multi-agent reinforcement learning model training environment; constructing a multi-agent reinforcement learning model; training a multi-agent reinforcement learning network in a mode of constructing a target network; according to the invention, the music features can be more comprehensively sorted, and the sorted music features are applied to the automatic matching device of the wind instrument.
Owner:CHENGDU UNIV OF INFORMATION TECH