Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

75 results about "MIDI" patented technology

MIDI (/ˈmɪdi/; short for Musical Instrument Digital Interface) is a technical standard that describes a communications protocol, digital interface, and electrical connectors that connect a wide variety of electronic musical instruments, computers, and related audio devices for playing, editing and recording music. A single MIDI link through a MIDI cable can carry up to sixteen channels of information, each of which can be routed to a separate device or instrument. This could be sixteen different digital instruments, for example.

Construction method of vem-token vocal emotion multimodal magic modification model

The construction method of the VEM-Token vocal emotion multi-modal magic modification model is different from the natural language processing model NLP-Token which interprets music through text. Instead, the VEM-Token sound-to-text innovation model comes with vocal emotion multi-modal information. The model captures and aligns the music tempo of sample songs and user learning songs, identifies various modalities of vocal emotion, divides the file into word units based on tempo, obtains VEM parameters through supervised learning and reinforcement learning, and decomposes songs into vocals, accompaniment, and emotion. The magic modification model provides a multi-modal magic modification method for imitating sample songs, including vocals, accompaniment, emotion overtones, emotion fluctuations, learning to sing, voice cloning, lyrics, pitch calibration, ornaments, tempo length, rhythm speed, tempo strength, freedom, and multiple sample magic modification. It also provides member management, mobile and PC application systems, dedicated support hardware, and communication protocols including MIDI, making it easy to access popular AI large models, reducing model hallucinations, and forming AI vocal agents and AI karaoke.
Owner:GREATER BAY AREA STAR BIOTECH (SHENZHEN) CO LTD

A note-level automatic singing transcription method based on target detection and language features

The application provides a note-level automatic singing transcription method based on target detection and language characteristics, comprising the following steps: step 1: converting a one-dimensional audio sequence into two-dimensional mel spectrum slice and phoneme posterior slice with similar width-height ratio through a preprocessing method of mel transformation, phoneme classification, linear intensity mapping and slicing; step 2: performing target detection on the mel spectrum slice and phoneme posterior slice, post-processing and time adjustment on the left and right boundaries of the boundary box obtained by target detection, and then obtaining the final start time and end time through decision screening; step 3: taking the lower boundary of the target detection boundary box as the fundamental frequency, obtaining the final fundamental frequency through peak value search, and then converting the final fundamental frequency to obtain the MIDI pitch value. The method can effectively improve the phoneme feature extraction effect, improve the phoneme posterior slice quality and improve the feature extraction and analysis effect, thereby improving the transcription accuracy.
Owner:XIDIAN UNIV

Music performance accuracy evaluation system based on multi-modal signal fusion

The invention relates to the technical field of artificial intelligence and music information processing, in particular to a multi-modal signal fusion music playing accuracy evaluation system. The system comprises a multi-modal signal synchronous acquisition module, a signal preprocessing and feature extraction module, a time domain alignment and event association engine, a hierarchical collaborative evaluation module and a comprehensive evaluation report generation module. Through high-precision synchronous acquisition of an MIDI instruction, an acoustic signal and a structural vibration signal, a nonlinear time domain alignment mechanism taking dynamic time warping as a core is constructed, accurate matching of a symbol event and a physical response is realized, and an alignment multi-modal event frame is formed; and collaborative quantitative evaluation of basic accuracy, technical skill and music expressive force is carried out through the hierarchical model. By adopting the technical scheme, full-link high-fidelity analysis from playing intention to physical presentation can be realized, and the accuracy, robustness and interpretability of evaluation are remarkably improved.
Owner:COMMUNICATION UNIVERSITY OF CHINA

Singing detection method and singing detection system using the same

A singing detection method and a singing detection system using the same are provided in the embodiments of the present invention. The singing detection method includes the following steps: performing packet trigger judgment to activate the detection function; conducting rhythm detection on the sound signal; performing pitch detection on the sound signal; quantizing the pitch detection data; converting the quantized data and rhythm detection data into MIDI note data; using the MIDI note data to judge rhythm and pitch changes; assigning weighted scores and comparing with a threshold score. Finally, it determines whether the sound signal is singing based on the total score. This method combines multiple musical feature analyses to identify singing voices through comprehensive scoring.
Owner:GENERALPLUS TECH INC

Interactive system for music performance

The invention describes an advanced musical control system based on the identification and tracking of physical objects on an infrared touch surface. Traditional musical controllers offer limited control possibilities. Existing controllers based on object tracking systems, on the other hand, are often bulky (when based on visual detection) or restricted to the use of conductive materials (in the case of capacitive controllers). In the proposed solution, object tracking is achieved by identifying groups of points (protrusions on the objects themselves) through an infrared detection system. The geometric configurations of these points uniquely identify and localize the objects, making it possible to detect data that, once processed through suitable protocols such as MIDI, TUIO, or equivalents, allow the user to control up to three sound parameters with each hand. The invention is intended for interactive musical performance and enables intuitive and compact control of sound parameters through physical objects.
Owner:PAVANELLI ISACCO

Method and system for converting image into music score

The invention relates to a method for converting an image into a music score, which comprises the following steps of: performing note identification on a note image to obtain an MIDI (musical instrument digital interface) file, and performing music score conversion on the MIDI file to obtain music score data. According to the method for converting the image into the music score, the recognized note data can be directly converted into the music score data in a standardized data format used by a music teaching system, the structural consistency of a score recognition result and an original music score is effectively improved, and the generated music score is closer to a manual input standard in visual and logic aspects; and the obtained music score data can also be used for conventional functions of a music teaching system.
Owner:BEIJING JINSANHUI TECH CO LTD

Vlog generation methods and related devices

This application discloses a Vlog generation method and related apparatus. The method includes: acquiring video footage, the duration of the target Vlog, reference transition points, video emotional information, and a reference audio signal; determining the edited video footage, the transition points of the target Vlog, and reference background music information based on the acquired information; performing track splitting and transcription processing on the reference audio signal to obtain the MIDI score and phrase / section segmentation points of the reference audio signal; processing the reference audio signal according to the transition points of the target Vlog, the reference background music information, the MIDI score of the reference audio signal, and the phrase / section segmentation points of the reference audio signal to obtain the background music of the target Vlog; and obtaining the target Vlog based on the edited video footage and the background music of the target Vlog. Using the method of this application, Vlogs that meet the personalized needs of users can be obtained.
Owner:HUAWEI TECH CO LTD

Piano playing audio and video joint generation model based on multi-modal MIDI guidance

The invention belongs to the technical field of multi-modal generation and audio and video synthesis, and relates to a piano playing audio and video joint generation model based on multi-modal MIDI guidance, and the overall structure of the piano playing audio and video joint generation model comprises three core modules: an MIDI to gesture (MIDI-to-Pose) module; a gesture-to-video (Pose-to-Video) generation module is used for generating a gesture (Pose-to-Video); the invention relates to a symbol-to-audio (Symbolic-to-Audio) generation module. The model provided by the invention can be applied to music education, virtual playing, music generation explaining, performance analysis and other scenes. According to the invention, the generated hand action, key knocking and piano sound are strictly aligned in time and semantics, and the real, synchronous and physically consistent piano playing video generation is realized.
Owner:GIANT MOBILE TECH CO LTD

Harmonic texture and sequence reasoning fused audio-to-MIDI method and system

PendingCN121687100ASpeech analysisFrequency spectrumNote value
The invention discloses a harmonic texture and sequence reasoning fused audio-to-MIDI method and system, and relates to the technical field of audio processing. Aiming at the problem of insufficient transcription accuracy of the existing polyphonic music, the method comprises the following steps: carrying out time-frequency transformation on the polyphonic music to generate a spectrogram; constructing a time-frequency target detection network, detecting notes based on harmonic comb texture features, and outputting a time-frequency bounding box; constructing a sequence inference network to extract time sequence context features; carrying out local alignment and weighted fusion on the time sequence context features by using a time-frequency bounding box by adopting an ROI-guided gating fusion mechanism; and inputting the fusion features into a decoder to generate an MIDI sequence, and correcting the note time value by using a time-frequency bounding box. According to the method, the visual texture features and the music sequence logic are effectively fused, and the time precision and robustness of polyphonic music transcription are remarkably improved.
Owner:XINJIANG UNIVERSITY

MIDI and DMX interconversion device

The utility model discloses an MIDI and DMX interconversion device, which comprises an MCU module, a control panel and a plurality of interface modules, wherein the control panel and the interface modules are connected with the MCU module; the interface module is provided with at least one DMX interface used for DMX information communication and at least one MIDI interface used for MIDI information communication, and the MCU module can edit the conversion relation between DMX information and MIDI information through the control panel; when the MCU module receives any one of the DMX information or the MIDI information, the DMX information or the MIDI information can be converted into the corresponding MIDI information or the DMX information according to a preset conversion relation. Compared with the prior art, the interconversion of MIDI and DMX can be realized, so that a music player can quickly operate stage equipment such as lighting equipment, and a stage lamp player can quickly play MIDI musical instruments.
Owner:HUASHI (SHENZHEN) TECH CO LTD

Method and system for extracting duration of singing voice phoneme using midi

There are provided a method and a system for extracting singing voice phoneme duration. A singing voice phoneme duration extraction system using a MIDI according to an embodiment may receive phonemes converted from a text as input, and may output a prior probability distribution, may receive acoustic features as input and may output a posterior probability distribution, may convert the probability distribution, may perform monotonic alignment search by using information on MIDI duration, and may output a waveform which is a voice digital signal, based on input reflecting a result of extracting the phoneme duration.
Owner:KOREA ELECTRONICS TECH INST

Enhanced earplug device with advanced acoustic filtering and smart lighting system

An enhanced earplug device for noise suppression and entertainment lighting comprises a tubular housing with a removable tip portion for ear canal insertion and a visible light element. The device incorporates a quartz acoustic filter for high-fidelity sound preservation while reducing volume, and a vibration sensor detecting musical beats in the 20-250 Hz range for music-reactive lighting patterns. Bluetooth connectivity enables inter-device synchronization and mobile app control. A tap sensor integrated into the housing exterior provides intuitive gesture-based mode switching between static color, music-reactive, and synchronized operational modes. The system includes RF capabilities for professional venue integration via centralized controllers with DMX / MIDI connectivity, enabling coordinated lighting displays across multiple users in entertainment environments.
Owner:DENNIS JOHN MAXWELL NORMAN

Transmission device, transmission method, reception device, and reception method

A novel interface is provided which supports simultaneous transmission of an audio signal and a MIDI signal. A signal continuous in predetermined units is transmitted to a receiving side via a prescribed transmission path. The signal continuous in predetermined units includes a signal of a first predetermined unit including an audio signal and a signal of a second predetermined unit including a MIDI signal. The audio signal is, for example, a linear PCM signal which constitutes a stereo 2-channel audio signal. The MIDI signal includes, for example, packet data having a prescribed length which is divided into a plurality of units, included in a plurality of signals in the second predetermined unit, and transmitted.
Owner:SONY GROUP CORP

Singing detection method and singing detection system using same

The invention provides a singing detection method and a singing detection system using the same. The singing detection method comprises the following steps: performing packet triggering judgment to start a detection function; performing rhythm detection on the sound signal; carrying out fundamental tone detection on the sound signal; carrying out quantization operation on the fundamental tone detection data; converting the quantized data and the rhythm detection data into MIDI note data; performing rhythm and pitch change judgment by using MIDI note data; weight scores are given and compared with threshold scores. And finally, judging whether the sound signal is singing according to the total score. According to the method, analysis of multiple music features is combined, singing sounds are recognized through comprehensive scores, singing sounds and non-singing sounds can be effectively distinguished without a high-performance processor or a large number of memories, and the application range of the singing detection technology in low-power-consumption and small-size equipment is greatly expanded.
Owner:GENERALPLUS TECH INC

Electronic musical instrument keyboard

The utility model provides an electronic musical instrument keyboard, including the face shell and bottom shell who butt joint into the casing, still include gyro wheel support structure, be located in the bottom shell inside, functional potentiometer, rotatable installation in gyro wheel support structure, cross section is semicircular functional gyro wheel, functional gyro wheel bottom direct surface sinks in the bottom shell inside, functional gyro wheel top stretches out the face shell upwards, functional gyro wheel center installation in functional potentiometer, gyro wheel reset spare, connect in gyro wheel support structure and functional gyro wheel between, reset buffer spare, be located in gyro wheel reset spare and gyro wheel support structure junction. The utility model provides electronic musical instrument keyboard under the diameter and rotation angle stroke limit optimization functional gyro wheel layout, realize MIDI keyboard thinning and portable, eliminate the larger rebound noise of shorter elastic arm in the limited space, satisfy individualized color selection while reducing the cost, realize assembly tolerance control, avoid partial deformation, warping in the assembly process, realize circuit board device integration module stacking and guarantee the plug assembly in place.
Owner:SHENZHEN ARTFAST TECH CO LTD

Multifunctional time code synchronization device

The utility model discloses a multifunctional time code synchronizer, which comprises an MCU (Microprogrammed Control Unit) module, a control panel connected with the MCU module, a power supply module connected with the MCU module, a plurality of interface modules and an interface circuit connected between the interface modules and the MCU module, the interface module at least comprises an MIDI IN interface and an MIDI OUT interface which can be used for transmitting an MIDI time code, an LTC IN interface and an LTC OUT interface which can be used for transmitting an SMPTE LTC time code, and an Ethernet port which can be used for transmitting an RTP MIDI time code and an ArtTimecode time code, and the MCU module can transmit and synchronize corresponding time code information through an interface circuit connected with the MIDI IN interface, the MIDI OUT interface, the LTC IN interface, the LTC OUT interface and the Ethernet port. Compared with the prior art, the device has the synchronization of the MIDI time code and the SMPTE LTC time code, and also has the network time codes of the ArtTimecode and the RTP-MIDI, so that the LTC time code can be transmitted more flexibly.
Owner:HUASHI (SHENZHEN) TECH CO LTD

Empathy-Based Wearable Systems for Perceptual and Communicative Assistance

PendingUS20260179644A1Speech recognitionHeadphonesEmpathy
The invention provides a unified assistive technology system, the ADA Empathy Wearable Bundle, enhancing perception, communication, and self-expression for individuals with visual, auditory, or speech impairments.It integrates Empathy Glasses, Empathy Headphones, Word Articulation Guidance (WAGS), Word Articulation Response Monitoring (W.A.R.M.), and Conversational MIDI (C-MIDI), using AI-driven narrative interpretation, augmented reality, and blockchain-based security.The system delivers vivid narrative audio for blind users, multimodal AR overlays for hearing-impaired users, and real-time visual and auditory articulation guidance for speech-impaired or non-verbal users. A tone-based trust protocol ensures operational integrity, while blockchain storage preserves privacy and consent.Grounded in ethical AI, empathy, and human dignity, the system enables users to perceive, understand, and engage with the world fully, offering a holistic, compassionate, and transformative solution in assistive technology.
Owner:BOWES DEAN

Symbol music generation method based on texture perception and adaptive representation alignment

The invention discloses a symbol music generation method based on texture perception and adaptive representation alignment. The method is characterized by comprising the steps of conditional decoupling and extraction of multi-scale music element features, conditional coding, texture perception FLUX music model construction, adaptive representation alignment, MIDI file generation and the like. Compared with the prior art, the method has the advantages that the structural advantage and multi-scale information of the symbol music are fully utilized, the texture consistency, chord accuracy and style controllability of the generated music are remarkably improved, the training and reasoning cost is reduced, and the generalization ability and practicability of the model are further improved. Through the three-in-one design of'conditional decoupling, texture perception diffusion and self-adaptive expression alignment ', the problems of semantic loss, inconsistent texture, uncontrollable chord and the like caused by neglecting of internal association of music in a traditional method are effectively solved, and the method is efficient, reliable and easy to implement and has good application prospects in the fields of automatic composition, interactive music creation and the like.
Owner:EAST CHINA NORMAL UNIV

MIDI controller

This is a MIDI (Musical Instrument Digital Interface) controller used for music production and performance. By using the buttons and dials on this device, you can play, record, and edit music using external sound sources, interacting with and creating music.
Owner:FOCUSRITE AUDIO ENG

AI light robot (piano version)

1. The name of the design product: AI light robot (piano version). 2. The use of the design product: can be connected to the keyboard through Bluetooth or midi for piano use, can also be played automatically according to the music score through AI, the corresponding piano keys will be lighted when playing, and can also be used in various entertainment places, birthday parties and other places to create light atmosphere and play music. 3. The design points of the design product: in shape. 4. The picture or photo that best indicates the design points: perspective view. 5. Other circumstances that need to be explained: A is transparent material.
Owner:江青

Providing MIDI controls within a virtual conferencing system

Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and method for providing Musical Instrument Digital Interface (MIDI) controls within a virtual conferencing system. The program and method provide, in association with designing a room for virtual conferencing, an interface for updating a property of an element within a room based on a received MIDI message; receiving, based on the interface, an indication of user input specifying to update the property of the element when the received MIDI message includes a predefined value; providing a virtual conference between plural participants within the room, the room including the element; receiving a MIDI message that includes the predefined value; and updating, in response to receiving the MIDI message, the property of the element.
Owner:SNAP INC

No-reference sound correction method based on music score big language model

The invention discloses a no-reference sound correction method based on a music score big language model, and relates to a mismatch song correction method in the technical field of intelligent audio processing. The objective of the invention is to solve the problems of poor pitch and rhythm collaborative correction effect, low note feature extraction precision and easy loss of emotional expression in a non-reference scene in the existing sound correction technology. The method comprises the following steps: S1, acquiring and preprocessing off-tune singing audio data; s2, carrying out note accurate segmentation on the preprocessed audio; s3, note level features are extracted, and Octuple-MIDI symbolization input is generated; s4, realizing non-reference pitch-rhythm joint error correction through a music score big language model, and outputting a standard MIDI sequence; s5, based on a standard MIDI sequence, completing cooperative accurate correction of the sound pitch and the rhythm of the song; and S6, outputting the natural corrected singing sound through a high-fidelity synthesis technology. The method is used in the field of off-tune singing processing such as karaoke entertainment, AI music creation, singing sound correction and the like.
Owner:HARBIN INST OF TECH

Multi-Host Redundant Audio Interface System with Real-Time Vocal Pitch Correction and Enhanced Signal Processing

An audio interface comprises a CPU (central processing unit) that includes a DSP (digital signal processor), a MIDI (musical instrument digital interface) processor, and two Voc FX (vocal effects) modules; MIDI IN, MIDI OUT, and MIDI HOST ports connected to the MIDI processor; first and second microphone ports connected to the first and second Voc FX modules, respectively; a plurality of OUTPUT ports connected to the DSP; and first and second HOST ports connected to the DSP; wherein the audio interface provides multi-host redundancy, real-time signal processing and MIDI control,
Owner:APOGEE ELECTRONICS CORP

Pendant (MIDI)

1. Name of the designed product: pendant (MIDI). 2. Use of the designed product: for wearing decoration. 3. Design points of the designed product: in shape. 4. Picture or photo best indicating the design points: perspective view 1.
Owner:SHENGGE (SHANGHAI) TECHNOLOGY CO LTD

Audio recognition method and device, computer device and computer readable storage medium

The application discloses an audio recognition method and device, computer equipment and a computer readable storage medium, and applies to the technical field of computers. The method comprises the following steps: after obtaining a cappella audio data, extracting a fundamental frequency sequence of the cappella audio data; obtaining a musical instrument digital interface (MIDI) template of a song corresponding to the cappella audio data, wherein the MIDI template is used for representing standardized music parameters of the song; adjusting a sound area corresponding to the fundamental frequency sequence and a sound production speed corresponding to the fundamental frequency sequence based on the MIDI template; calculating a matching degree between the adjusted fundamental frequency sequence and the MIDI template; and determining a recognition result of the cappella audio data according to the matching degree, wherein the recognition result is used for indicating whether the cappella audio data is qualified. Through the application, the accuracy of audio evaluation can be improved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

Audio rendering method, system, device and medium

The invention discloses an audio rendering method, system and device and a medium. The method comprises the following steps: receiving an audio rendering request sent by a client; obtaining timbre type information and a storage path of a digital interface file of music equipment based on the audio rendering request; generating an engine control script based on a storage path of the digital interface file of the music equipment and a target control script template matched with the tone type information, the target control script template including parameter rendering setting logic; executing the engine control script to control an audio engine to load the music equipment digital interface file based on the storage path, loading a timbre file corresponding to the timbre type information from a preset timbre library, and performing audio rendering on the music equipment digital interface file based on the timbre file, and obtaining a target audio file corresponding to the tone type information. Therefore, the audio rendering efficiency can be improved.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

A music generation method based on a large pre-training language model

The present application relates to the technical field of language processing, and more particularly to a music generation method based on a large pre-training language model. The GPT-3 is fine-tuned by using thousands of melody MIDI files, and then the fine-tuned model is used for melody generation. The main advantages of the method are as follows: (1) the algorithm can learn the long-term dependency structure of the melody and generate music with long-term structure and musicality; (2) the algorithm can simulate different melody generation methods by adjusting the fine-tuning data format; (3) the algorithm allows only a small amount of data to be used to generate similar style melodies.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Electronic device with graphical user interface for editing a polyphonic panflute panel of a stringed instrument

1. The name of the design product: the graphical user interface of the electronic device with the riffer MIDI editor panel of stringed musical instrument. 2. The use of the design product: an electronic device. 3. The design points of the design product: the graphical user interface. 4. The picture or photo that best shows the design points: the front view. 5. The use of the graphical user interface: to display the relevant information of the riffer MIDI editor panel of stringed musical instrument. 6. The human-computer interaction mode of the graphical user interface: the main view is the main interface view, and clicking the first blue "switch mode" button from left to right in the upper right corner of the main view page can enter the interface change state diagram.
Owner:BEIJING BOSHENGYUAN TECH CO LTD

Graphical user interface [computer screen layout]

This image is mainly used as an operation and display image for the music editing view of music composition software. At the bottom, the song track displays the song representation and provides song editing tools, and the groove (MIDI file and chords) is represented by blocks. The buttons and tools in the arrangement section allow the user to arrange the song by groove editing. The middle section allows a preview, and allows the user to replace the selected song track MIDI groove with the groove selected within the song track, while maintaining the song track chords. Figure 9.2 is a reference diagram showing the names of the parts. 9) Image; 9.2) Reference diagram
Owner:TOONTRACK MUSIC

A MIDI music generation method, device and terminal equipment

The application is suitable for the technical field of computers, and provides a MIDI music generation method, device and terminal equipment, which comprises the following steps: preprocessing original MIDI music data to obtain a note track and a chord track; obtaining MCST elements and chord elements based on the note track and the chord track; obtaining a first prediction model and a second prediction model based on the MCST elements and the chord elements; inputting a random MIDI music segment into the first prediction model to generate an intermediate MCST segment; inputting the generated intermediate MCST segment into the second prediction model to perform chord replacement and generate new MIDI music. The application can improve the fluency and harmony of music and also enrich the diversity of music.
Owner:HEBEI UNIV OF SCI & TECH +1