Guqin AI interactive teaching system and method based on multi-modal perception

By designing a guqin AI interactive teaching system that integrates computer vision, tactile sensing and acoustic analysis, the problems of complex fingering methods, tuning dependence on experience, feedback lag, and weak interaction in traditional guqin teaching are solved, and the precise capture and real-time feedback of guqin fingering methods are achieved, which improves learning efficiency and creative interest.

CN120014892APending Publication Date: 2025-05-16GUANGZHOU XIDAO CULTURE COMMUNICATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510220072.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In traditional guqin teaching, the fingering method is complex, the pitch depends on experience, the feedback is lagging, the interaction is weak, and the dynamic visual guidance and personalized learning paths are lacking.

Method used

Design an interactive guqin AI teaching system that integrates computer vision, tactile sensing and acoustic analysis. Through the collaborative design of hardware modules and intelligent software, real-time capture, tone analysis and personalized teaching feedback of guqin performance movements is realized.

Benefits of technology

It realizes accurate capture and real-time feedback of the guqin fingering method, improves learning efficiency, reduces the difficulty of teacher monitoring, provides dynamic visual guidance and personalized learning paths, and stimulates students' creative interest.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014892A_ABST
    Figure CN120014892A_ABST
Patent Text Reader

Abstract

The invention provides a guqin AI interactive teaching system and method based on multi-modal perception, and the system is characterized in that the system comprises a visual capturing unit (a side-view / top-view camera), a tactile sensing unit (a hui piezoelectric film and a string vibration sensor), and an acoustic collection unit (a microphone array and a vibration pickup); the multi-modal signal fusion engine is used for aligning visual, tactile and acoustic data by adopting a DTW (Dynamic Time Warning) algorithm; and the intelligent mirror surface display screen integrates a 3D subtraction character spectrum, a real-time thermodynamic diagram and an AI correction suggestion. A real-time feedback mechanism is combined with abbreviated character spectrum visualization and holographic projection, a traditional teaching mode is broken through, non-screen assistance is achieved through an LED lamp strip and vibration feedback, and attention distraction of students is reduced; the modular design adapts to Guqin education institutions and training schools, Guqin societies, cultural and cultural centers and Guqin fans, and the teaching cost is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of music education technology, and in particular to a guqin AI interactive teaching system that integrates computer vision, tactile sensing and acoustic analysis. Through the collaborative design of hardware modules and intelligent software, real-time capture of guqin playing movements, timbre analysis and personalized teaching feedback are achieved, solving the problems of difficult fingering and reliance on experience for intonation in traditional guqin teaching. Background Art

[0002] Traditional guqin teaching faces the following pain points: complex fingering: it is difficult to read the reduced notation, and the positioning of the frets depends on hand feel experience; delayed feedback: it is difficult for teachers to monitor students' fingering errors and force deviations in real time; weak interactivity: traditional teaching lacks dynamic visual guidance and personalized learning paths. Existing technologies such as CN113314087A (intelligent guqin system) use LEDs to prompt fingering, but lack motion capture and multimodal data fusion capabilities; CN112767557A (virtual guqin system) relies on virtual reality equipment, which is costly and detached from the real guqin operation experience. Summary of the invention

[0003] This system consists of a hardware perception layer, a data processing layer and an interactive software layer.

[0004] The hardware perception layer includes a visual capture unit, a tactile sensing unit, an acoustic collection unit, and an interactive display unit.

[0005] The side-view camera in the visual capture unit is installed on the right side of the piano body, and shoots the right-hand plucking action at a 30° angle, with a frame rate of ≥120fps, supporting tracking of key hand points (such as the angle of the fingernails touching the strings). The top-view camera in the visual capture unit (optional) shoots the piano surface from above and identifies the deviation of the left-hand fret position (accuracy ±1mm).

[0006] The tactile sensing unit includes a gesture sensor patch at the tactile position and a string vibration sensor. The gesture sensor patch at the tactile position is a strip-shaped flexible piezoelectric film (thickness ≤ 0.5mm), which is temporarily laid on the thirteenth tactile position to detect the pressing force (range 0-5N, accuracy ±0.1N); the string vibration sensor is a miniature MEMS accelerometer (size 5×5mm), which is temporarily clamped at the exposure position to monitor the vibration frequency and energy attenuation curve of the seven strings.

[0007] The acoustic collection unit includes a cardioid microphone array and a vibration pickup. The cardioid microphone array is placed 10cm below the piano body and combines beamforming technology to suppress ambient noise (SNR ≥ 70dB); the vibration pickup is embedded in the piano body resonance box to capture low-frequency resonance details (20Hz-1kHz).

[0008] The interactive display unit includes an LED light strip and a smart mirror display. LED light strip: A programmable RGB light strip is attached along the frets and the ridges to provide real-time feedback on the fingering status (such as correct fingering → green, insufficient force → yellow, misalignment → red); Smart mirror display: Integrated in front of the teaching desk, it displays 3D reduced notation, real-time fingering heat map and AI correction suggestions. In multi-person teaching mode or to save space, the smart mirror display can also be replaced by a smartphone or smart tablet placed on the piano desk.

[0009] The data processing layer includes a multimodal signal fusion engine and an AI analysis module. Multimodal signal fusion engine: visual data (hand coordinates) + tactile data (pressure) + acoustic data (pitch / timbre) → generate a comprehensive scoring matrix; use a time alignment algorithm (dynamic time warping DTW) to solve the multi-sensor delay problem. AI analysis module: Fingering recognition model: classify 21 basic fingerings (such as "pick", "hook", and "pinch") based on a convolutional neural network (CNN); timbre optimization model: compare the spectrum differences between student performances and recordings of famous artists through a generative adversarial network (GAN) to generate timbre correction suggestions.

[0010] The interactive software layer includes teaching auxiliary software functions, which have three modes: real-time feedback mode: the intelligent mirror synchronously displays the fingering trajectory and the reduced score cursor, and triggers vibration feedback when an error occurs (the piano body has a built-in linear motor); virtual accompaniment mode: AI generates a holographic projection of a famous performance, and students can follow the projection for "shadow practice"; personalized evaluation system: generates a "weakness analysis report" based on historical data (such as "the stability of the scattered sound force is insufficient, it is recommended to practice the third paragraph of "Xianweng Cao"").

[0011] The patent of this invention adopts multimodal perception fusion and pioneered the "vision-tactile-acoustic" ternary data fusion architecture to solve the problem of single sensor misjudgment (such as vision alone is easily affected by occlusion); the dynamic calibration algorithm ensures that the precision of the position positioning reaches ±0.5mm, which is better than the ±2mm of the existing patent CN112767557A. Low-cost detachable design: the sensor patch adopts magnetic adsorption or electrostatic bonding, which does not damage the guqin paint surface and is suitable for guqins of different standards; the hardware modular design (such as the camera can be folded and stored) is suitable for multiple scenarios in classrooms and homes. Intelligent teaching algorithm: based on the forgetting curve model (Ebbinghaus curve), it recommends review plans to improve learning efficiency by 30%; supports the "AI composition" function: students' improvisation fragments can automatically generate complete tracks to stimulate creative interest. For the first time, the real-time feedback logic of the fitness smart magic mirror was transferred to Guqin teaching, combining the visualization of the reduced notation and holographic projection, breaking through the traditional teaching mode, and realizing "screenless" assistance through LED light strips and vibration feedback to reduce students' distraction; the modular design is suitable for Guqin education institutions and training schools, Guqin clubs, cultural centers and Guqin enthusiasts, greatly reducing teaching costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 : System hardware layout diagram. In the figure: 1. Guqin panel; 2. Strings; 3. Qinwei; 4. Weiwei position sensor patch; 5. Yueshan; 6. String vibration sensor; 7. Vibration pickup; 8. Cardioid directional microphone array; 9. Side view camera; 10. Top view camera; 11. Smart mirror display

[0013] Figure 2 :Guqin interactive teaching flow chart based on multimodal perception Implementation

[0014] Hardware installation: Temporarily lay the fret position sensor patches along the thirteen frets of the piano surface, and clamp the string sensor on the outside of the nut; fix the side-view camera on the piano table bracket, and place the smart mirror display 1.5m in front of the student.

[0015] Software operation process: students select the song "Flowing Water", and the system loads 3D reduced notation; when playing, the LED light strip highlights the target frets in real time, and the intelligent mirror displays the left hand pressure curve; AI detects the rhythm deviation of the "rolling and brushing" technique, and the mirror pops up a slow-motion decomposition teaching video and adjusts the LED to pulse prompt mode.

[0016] Data closed-loop optimization: students’ performance data is uploaded to the cloud to generate a personal skill map; the teacher-side APP can remotely view class data and adjust the teaching plan in a targeted manner.

Claims

1. A Guqin AI interactive teaching system and method based on multimodal perception, characterized in that: Visual capture unit (side / top view camera), tactile sensing unit (microphone film + string vibration sensor), acoustic acquisition unit (microphone array + vibration pickup); multimodal signal fusion engine, using DTW algorithm to align visual, tactile, and acoustic data; Smart mirror display screen with integrated 3D score reduction, real-time heat map and AI correction suggestions.

2. The guqin AI interactive teaching system and method based on multimodal perception according to claim 1 is characterized by: The side-view camera in the visual capture unit is installed on the right side of the piano body, capturing the right-hand plucking action at a 30° angle, supporting tracking of key hand points (such as the angle at which the fingernails touch the strings). The top-view camera in the visual capture unit (optional) takes a bird's-eye view of the piano surface and identifies the deviation of the left-hand fret position (accuracy ±1mm).

3. The guqin AI interactive teaching system and method based on multimodal perception according to claim 1 is characterized by: The tactile sensing unit includes a gesture sensor patch at the tactile position and a string vibration sensor. The gesture sensor patch at the tactile position is a strip-shaped flexible piezoelectric film (thickness ≤ 0.5mm), which is temporarily laid on the thirteenth tactile position to detect the pressing force (range 0-5N, accuracy ±0.1N); the string vibration sensor is a miniature MEMS accelerometer (size 5×5mm), which is temporarily clamped at the exposure position to monitor the vibration frequency and energy attenuation curve of the seven strings.

4. The guqin AI interactive teaching system and method based on multimodal perception according to claim 1 is characterized by: The acoustic collection unit includes a cardioid microphone array and a vibration pickup. The cardioid microphone array is placed 10cm below the piano body and combines beamforming technology to suppress ambient noise (SNR ≥ 70dB); the vibration pickup is embedded in the piano body resonance box to capture low-frequency resonance details (20Hz-1kHz).

5. The guqin AI interactive teaching system and method based on multimodal perception according to claim 1 is characterized by: The interactive display unit includes an LED light strip and an intelligent mirror display. LED light strip: A programmable RGB light strip is attached along the fingering position and the yoke to provide real-time feedback on the fingering status (such as correct fingering → green, insufficient force → yellow, misalignment → red).

6. The guqin AI interactive teaching system and method based on multimodal perception according to claim 1 is characterized by: Smart mirror display: integrated in front of the teaching desk, it displays 3D reduced notation, real-time fingering heat map and AI correction suggestions. In multi-person teaching mode or to save space, the smart mirror display can also be replaced by a smartphone or smart tablet placed on the piano table.

7. The guqin AI interactive teaching system and method based on multimodal perception according to claim 1 is characterized by: The data processing layer includes a multimodal signal fusion engine and an AI analysis module. Multimodal signal fusion engine: visual data (hand coordinates) + tactile data (pressure) + acoustic data (pitch / timbre) → generate a comprehensive scoring matrix; use a time alignment algorithm (dynamic time warping DTW) to solve the multi-sensor delay problem. AI analysis module: Fingering recognition model: classify 21 basic fingerings based on a convolutional neural network (CNN); timbre optimization model: compare the spectrum differences between the student's performance and the recording of a famous artist through a generative adversarial network (GAN) to generate timbre correction suggestions.

8. The guqin AI interactive teaching system and method based on multimodal perception according to claim 1 is characterized by: The AI ​​analysis module includes a fingering recognition CNN model and a timbre optimization GAN model; and supports personalized review plan recommendations based on forgetting curves.

9. The Guqin AI interactive teaching system and method based on multimodal perception according to claim 1 is characterized by: Supports "AI Composition" function: students' improvisation fragments can automatically generate complete songs, stimulating their interest in creation.

10. The Guqin AI interactive teaching system and method based on multimodal perception according to claim 1 is characterized by: Data closed-loop optimization: students’ performance data is uploaded to the cloud to generate a personal skill map; the teacher-side APP can remotely view class data and adjust the teaching plan in a targeted manner.

Citation Information

Patent Citations

  • Design method of immersive virtual Guqin playing multi-channel interactive experience system

    CN112767557A

  • Intelligent Guqin (seven-stringed plucked instrument) system and use method

    CN113314087A