A cloud database-based music teaching man-machine interaction implementation system and method

By using a cloud database and a digital music generation engine, music teaching content is dynamically adjusted based on students' real-time music beats and body expressions, solving the problem of the lack of interactivity in existing music teaching and achieving personalized and efficient music teaching.

CN118506655BActive Publication Date: 2026-05-08JINING POLYTECHNIC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JINING POLYTECHNIC
Filing Date
2024-06-19
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing music teaching processes, AI-assisted solutions lack interactivity, fail to fully consider the user's body movements and facial expressions, and are unable to achieve personalized music teaching.

Method used

The system employs a cloud-based database for music teaching. It uses sound and image sensors to collect students' real-time music beat sequences and sequences of changes in body movements and facial expressions. The system then uses a cloud-based matching unit to match target music in the database and adjusts the teaching content through a visual interactive terminal. In addition, it uses a digital music generation engine to generate music autonomously, thus achieving personalized teaching.

Benefits of technology

It enables dynamic adjustment of teaching content based on students' individual needs, improves students' initiative and interest in learning, enhances their musical literacy, and realizes human-computer interactive teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118506655B_ABST
    Figure CN118506655B_ABST
Patent Text Reader

Abstract

The application provides a music teaching man-machine interaction implementation system and method based on a cloud database, and belongs to the technical field of man-machine interaction and artificial intelligence. The system comprises a sound sensor, an image sensor, a cloud matching unit and a visual interaction terminal. The sound sensor comprises a music collection unit and a music playing unit. The music playing unit first plays first target teaching music. Then, the music collection unit collects real-time music rhythm sequences generated by a target person at present. The image sensor synchronously captures a body action sequence and an expression change sequence of the target person. The cloud matching unit matches at least one target matching music in the cloud database based on the real-time music rhythm sequences of the target person, the body action sequence or the expression change sequence. The visual interaction terminal compares the first target teaching music with the target matching music, and adjusts the first target teaching music. The technical scheme of the application can realize personalized music teaching based on artificial intelligence assistance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human-computer interaction and artificial intelligence technology, and particularly relates to a system and method for implementing human-computer interaction in music teaching based on a cloud database, a portable terminal device, an electronic device, and a computer-readable storage medium for implementing the method. Background Technology

[0002] Traditional music teaching resources are mostly static, one-way, and paper-based, often relying on textbooks. However, textbooks typically undergo a long "experimentation period" from design and development to revision and publication. Coupled with limited content and the fact that they can only be presented in print, music teaching resources are relatively scarce and dull. Music education in the new educational landscape advocates changing the traditional, tedious "disengaged learning" process, emphasizing embodied learning where students are physically and mentally present. This goal is gradually becoming possible with the widespread adoption of artificial intelligence technology.

[0003] Artificial intelligence (AI) technology has expanded the scope of music teaching and learning beyond traditional classroom lectures and face-to-face instruction. Intelligent teaching systems can tailor learning plans to each student's progress and interests, significantly enhancing their initiative and focus, and improving their musical literacy. Furthermore, the application of AI in music analysis, sound recognition, and simulation provides students with ample opportunities for hands-on practice, deepening their understanding and appreciation of music.

[0004] Chinese patent publication CN112507294B discloses an English teaching system and method based on human-computer interaction. The patent classification number is G06F. The human-computer interaction-based English teaching system includes: a registration module, an identity verification and login module, a course selection module, a central control module, a video teaching module, an annotation module, a translation module, a voice recording module, a voice analysis module, a question-and-answer module, a storage module, and a comprehensive evaluation module. The login verification method provided by this invention enhances user login security and solves the problem of low security in existing login methods.

[0005] Chinese patent CN108897879B discloses a method for personalized teaching through human-computer interaction. Classified as G06F, this method provides a personalized teaching approach that integrates learning task decomposition, automatic question recommendation, error location, and explanation of questions via human-computer interaction. Aside from data preparation and parameter setting in the first step, no human teacher is required in the remaining steps, minimizing the use of human resources, especially educational resources. Compared to traditional explanation methods that average 15-20 minutes, this patent, by identifying error causes and providing step-by-step explanations, not only skips unnecessary content based on the user's specific learning situation, focusing on the user's questions and delving into the root causes, making it easier for users to understand the knowledge points, but also reduces the explanation time to 2-3 minutes, significantly improving the user's learning efficiency.

[0006] However, practical applications have revealed that existing AI-assisted music teaching methods still have significant room for improvement. For example, they lack sufficient interactivity with users and fail to consider other emotional factors, thus failing to truly achieve personalized music teaching. Summary of the Invention

[0007] To address the aforementioned technical problems, this invention proposes a system and method for human-computer interaction in music teaching based on a cloud database, a portable terminal device for implementing the method, an electronic device, and a computer-readable storage medium.

[0008] In a first aspect of the invention, a human-computer interaction system for music teaching based on a cloud database is proposed, the system comprising a sound sensor, an image sensor, a cloud matching unit, and a visual interactive terminal.

[0009] The sound sensor includes a music acquisition unit and a music playback unit;

[0010] The music playback unit first plays the first target teaching music;

[0011] After the first target teaching music finishes playing, the music acquisition unit acquires the real-time music beat sequence currently generated by the target character;

[0012] The image sensor simultaneously captures the sequence of body movements and facial expression changes of the target person;

[0013] The cloud matching unit matches at least one target music track in the cloud database based on the real-time music beat sequence of the target person, as well as the body movement sequence or facial expression change sequence.

[0014] The visual interactive terminal compares the first target teaching music with the target matching music, and adjusts the first target teaching music based on the comparison results.

[0015] The visual interactive terminal compares the real-time music beat sequence with the target matching music, and adjusts the first target teaching music based on the comparison results, specifically including:

[0016] The first spectral diagram of the first target teaching music and the second spectral diagram of the target matching music are displayed aligned on the human-computer interaction interface of the visual interactive terminal.

[0017] Mark the same trend spectral segments in the first and second spectral plots;

[0018] The musical sub-segment corresponding to the same trend section in the first target teaching music is used as the adjusted first target teaching music, and the adjusted first target teaching music is played to the target person to start teaching.

[0019] The cloud-based matching unit matches at least one target music track from the cloud database based on the target person's real-time music beat sequence, as well as their body movement sequence or facial expression change sequence. Specifically, this includes one of the following methods:

[0020] (1) The cloud matching unit first matches multiple candidate matching music in the cloud database based on the real-time music beat sequence, and then filters out at least one target matching music from the multiple candidate matching music based on the body movement sequence;

[0021] (2) The cloud matching unit first matches multiple candidate matching music in the cloud database based on the real-time music beat sequence, and then filters out at least one target matching music from the multiple candidate matching music based on the facial expression change sequence;

[0022] (3) The cloud matching unit matches at least one candidate matching music in the cloud database based on the real-time music beat sequence, and then uses the body movement sequence or facial expression change sequence and the candidate matching music as input to the digital music generation engine, and the digital music generation engine outputs at least one target matching music.

[0023] The cloud-based matching unit includes the digital music generation engine;

[0024] The digital music generation engine includes a generative adversarial network unit and a recurrent neural network unit;

[0025] The sequence of body movements or facial expression changes serves as the sample input for the generative adversarial network unit.

[0026] The cloud-based matching unit includes an audio extraction unit, an audio analysis unit, and a conversion unit;

[0027] The audio extraction unit extracts audio sequence data from the real-time music beat sequence;

[0028] The audio analysis unit identifies the pitch, volume, and timing in the audio sequence data;

[0029] The conversion unit converts the audio sequence data into MIDI data based on the analysis results of the audio analysis unit.

[0030] In a second aspect of the present invention, a method for implementing human-computer interaction in music teaching based on a cloud database is proposed. This method is implemented on an electronic device including a visual interactive terminal and includes the following steps:

[0031] S10: Play the music for teaching the first objective;

[0032] S20: After the first target teaching music finishes playing, collect the real-time music beat sequence generated by the target person, and simultaneously capture the target person's body movement sequence and facial expression change sequence.

[0033] S30: Extract audio sequence data from the real-time music beat sequence, identify the pitch, volume, and timing in the audio sequence data, and convert the audio sequence data into MIDI data;

[0034] S40: Based on the MIDI data, match multiple candidate matching music tracks in the cloud database, and then filter out at least one target matching music track from the multiple candidate matching music tracks based on the body movement sequence or facial expression change sequence;

[0035] S50: Compare the first target teaching music with the target matching music, adjust the first target teaching music based on the comparison result, play the adjusted first target teaching music to the target person and start teaching.

[0036] Step S50 specifically includes:

[0037] The first spectral diagram of the first target teaching music and the second spectral diagram of the target matching music are displayed aligned on the human-computer interaction interface of the visual interactive terminal.

[0038] Mark the same trend spectral segments in the first and second spectral plots;

[0039] The musical sub-segments corresponding to the same trend section in the first target teaching music are used as the adjusted first target teaching music.

[0040] In a third aspect of the present invention, a method for implementing human-computer interaction in music teaching based on a cloud database is proposed. This method is implemented using an electronic device containing a visual interactive terminal. The electronic device includes a digital music generation engine. The method comprises the following steps:

[0041] S11: Play the music for teaching the first objective;

[0042] S21: After the first target teaching music is played, collect the real-time music beat sequence generated by the target person, and simultaneously capture the target person's body movement sequence and facial expression change sequence.

[0043] S31: Extract audio sequence data from the real-time music beat sequence, identify the pitch, volume, and timing in the audio sequence data, and convert the audio sequence data into MIDI data;

[0044] S41: Based on the MIDI data, at least one candidate matching music is matched in the cloud database. Then, the body movement sequence or facial expression change sequence and the candidate matching music are used as input to the digital music generation engine, and the digital music generation engine outputs at least one target matching music.

[0045] S51: Compare the first target teaching music with the target matching music, adjust the first target teaching music based on the comparison result, play the adjusted first target teaching music to the target person and start teaching.

[0046] Step S50 specifically includes:

[0047] The first spectral diagram of the first target teaching music and the second spectral diagram of the target matching music are displayed aligned on the human-computer interaction interface of the visual interactive terminal.

[0048] Mark the same trend spectral segments in the first and second spectral plots;

[0049] The musical sub-segments corresponding to the same trend section in the first target teaching music are used as the adjusted first target teaching music.

[0050] The digital music generation engine includes a generative adversarial network unit and a recurrent neural network unit;

[0051] The sequence of body movements or facial expression changes serves as the sample input for the generative adversarial network unit.

[0052] The method described in the third or second aspect can be implemented automatically by remote server hosts, host clusters, local user terminal devices, etc., or by using electronic devices or virtual devices such as containers, runtime containers, virtual machines, physical machines, etc., as the execution subject, and automatically executing the method by executing program instructions stored in at least one storage medium.

[0053] Therefore, in a fourth aspect of the present invention, a computer-readable storage medium is also provided for storing computer instructions that, when executed on an electronic device, cause the electronic device to perform a cloud-based human-computer interaction method for music teaching as described in the third or second aspect.

[0054] In a fifth aspect of the invention, an electronic device is also provided, the electronic device comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to execute the program stored in the memory to implement the steps of the cloud database-based human-computer interaction method for music teaching described in the third or second aspect.

[0055] In specific implementation, the electronic device can be a portable terminal device, such as a PDA. The portable terminal integrates a sound sensor and an image sensor and communicates with a cloud-based artificial intelligence server to realize the human-computer interaction method for music teaching based on a cloud database as described in the third or second aspect.

[0056] The cloud-based database-based music teaching human-computer interaction system and method proposed in this invention changes the traditional mechanical, didactic, one-sided teaching method of music teaching. It fully considers the different body movements and facial expressions of each target audience (referred to as the target person, i.e., the student) to match the most suitable music fragment to start the teaching process. It can mobilize the target person's movements and emotions as much as possible and quickly integrate into the music rhythm to achieve human-computer interactive teaching. Furthermore, by introducing an artificial intelligence digital music generation engine, when suitable music cannot be matched with the existing database, the artificial intelligence engine will autonomously generate new digital music to maximize the human-computer interactive teaching matching.

[0057] Further advantages of the present invention will be further detailed in the Specific Embodiments section in conjunction with the accompanying drawings. Attached Figure Description

[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0059] Figure 1 This is a schematic diagram of the structural functional units of a cloud-based database-based human-computer interaction system for music teaching, according to an embodiment of the present invention.

[0060] Figure 2 yes Figure 1 A schematic diagram of a further preferred embodiment of the cloud-based database-based music teaching human-computer interaction system

[0061] Figure 3 yes Figure 1 A schematic diagram of an embodiment of a human-computer interaction system that further incorporates an artificial intelligence engine.

[0062] Figure 4 This is a schematic diagram comparing the trends of different spectral lines in various embodiments of the present invention.

[0063] Figure 5 This is a schematic diagram of the main flow of a cloud-based database-based human-computer interaction method for music teaching, according to an embodiment of the present invention.

[0064] Figure 6 yes Figure 5 The method further incorporates a preferred embodiment of an artificial intelligence engine. Detailed Implementation

[0065] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments identical to those described in this application. Rather, they are merely examples of apparatuses and methods identical to some aspects of this application as detailed in the appended claims.

[0066] In specific embodiments of this application, if user-related data is involved, user permission or consent must be obtained when the embodiments of this application are applied to specific products or technologies, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0067] Figure 1 This is a schematic diagram of the structural functional units of a cloud-based database-based music teaching human-computer interaction system according to an embodiment of the present invention.

[0068] exist Figure 1 The diagram shows that the cloud-based music teaching human-computer interaction system includes a sound sensor, an image sensor, a cloud matching unit, and a visual interactive terminal.

[0069] As a specific illustrative example, a sound sensor can be a sound sensing module that includes a microphone and a speaker, with the microphone serving as the music acquisition unit and the speaker serving as the music playback unit.

[0070] In one embodiment, the sound sensor may belong to a music teaching platform, that is, a fixed podium device embedded in the teaching site (classroom);

[0071] In another embodiment, the handheld portable electronic device may include the sound sensor, which includes a music acquisition unit and a music playback unit.

[0072] In the specific teaching process, the teacher first operates the music playback unit of the sound sensor to play the first target teaching music.

[0073] It is understandable that in actual teaching, when the teacher operates the music playback unit of the sound sensor, he usually makes it play the first target teaching music from the beginning, which can be the whole song or a part of the song.

[0074] After the playback is finished, the students (hereinafter referred to as the target students in this embodiment) can try singing (sing along).

[0075] At this time, after the first target teaching music has finished playing, in response to the detection that the student has started to sing the first target teaching music, the music acquisition unit is activated and begins to acquire the real-time music beat sequence currently generated by the target person.

[0076] At the same time, after the first target teaching music finishes playing, in response to detecting that a student has started to sing the first target teaching music, the image sensor is activated, and the image sensor simultaneously captures the target person's body movement sequence and facial expression change sequence.

[0077] The image sensor includes a pose detection sensor and an expression capture sensor;

[0078] The posture detection sensor is mainly used to detect the body posture and movements of a target person, preferably to capture the limb movement sequence of the target person;

[0079] The facial expression capture sensor is used to locate the facial area of ​​the target person, capture the facial expression changes in the facial area, and obtain the expression change sequence.

[0080] It is understood that the target of this application embodiment may be more than one person. When there are multiple students, the technical solution of this application can be applied to each student at the same time for teaching and interaction.

[0081] Therefore, for ease of description, the following embodiments all assume that there is only one target person, that is, interactive teaching is carried out for one student; however, it is understood that when there are multiple students, those skilled in the art can obtain the corresponding implementation scheme without creative effort.

[0082] Next, the cloud matching unit matches at least one target music track in the cloud database based on the real-time music beat sequence of the target person, as well as the body movement sequence or facial expression change sequence.

[0083] The improvement of this process includes two scenarios:

[0084] Scenario 1: At least one matching candidate music can be found in the existing music database of the cloud database.

[0085] Scenario 2: No matching music can be found in the existing music database of the cloud database.

[0086] In the following scenario, it is clear that at least one target matching music can be found from the candidate target matching music to meet the requirements;

[0087] In scenario two, since the existing music database in the cloud database does not contain any candidate music that meets the requirements, the technical solution of this application solves the problem by introducing an artificial intelligence engine to adopt an "active generation" approach.

[0088] Therefore, the cloud-based matching unit matches at least one target music track in the cloud database based on the target person's real-time music beat sequence, as well as their body movement sequence or facial expression change sequence. Specifically, this includes one of the following methods:

[0089] (1) The cloud matching unit first matches multiple candidate matching music in the cloud database based on the real-time music beat sequence, and then filters out at least one target matching music from the multiple candidate matching music based on the body movement sequence;

[0090] (2) The cloud matching unit first matches multiple candidate matching music in the cloud database based on the real-time music beat sequence, and then filters out at least one target matching music from the multiple candidate matching music based on the facial expression change sequence;

[0091] (3) The cloud matching unit matches at least one candidate matching music in the cloud database based on the real-time music beat sequence, and then uses the body movement sequence or facial expression change sequence and the candidate matching music as input to the digital music generation engine, and the digital music generation engine outputs at least one target matching music.

[0092] It is understandable that methods (1) and (2) are applicable to case one, while method (3) is applicable to case two.

[0093] The visual interactive terminal compares the first target teaching music with the target matching music, and adjusts the first target teaching music based on the comparison results.

[0094] Next, based on Figure 2 and Figure 3 Further examples of different human-computer interaction implementation systems are shown, and specific embodiments of the technical solutions under the above-mentioned scenarios one and two are further described.

[0095] Figure 2 yes Figure 1 A schematic diagram of a further preferred embodiment of the cloud-based database-based music teaching human-computer interaction system.

[0096] exist Figure 2 In the middle, the sound sensor integrates a music acquisition unit and a music playback unit; the cloud matching unit includes an audio extraction unit, an audio analysis unit, and a conversion unit.

[0097] The audio extraction unit extracts audio sequence data from the real-time music beat sequence;

[0098] The audio analysis unit identifies the pitch, volume, and timing in the audio sequence data;

[0099] The conversion unit converts the audio sequence data into MIDI data based on the analysis results of the audio analysis unit.

[0100] Based on this, the cloud matching unit matches multiple candidate matching music tracks in the cloud database based on the MIDI data, and then filters out at least one target matching music track from the multiple candidate matching music tracks based on the body movement sequence or facial expression change sequence.

[0101] Specifically, the cloud matching unit first matches multiple candidate matching music tracks in the cloud database based on the MIDI data, and then filters out at least one target matching music track from the multiple candidate matching music tracks based on the body movement sequence;

[0102] Alternatively, the cloud matching unit first matches multiple candidate matching music tracks in the cloud database based on the real-time music beat sequence, and then filters out at least one target matching music track from the multiple candidate matching music tracks based on the facial expression change sequence.

[0103] MIDI, or Musical Instrument Digital Interface, is a real-time communication protocol for electronic musical instruments, synthesizers, and other performance devices. Developed by the MIDI Manufacturers Association (MMA), it became an international standard in 1983. MIDI has become an indispensable part of music production and performance, widely used in music production software, electronic musical instruments, and stage performance. It allows different electronic instruments and computers to communicate via cables or radio waves, enabling them to work together and perform complex musical compositions. The protocol defines a series of digital information used to describe musical performance.

[0104] The MIDI protocol transmits various actions, such as key presses, playing, volume control, and pitch adjustment. This information is encoded into MIDI data and transmitted. The receiving electronic instrument or computer can decode this data and convert it into actual musical performance. The advantage of this protocol is that it provides a standardized communication method between different electronic instruments and computers, enabling different devices to work together and perform complex musical pieces. Furthermore, the MIDI protocol supports multi-channel communication, allowing multiple devices to play different tracks simultaneously, thus creating richer musical effects.

[0105] It is understood that the audio sequence data extracted by the audio extraction unit from the real-time music beat sequence is usually in other formats (e.g., .MP4 format), while the existing music files stored in the cloud database are usually in MIDI format. Therefore, the above embodiment first needs to perform conversion from other formats to MIDI format.

[0106] When a music file is transferred to a cloud database for storage, the specific implementation of this application is as follows:

[0107] Input: Converting musical actions (such as pressing keys, playing, volume, pitch, etc.) into digital information. This information can be input through devices such as electronic musical instruments, computer keyboards, mice, and touchscreens.

[0108] Encoding: Encoding the input digital information into MIDI data. MIDI data consists of a series of bytes, each byte representing a specific piece of information, such as a note, volume, or pitch.

[0109] Transmission: The encoded MIDI data is transmitted to a cloud database at the receiving end via cable or radio waves. Transmission can use a MIDI interface or wireless technologies such as Bluetooth.

[0110] When the music is to be played, the following playback decoding stage is entered:

[0111] Decoding: At the receiving end, the received MIDI data is decoded into digital information. The decoded digital information can then be converted into actual musical performance by electronic musical instruments or computers.

[0112] Output: Outputs the decoded digital information to the speakers or headphones of an electronic musical instrument or computer.

[0113] Therefore, the aforementioned method of the teacher first operating the music playback unit of the sound sensor to play the first target teaching music can also involve the teacher retrieving the corresponding music MIDI file from the cloud database and performing the above-mentioned decoding-output process.

[0114] However, it's important to note that while the MIDI format can accurately record information such as pitch, volume, and timing in music, it cannot capture details like timbre and emotional nuances. Therefore, the audio analysis unit identifies pitch, volume, and timing information in the audio sequence data, and the subsequent cloud matching unit also uses this information to match multiple candidate tracks in the cloud database. The pitch, volume, and timing information in the MIDI data of these candidate tracks is the same as or similar to that in the audio sequence data.

[0115] As can be seen, because the cloud database matching is performed using MIDI data format, only information such as pitch, volume, and timing is considered. Therefore, in most cases, multiple candidate music tracks can be initially matched.

[0116] However, while this matching is fast, it is not accurate enough. Therefore, based on the initial matching, the embodiments of this application need to be further screened.

[0117] The screening criteria include screening based on body movement sequences and screening based on facial expression change sequences.

[0118] In one embodiment, after initially identifying multiple candidate music tracks, video footage of these candidate tracks is further obtained, including music video footage, singer performance footage, solo performance footage, and chorus footage. From these video footage, the singer's body movement sequences and facial expression change sequences are identified. Then, based on a comparison of the collected body movement sequences of the target person and the singer's body movement sequences, or based on a comparison of the collected facial expression change sequences of the target person and the singer's facial expression changes, at least one target music track is filtered out from the multiple candidate tracks.

[0119] The above is the handling method for scenario one, which is the case where there is at least a candidate matching music and at least a video frame for the candidate matching music.

[0120] Next, Figure 3 yes Figure 1 A schematic diagram of an embodiment of a human-computer interaction system that further incorporates an artificial intelligence engine.

[0121] In some extreme cases, based on the real-time music beat sequence, only one candidate music may be matched in the cloud database, or no candidate music may be matched at all, or the output candidate music may not exist in the video frame and cannot be filtered later. In such cases, the above method may no longer be applicable.

[0122] The technical solution of this application first solves the problem of "being unable to match the candidate music at all".

[0123] As described above, the principle of the matching output is that "the cloud matching unit matches multiple candidate matching music in the cloud database based on the pitch, volume, timing and other information of the MIDI data of the multiple candidate matching music. The pitch, volume, timing and other information of the MIDI data of these multiple candidate matching music are the same as or similar to the pitch, volume, timing and other information in the audio sequence data", that is, similarity matching.

[0124] If the similarity matching threshold is set too high (e.g., the similarity between the two is greater than 90%), it may result in no matching of candidate music at all. Therefore, if the current similarity matching threshold causes it to be unable to match candidate music at all, the current similarity matching threshold should be lowered until at least one candidate music is output. Preferably, the similarity matching threshold should be further lowered so that it outputs multiple candidate music.

[0125] Therefore, subsequent embodiments of this application can only consider the problem that the output candidate matching music does not have video footage, thus preventing subsequent filtering.

[0126] At this point, an improved embodiment of this application introduces a digital music generation engine.

[0127] A digital music generation engine is essentially an artificial intelligence-based digital music generation tool.

[0128] In this embodiment of the application, the digital music generation engine includes a generative adversarial network unit and a recurrent neural network unit, which can autonomously generate music by combining rule-based methods and machine learning-based methods.

[0129] Specifically, firstly, based on existing music files, such as the aforementioned MIDI data format files, they are converted into note sequences and used as training data to input into the recurrent neural network training model, which then autonomously outputs melodies similar to classical music.

[0130] Then, based on the second text sequence contained in the real-time music beat sequence and the third text sequence of at least one candidate matching music matched in the cloud database based on the real-time music beat sequence, the Locality Sensitive Hashing (LSH) algorithm is used to perform music retrieval to obtain multiple candidate music text spectra;

[0131] Finally, the digital music generation engine continuously trains and optimizes the music rhythm output by the recurrent neural network training model and the candidate music syllables using the generator and discriminator included in the adversarial network unit, ultimately enabling the generator to generate images and audio similar to real data.

[0132] In this process, the body movement sequence or facial expression change sequence is used as the sample input of the generative adversarial network unit, that is, the body movement sequence or facial expression change sequence is used as the filter sample of two competing generators and discriminators.

[0133] In another embodiment, the digital music generation engine can also employ deep learning to perform content-based recommendation and collaborative filtering algorithms. For example, by using deep learning techniques such as convolutional neural networks, it can analyze music and the user's historical playback records to recommend new music to the user.

[0134] It is understandable that AI-based digital music generation is an existing technology. For example, see the following literature, which introduces the technical implementation principles and general training process of related models:

[0135] MILLER A I. 17 Project Magenta: AI Creates Its Own Music[J], 2019.

[0136] Chen Zhuanghao, et al. Intelligent evaluation algorithm for music lyrics and music matching degree based on sequence model [J]. Journal of Intelligent Systems, 2020, 15(1): 67–73.

[0137] GOODFELLOW I, et al. Generative adversarial networks[J]. Communications of the ACM, 2020, 63(11): 139–144.

[0138] LECUN Y, et al. Gradientbased learning applied to document recognition[J]. Proceedings of the IEEE, 1998, 86(11): 2278–2324.

[0139] YU

[0140] The improvement introduced in this application embodiment is that, when training and using these existing technology models, the generator and discriminator included in the adversarial network unit are continuously trained and optimized, and the samples used therein first consider the body movement sequence or facial expression change sequence as the sample input of the generative adversarial network unit, and the output is an image and audio similar to real data (i.e., the real-time music beat sequence currently generated by the target person + the target person's body movement sequence and facial expression change sequence), which can better meet the personalized needs of the current target user.

[0141] Based on this, the technical solution of this application can ensure, under any circumstances, that the cloud matching unit matches at least one target matching music in the cloud database based on the real-time music beat sequence of the target person, as well as the body movement sequence or facial expression change sequence.

[0142] Then, the visual interactive terminal compares the first target teaching music with the target matching music, adjusts the first target teaching music based on the comparison result, and plays the adjusted first target teaching music to the target person to start teaching.

[0143] Specifically, the visual interactive terminal compares the real-time music beat sequence with the target matching music, and adjusts the first target teaching music based on the comparison result, specifically including:

[0144] The first spectral diagram of the first target teaching music and the second spectral diagram of the target matching music are displayed aligned on the human-computer interaction interface of the visual interactive terminal.

[0145] Mark the same trend spectral segments in the first and second spectral plots;

[0146] The musical sub-segments corresponding to the same trend section in the first target teaching music are used as the adjusted first target teaching music.

[0147] Figure 4 This is a schematic diagram comparing the trends of different spectral lines in various embodiments of the present invention.

[0148] The red part is a schematic diagram of the first staff (beat diagram) of the current first target teaching music, and the blue part is a schematic diagram of the second staff (beat diagram) of the matched target music.

[0149] like Figure 4 The diagram shows how the first spectral graph of the first target teaching music and the second spectral graph of the target matching music are aligned and displayed on the human-computer interaction interface of the visual interactive terminal (red line and blue line).

[0150] Then, mark the same trend segments in the first and second spectral plots;

[0151] Figure 4 The dotted lines in the diagram indicate the common trend segments of the two (red FDBG and blue CAFD, and red FDBG and blue BGEC).

[0152] Finally, the musical sub-segments corresponding to the same trend segment in the first target teaching music are used as the adjusted first target teaching music.

[0153] In other words, the adjusted first target teaching music only includes the musical sub-fragments corresponding to the red FDBG spectral section. The next time it is played, the sub-fragment of that section will be played directly instead of starting from the beginning.

[0154] Understandable. Figure 4 The diagram shown is merely a schematic. Actual musical beat diagrams (spectral graphs) may take many forms, but based on graphic analysis techniques, it is possible to obtain and mark the same trend segments in the first and second spectral graphs. Alignment, trend analysis, and marking are all existing technologies.

[0155] exist Figures 1-4 Based on product implementation examples or principle descriptions Figures 5-6 Different embodiments of the technical solution of the method of this application are further illustrated.

[0156] Figure 5 This is a schematic diagram of the main process of a cloud database-based human-computer interaction method for music teaching according to an embodiment of the present invention;

[0157] Figure 6 yes Figure 5 The method further incorporates a preferred embodiment of an artificial intelligence engine.

[0158] The embodiments of the present invention can be divided into embodiments in the form of method flow, embodiments in the form of system functional architecture, and embodiments in the form of computer program code pseudo-language. It is understood that, unless otherwise specified, the same or corresponding steps between the different forms of embodiments are corresponding. For example, in the embodiment that provides a system functional architecture, "XX unit, used to implement XX function" is provided, which in fact corresponds to "XX step, used to implement XX function" in the embodiment in the form of method flow.

[0159] Therefore, after detailing the embodiments in the form of system functional architecture, the corresponding embodiments in the form of method flow do not need to be repeated. However, those skilled in the art can directly and unequivocally determine that the embodiments in the form of method flow also contain all the content of the embodiments in the form of system functional architecture. The same applies to the embodiments in the form of computer program code pseudo-language.

[0160] Figure 5 This paper presents a method for implementing human-computer interaction in music teaching based on a cloud database. The method is implemented using an electronic device that includes a visual interactive terminal.

[0161] The method for implementing human-computer interaction in music teaching based on a cloud database includes the following steps:

[0162] S10: Play the music for teaching the first objective;

[0163] S20: After the first target teaching music finishes playing, collect the real-time music beat sequence generated by the target person, and simultaneously capture the target person's body movement sequence and facial expression change sequence.

[0164] S30: Extract audio sequence data from the real-time music beat sequence, identify the pitch, volume, and timing in the audio sequence data, and convert the audio sequence data into MIDI data;

[0165] S40: Based on the MIDI data, match multiple candidate matching music tracks in the cloud database, and then based on the body movement sequence or facial expression change sequence, filter out at least one target matching music track from the multiple candidate matching music tracks;

[0166] S50: Compare the first target teaching music with the target matching music, adjust the first target teaching music based on the comparison result, play the adjusted first target teaching music to the target person and start teaching.

[0167] Step S50 specifically includes:

[0168] The first spectral diagram of the first target teaching music and the second spectral diagram of the target matching music are displayed aligned on the human-computer interaction interface of the visual interactive terminal.

[0169] Mark the same trend spectral segments in the first and second spectral plots;

[0170] The musical sub-segments corresponding to the same trend section in the first target teaching music are used as the adjusted first target teaching music.

[0171] Understandable. Figure 5 The method implementation can correspond to Figure 2 The system architecture embodiment described above.

[0172] Figure 6 This paper presents a method for implementing human-computer interaction in music teaching based on a cloud database. The method is implemented on an electronic device that includes a visual interactive terminal and a digital music generation engine.

[0173] The method for implementing human-computer interaction in music teaching based on a cloud database includes the following steps:

[0174] S11: Play the music for teaching the first objective;

[0175] S21: After the first target teaching music is played, collect the real-time music beat sequence generated by the target person, and simultaneously capture the target person's body movement sequence and facial expression change sequence.

[0176] S31: Extract audio sequence data from the real-time music beat sequence, identify the pitch, volume, and timing in the audio sequence data, and convert the audio sequence data into MIDI data;

[0177] S41: Based on the MIDI data, at least one candidate matching music is matched in the cloud database. Then, the body movement sequence or facial expression change sequence and the candidate matching music are used as input to the digital music generation engine, and the digital music generation engine outputs at least one target matching music.

[0178] S51: Compare the first target teaching music with the target matching music, and adjust the first target teaching music based on the comparison results;

[0179] Step S50 specifically includes:

[0180] The first spectral diagram of the first target teaching music and the second spectral diagram of the target matching music are displayed aligned on the human-computer interaction interface of the visual interactive terminal.

[0181] Mark the same trend spectral segments in the first and second spectral plots;

[0182] The musical sub-segments corresponding to the same trend section in the first target teaching music are used as the adjusted first target teaching music.

[0183] The digital music generation engine includes a generative adversarial network unit and a recurrent neural network unit;

[0184] The sequence of body movements or facial expression changes serves as the sample input for the generative adversarial network unit.

[0185] Understandable. Figure 6 The method implementation can correspond to Figure 3 The system architecture embodiment described above.

[0186] Furthermore, as can be seen from the foregoing description of the embodiments, Figure 5 and Figure 6 The methods described above can handle situations one and two respectively, but the technical solution of this application can also handle all situations simultaneously as a whole.

[0187] Therefore, a more complete preferred embodiment of the method of this application can be described as follows:

[0188] A method for implementing human-computer interaction in music teaching based on a cloud database, wherein the method is implemented on an electronic device including a visual interactive terminal, and the electronic device includes a digital music generation engine;

[0189] The method includes the following steps:

[0190] S1: Play the music for teaching the first objective;

[0191] S2: After the first target teaching music finishes playing, collect the real-time music beat sequence generated by the target person, and simultaneously capture the target person's body movement sequence and facial expression change sequence.

[0192] S3: Extract audio sequence data from the real-time music beat sequence, identify the pitch, volume, and timing in the audio sequence data, and convert the audio sequence data into MIDI data;

[0193] S4: Determine whether multiple candidate matching music tracks can be matched in the cloud database based on the current preset MIDI data similarity threshold;

[0194] If so, proceed to step S6;

[0195] Otherwise, proceed to step S5;

[0196] S5: Lower the similarity threshold and return to step S4;

[0197] S6: Determine whether the candidate matching music has a corresponding video frame. If so, proceed to step S7.

[0198] Otherwise, proceed to step S8;

[0199] S7: Based on the comparison between the body movement sequence or facial expression change sequence and the body movement sequence or facial expression change sequence in the video frame, filter out at least one target matching music from the multiple candidate matching music tracks, and proceed to step S9;

[0200] S8: The body movement sequence or facial expression change sequence, and the candidate matching music are used as input to the digital music generation engine. The digital music generation engine outputs at least one target matching music and proceeds to step S9.

[0201] S9: Compare the first target teaching music with the target matching music, and adjust the first target teaching music based on the comparison results.

[0202] Step S9 specifically includes:

[0203] The first spectral diagram of the first target teaching music and the second spectral diagram of the target matching music are displayed aligned on the human-computer interaction interface of the visual interactive terminal.

[0204] Mark the same trend spectral segments in the first and second spectral plots;

[0205] The musical sub-segments corresponding to the same trend section in the first target teaching music are used as the adjusted first target teaching music.

[0206] Other principles or functional units involved in each step can be referred to in the different embodiments described above.

[0207] In the above embodiments, the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device).

[0208] Therefore, further embodiments include a computer-readable storage medium for storing computer instructions that, when executed on an electronic device, cause the electronic device to perform a cloud-based database-based human-computer interaction method for music teaching described in the foregoing embodiments.

[0209] Other embodiments also propose an electronic device, which includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; the memory is used to store computer programs; and the processor is used to execute the program stored in the memory to implement the steps of the cloud database-based music teaching human-computer interaction implementation method described in the foregoing embodiments.

[0210] In a specific implementation, the electronic device may be a portable terminal device, such as a PDA. The portable terminal integrates a sound sensor and an image sensor and communicates with a cloud-based artificial intelligence server for the human-computer interaction implementation method for music teaching based on a cloud database described in the foregoing embodiments.

[0211] The human-computer interaction system and method proposed in this invention changes the mechanical, didactic, one-sided teaching method of traditional music teaching. It fully considers the different body movements and facial expressions of each target audience (referred to as the target person, i.e., the student) to match the most suitable music segment to start the teaching process. It can mobilize the target person's movements and emotions as much as possible and quickly integrate into the music rhythm to achieve human-computer interactive teaching. Furthermore, by introducing an artificial intelligence digital music generation engine, when suitable music cannot be matched with the existing database, the artificial intelligence engine will autonomously generate new digital music to maximize the human-computer interactive teaching matching.

[0212] Other functional components involved in the various embodiments, unless otherwise described in detail, are all prior art. Furthermore, the claims of this application have an open scope of protection; in addition to the essential technical features and means defined in the claims, other necessary computer control components or auxiliary components may also be included, which can be freely selected by those skilled in the art according to the actual situation.

Claims

1. A human-computer interaction system for music teaching based on a cloud database, the system comprising a sound sensor, an image sensor, a cloud matching unit, and a visual interactive terminal; Its features are: The sound sensor includes a music acquisition unit and a music playback unit; The music playback unit first plays the first target teaching music; After the first target teaching music finishes playing, the music acquisition unit acquires the real-time music beat sequence currently generated by the target character; The image sensor simultaneously captures the sequence of body movements and facial expression changes of the target person; The cloud matching unit matches at least one target music track in the cloud database based on the real-time music beat sequence of the target person, as well as the body movement sequence or facial expression change sequence. The visual interactive terminal compares the first target teaching music with the target matching music, and adjusts the first target teaching music based on the comparison results; The cloud-based matching unit matches at least one target music track from the cloud database based on the target person's real-time music beat sequence, as well as their body movement sequence or facial expression change sequence. Specifically, this includes: The cloud matching unit matches at least one candidate matching music in the cloud database based on the real-time music beat sequence, and then uses the body movement sequence or facial expression change sequence and the candidate matching music as input to the digital music generation engine, which outputs at least one target matching music. The cloud-based matching unit includes the digital music generation engine; The digital music generation engine includes a generative adversarial network unit and a recurrent neural network unit; The sequence of body movements or facial expression changes serves as the sample input for the generative adversarial network unit; The cloud-based matching unit includes an audio extraction unit, an audio analysis unit, and a conversion unit; The audio extraction unit extracts audio sequence data from the real-time music beat sequence; The audio analysis unit identifies the pitch, volume, and timing in the audio sequence data; The conversion unit converts the audio sequence data into MIDI data based on the analysis results of the audio analysis unit.

2. A method for implementing human-computer interaction in music teaching based on a cloud database, wherein the method is implemented using an electronic device, the electronic device including a digital music generation engine; Its features are, The human-computer interaction implementation method includes the following steps: S11: Play the music for teaching the first objective; S21: After the first target teaching music is played, collect the real-time music beat sequence generated by the target person, and simultaneously capture the target person's body movement sequence and facial expression change sequence. S31: Extract audio sequence data from the real-time music beat sequence, identify the pitch, volume, and timing in the audio sequence data, and convert the audio sequence data into MIDI data; S41: Based on the MIDI data, at least one candidate matching music is matched in the cloud database. Then, the body movement sequence or facial expression change sequence and the candidate matching music are used as input to the digital music generation engine, and the digital music generation engine outputs at least one target matching music. S51: Compare the first target teaching music with the target matching music, and adjust the first target teaching music based on the comparison results; The digital music generation engine includes a generative adversarial network unit and a recurrent neural network unit; The sequence of body movements or facial expression changes serves as the sample input for the generative adversarial network unit.

3. A portable terminal device, wherein the portable terminal integrates a sound sensor and an image sensor and communicates with a cloud-based artificial intelligence server, for implementing the human-computer interaction method for music teaching based on a cloud database as described in claim 2.

4. An electronic device, characterized in that, The electronic device includes: a memory and one or more processors; the memory is coupled to the processors; The memory is used to store computer program code, which includes computer instructions; when the computer instructions are executed by the processor, the electronic device performs the music teaching human-computer interaction method based on a cloud database as described in claim 2.

Citation Information

Patent Citations

  • A method for achieving personalized teaching through human-computer interaction

    CN108897879B

  • A human-computer interaction-based English teaching system and method

    CN112507294B

  • Method for realizing matching music playing by detecting motion frequency of human body

    CN105304101A

  • Tune adjusting method and device and storage medium

    CN108074557A