Music mutual sensing assisted beat file automatic generation method and system
Patent Information
- Application Number
- CN202610824598.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请通过提供音乐互感辅助的律动文件自动生成方法及系统,解决了现有技术中存在的缺乏有效的触觉和体感反馈机制、无法进行精准的律动感知和节奏掌握,进而导致难以有效将律动和节奏内化为肌肉记忆,使得用户参与感不足、按摩恢复效果不理想的技术问题,达到了提升用户沉浸感与互动性、增强用户参与感及理疗效果的技术效果
[0013]拟通过本申请提出的音乐互感辅助的律动文件自动生成方法及系统,获取目标用户的音乐体感跟随记录,挖掘体感关系;调取用户兴趣特征并结合体感关系,协助振动模块,构建互感决策模型;接收律动教学目标;将律动教学目标传输至互感决策模型,通过进行音乐与振动协同下的目标决策,输出目标生成文件。解决了现有技术中存在的缺乏有效的触觉和体感反馈机制、无法进行精准的律动感知和节奏掌握,进而导致难以有效将律动和节奏内化为肌肉记忆,使得用户参与感不足、按摩恢复效果不理想的技术问题,达到了提升用户沉浸感与互动性、增强用户参与感及理疗效果的技术效果。
Smart Images

Figure CN122805939A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, specifically to a method and system for automatically generating rhythm files with music interaction assistance. Background Technology
[0002] With the rapid development of smart technology and artificial intelligence, the integration of music and health has become a new trend, especially in the field of combining music rhythm and kinesthetic feedback. Music can not only regulate emotions but also, to a certain extent, regulate the body's physiological state. This is particularly beneficial for users with specific health needs, as the synergistic effect of music and vibration can help them regulate their movement and enhance their responsiveness. Currently, traditional teaching methods for kinesthetic feedback in music rhythm instruction focus primarily on visual and auditory guidance, paying less attention to individual differences in kinesthetic perception. This impacts the user's learning experience and overall training effectiveness, failing to effectively improve the user's physiological and psychological state during training. Kinesthetic technology, especially vibration feedback technology, as a new auxiliary learning tool, has been widely applied in fitness, medical rehabilitation, and other fields. By recording the user's limb movement data through kinesthetic sensors and combining it with real-time vibration feedback, learners can enhance their training effectiveness through tactile and kinesthetic perception, in addition to auditory perception.
[0003] Therefore, current technologies suffer from a lack of effective tactile and somatosensory feedback mechanisms, an inability to accurately perceive and master rhythm, and consequently, difficulty in effectively internalizing rhythm and cadence into muscle memory. This results in insufficient user engagement and unsatisfactory massage recovery effects. Summary of the Invention
[0004] This application provides a method and system for automatically generating rhythm files with music interaction assistance, which solves the technical problems existing in the prior art, such as the lack of effective tactile and somatosensory feedback mechanisms, the inability to accurately perceive rhythm and master rhythm, and the difficulty in effectively internalizing rhythm and tempo into muscle memory, resulting in insufficient user participation and unsatisfactory massage recovery effects. It achieves the technical effect of improving user immersion and interactivity, enhancing user participation and physiotherapy effects.
[0005] This application provides a method for automatically generating rhythm files with music-assisted interaction. The method includes: acquiring the music-based motion tracking records of a target user and mining motion relationships, wherein the motion relationships represent the motion tracking state of the target user in a music scene; retrieving user interest features and combining them with the motion relationships to assist a vibration module in constructing an interaction decision model, wherein the interaction decision model is trained based on the principle of adversarial networks and retains the generator construction after training; receiving a rhythm teaching objective, wherein the rhythm teaching objective is identified by a time interval; transmitting the rhythm teaching objective to the interaction decision model, and outputting a target-generated file by performing target decision-making under the coordination of music and vibration, wherein the target-generated file contains transition compensation for instantaneous rhythmic transitions.
[0006] In a possible implementation, the construction of the mutual sensing decision model further includes the following processing: based on the music haptic follow-up records, organizing and determining training samples, wherein the training samples include mapped sample teaching objectives and sample music strategies; based on the user interest features and the haptic relationship, performing alternating iterative training driven by the training samples to construct a generator-discriminator architecture; and splitting the generator of the generator-discriminator architecture to construct a first mutual sensing decision unit.
[0007] In a possible implementation, the construction of the mutual inductance decision model further includes the following processes: introducing a vibration module and obtaining the underlying control logic of the vibration module; supervising the training of a second mutual inductance decision unit, wherein the first mutual inductance decision unit is the main unit and the second mutual inductance decision unit assists, and the vibration module performs vibration control of the massage bed; integrating the first mutual inductance decision unit and the second mutual inductance decision unit to generate the mutual inductance decision model.
[0008] In a possible implementation, based on the mutual induction decision model, a target generated file is output, and the following processing is performed: identifying the rhythm teaching objective, generating a first file in conjunction with the first mutual induction decision unit, wherein the first file performs music rhythm management; based on the second mutual induction decision unit, performing user rhythm state compensation on the first file to determine a second file, wherein the second file performs vibration rhythm management; and based on a synchronization timestamp, coordinating the first file and the second file, and performing instantaneous transition processing to obtain the target generated file.
[0009] In a possible implementation, the instantaneous transition processing further includes the following steps: coordinating the first and second files to obtain an initial generated file; identifying the initial generated file and locating rhythmic transition nodes, wherein the rhythmic transition nodes are file nodes in the neighborhood time phase that are greater than a preset rhythmic variable; identifying the rhythmic transition nodes and, using the preset rhythmic variable as a constraint, performing a multi-step conversion of the rhythmic transition amount to determine the rhythmic transition time zone; and replacing the rhythmic transition nodes of the initial generated file with the rhythmic transition nodes based on the rhythmic transition time zone, using the target generated file as the target file.
[0010] In a possible implementation, after the target file is generated, the following processing is performed: controlling the music module and vibration module to respond to the target file; monitoring and determining the real-time status data of the target user based on a sensor array, wherein the sensor array is mounted on the massage bed and the sensor array has a communication connection with the mutual induction decision model; determining the rhythmic state deviation of the real-time status data, locating the file node of the target file, and adjusting the rhythmic feedback of the music and vibration.
[0011] In a possible implementation, the music-assisted rhythm file automatic generation method further performs the following processing: recording rhythm feedback information and storing it in a temporary database; based on a preset period, retrieving the periodic rhythm feedback information from the temporary database to update and learn the mutual sensing decision model.
[0012] This application also provides an automatic rhythm file generation system assisted by music interaction, comprising: a somatosensory relationship mining module, used to acquire the music somatosensory following records of the target user and mine somatosensory relationships, wherein the somatosensory relationships represent the somatosensory following rhythm state of the target user in a music scene; a somatosensory decision model construction module, used to retrieve user interest features and combine them with the somatosensory relationships to assist the vibration module in constructing a somatosensory decision model, wherein the somatosensory decision model is trained based on the principle of adversarial networks and retains the generator construction after training; a rhythm teaching target receiving module, used to receive rhythm teaching targets, wherein the rhythm teaching targets are identified by time intervals; and a target generation file output module, used to transmit the rhythm teaching targets to the somatosensory decision model, and output a target generation file by performing target decision under the coordination of music and vibration, wherein the target generation file has transition compensation for instantaneous rhythmic transitions.
[0013] This application proposes a method and system for automatically generating rhythmic files using music-assisted interaction. The method acquires the target user's music-based haptic feedback records and mines haptic relationships. It then retrieves user interest features and combines them with haptic relationships to assist the vibration module in constructing a haptic decision-making model. The system receives rhythmic teaching objectives and transmits these objectives to the haptic decision-making model. Through objective decision-making under the synergy of music and vibration, the model outputs a target-generated file. This solves the technical problems in existing technologies, such as the lack of effective tactile and haptic feedback mechanisms, the inability to accurately perceive and master rhythm, and the resulting difficulty in effectively internalizing rhythm and tempo into muscle memory. This leads to insufficient user engagement and unsatisfactory massage recovery effects. The proposed method achieves the technical effects of enhancing user immersion and interactivity, increasing user participation, and improving therapeutic efficacy. Attached Figure Description
[0014] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments of this disclosure will be briefly described below. Flowcharts are used in this application to illustrate the operations performed by the system according to the embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from these processes.
[0015] Figure 1 A schematic flowchart of the method for automatically generating rhythm files with music mutual sensing assistance provided in an embodiment of this application; Figure 2 A schematic diagram of the structure of the music-assisted rhythm file automatic generation system provided in this application embodiment.
[0016] Figure labeling: Module 10 for somatosensory relationship mining, Module 20 for mutual intuition decision-making model construction, Module 30 for rhythmic teaching objective receiving, and Module 40 for objective generation file output. Detailed Implementation
[0017] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application.
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description of this application will be provided in conjunction with the accompanying drawings. The described embodiments should not be considered as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. The terms "first" and "second" are used merely to distinguish similar objects and do not represent a specific ordering of objects. The terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only.
[0020] This application provides an embodiment of a method for automatically generating rhythm files with music-assisted resonance, such as... Figure 1 As shown, the method includes: Step S100: Obtain the music motion tracking record of the target user and mine the motion relationship, wherein the motion relationship represents the motion tracking rhythm state of the target user in the music scene.
[0021] Preferably, the music-based motion tracking data of the target user is collected based on sensor devices, visual analysis, and tactile feedback recording. The sensor devices include wearable devices (such as wristbands, smart clothing, etc.) or other motion capture devices (such as accelerometers, gyroscopes, force sensors, etc.) capable of accurately capturing the user's body posture and movement characteristics when performing rhythmic movements. Visual analysis refers to analyzing the user's limb movements through cameras or visual sensors (such as RGB cameras, depth cameras, etc.) to extract movement trajectories and posture changes. Tactile feedback recording refers to recording the user's specific tactile feedback using vibration devices, including vibration intensity, duration, and reaction speed. Specifically, motion tracking data refers to the data collected by these sensors or sensing devices during music teaching when the target user (learner) performs limb movements related to music rhythm. This data typically includes information such as the user's body movement, posture changes, movement frequency, intensity, and coordination.
[0022] Preferably, mining somatosensory relationships refers to analyzing and identifying the information contained in somatosensory following records, exploring the correlation between user body movements and musical rhythms and melodies. Specifically, somatosensory relationships refer to how user somatosensory feedback (such as movements, postures, and reactions) matches specific musical features (such as rhythm, beat, and note changes) in a musical rhythmic scenario. Through data analysis and model training, it is possible to identify the specific reaction patterns of users when performing different types of rhythmic movements, and further deduce the following information, including the degree of matching between the user's somatosensory experience and the musical rhythm (such as the coordination, speed, and intensity of the learner's limb movements when each beat or rhythmic change in the music), the difficulty of rhythmic learning and the user's adaptation (such as some users may have difficulty maintaining a stable movement frequency or precise rhythmic synchronization in high-frequency rhythms; the system can use this somatosensory data to determine the user's adaptability and learning progress in different types of musical scenarios), and personalized somatosensory patterns (the user's movement characteristics and response patterns in rhythmic following; for example, some users may rely more on visual cues when the rhythm changes, while others follow through tactile perception or auditory memory).
[0023] Preferably, the somatosensory relationship characterizes the target user's somatosensory rhythmic state in a music scene. Without external intervention, it represents the relative relationship of the user's body movements in accordance with the music. Specifically, by mining and analyzing the target user's somatosensory tracking records, a state model describing how the user rhythmically follows the music in a specific music scene is constructed. This model comprehensively reflects the user's overall state during rhythmic learning, including rhythmic precision and consistency (whether the user's movements are consistent with the rhythm and melody of the music, and whether they can consistently perform the corresponding movements at the correct time); reaction speed and sensitivity (whether the user's reaction speed in rhythmic learning is fast enough, and whether they can quickly adapt to changes in the music's rhythm); and somatosensory coordination (whether the user's rhythmic movements are coordinated and smooth, and whether different parts of the body can follow the rhythmic changes of the music). Through precise somatosensory feedback and motion analysis, it helps users more effectively perceive, master, and internalize the rhythm of the music.
[0024] Step S200: Retrieve user interest features and combine them with the somatosensory relationship to assist the vibration module in constructing a mutual sensing decision model. The mutual sensing decision model is trained based on the principle of adversarial networks, and the generator after training is retained.
[0025] Preferably, user interest characteristics refer to understanding users' personalized needs and interests in music learning by analyzing their preferences, behavioral data, and historical records. This includes analyzing historical behavioral data (such as users' historical learning records on the platform, the types of music they have participated in, and their preferred music tracks), user questionnaires or preference settings (user-provided preference settings, such as favorite music styles, rhythm intensity, etc.), and real-time behavioral analysis (such as users' immediate reactions during the learning process, whether they prefer a certain type of vibration feedback or rhythm training method). For example, users may prefer a certain type of music, a specific rhythm pattern or rhythm style, or even their performance patterns during music interaction (such as whether they prefer fast or slow music). Combining user interest characteristics with somatosensory relationships, that is, integrating user preferences and bodily response patterns, provides users with a personalized music rhythm learning experience. For example, some users may respond better to fast-paced music, while others may be more adapted to slow-paced training.
[0026] Preferably, the vibration module refers to the integration of vibration feedback technology with music teaching. Through vibration devices (such as wearable devices or smart clothing), it provides learners with tactile feedback, offering real-time bodily feedback during music rhythm instruction. This helps learners perceive rhythmic changes and coordinate body movements. For example, when learning rhythm, the system uses vibration to prompt the user when to perform a specific action or change the rhythm. Combining the user's interests and somatosensory relationships, the vibration module can be appropriately adjusted based on the user's preferences and bodily responses (such as rhythm speed and bodily reaction sensitivity), making the feedback more personalized and precise. Then, a mutual sensing decision model is constructed, which can dynamically generate personalized learning strategies and feedback schemes based on the user's somatosensory data, interests, and vibration feedback.
[0027] Preferably, the training is based on the principle of adversarial networks, retaining the trained generator to construct a mutual perception decision model. Here, a Generative Adversarial Network (GAN) is a deep learning model, typically composed of two neural networks: a generator and a discriminator. The generator aims to generate data as realistic as possible, while the discriminator aims to identify whether the generated data is authentic. During training, they compete against each other and continuously optimize to achieve the best results. Specifically, through the adversarial training process of the generator and discriminator, the mutual perception decision model is constructed and optimized. The generator is responsible for generating personalized teaching content (such as vibration feedback strategies or musical rhythm goals) based on the user's sensory feedback, interest characteristics, and other data. The discriminator evaluates whether the generated teaching content meets the user's expectations or actual effects. Through repeated training, the generator, as the mutual perception decision model, can more accurately generate more adaptive and effective teaching content based on user needs. After multiple rounds of adversarial network training… After training, the generator part of the model is optimized and saved for practical applications. This trained generator can generate music rhythm teaching content or vibration feedback strategies most suitable for the current user in real time based on the user's interest features, haptic feedback, and the synergistic effect of the vibration module, thereby improving the personalization and accuracy of teaching and ultimately achieving better music learning results. Among them, the generator (G) and discriminator (D) in the mutual induction decision model adopt a multilayer perceptron structure. The input is the user's haptic feature vector and the teaching target vector, and the output is the music-vibration synergistic strategy vector. During training, the Adam optimizer is used with a learning rate of 0.0002 and a batch size of 64. The generator and discriminator are trained alternately. In each iteration, the discriminator is trained twice and the generator is trained once, for a total of 200 rounds, until the discriminator accuracy stabilizes at around 0.5.
[0028] Furthermore, step S200 also includes step S210, which involves organizing and determining training samples based on the music-sensory follow-up records, wherein the training samples include mapped sample teaching objectives and sample music strategies; step S220, which involves performing alternating iterative training driven by the training samples based on the user interest features and the sensory relationship, thereby constructing a generation-discrimination architecture; and step S230, which involves splitting the generator of the generation-discrimination architecture to construct a first mutual sensing decision unit.
[0029] Preferably, training samples are generated by collecting and analyzing the music-based somatosensory sync records of target users. These training samples include mapped sample teaching objectives and sample music strategies, encompassing the user's somatosensory feedback under specific musical rhythms. This reflects the user's physiological or psychological reactions under different music and rhythms, such as the user's movement or rhythmic responses (e.g., whether they synchronously follow the music rhythm), the user's physiological state (e.g., heart rate, respiratory rate, especially under vibrational feedback), and the user's emotional responses (e.g., pleasure, relaxation). Specifically, the sample teaching objective refers to each sample being associated with a clear teaching objective, such as improving sleep, relaxation, or promoting recovery. The sample music strategy refers to the specific music strategy used to achieve the teaching objective, including characteristics such as rhythm, melody, frequency, and volume of the music.
[0030] Preferably, based on user interest features and somatosensory relationships, alternating iterative training driven by training samples is performed. That is, by combining these features, the alternating iterative training method (a training approach for generative adversarial networks) optimizes two main parts: the generator is responsible for generating music rhythm strategies (i.e., how to adjust the rhythm of the music, vibration feedback, etc.) based on the input (such as user interest features, somatosensory relationships, etc.); the discriminator is responsible for judging whether the generated music rhythm strategies meet the target effect and whether they can bring the user's expected feedback (such as relaxation, improved sleep, etc.). The somatosensory relationship is achieved through time-series feature extraction and regression modeling. Music features are extracted, including beat intensity, spectral centroid, and Mel frequency cepstral coefficients; somatosensory features are extracted, including acceleration amplitude, angular velocity, and root mean square of electromyography signals. Dynamic time warping (DTW) is used to align the music and somatosensory sequences. An LSTM regression model is constructed, taking music features as input and predicting haptic features as output. Error backpropagation is used for optimization, and training is performed until the mean squared error is less than 0.05. A generator-discriminator architecture is constructed, where the generator and discriminator continuously compete and optimize. Ultimately, the generator can gradually improve the quality of its output based on feedback from the discriminator, making it more in line with the target. Then, the generator of the generator-discriminator architecture is decomposed, and a first mutual sensing decision unit is constructed based on the generator. This unit is used to decide how to coordinate between music and vibration feedback. Specifically, it is responsible for translating the user's interest features and haptic relationship into specific decisions, guiding the music and vibration modules on how to cooperate, adjust feedback intensity, rhythm, and frequency under specific goals, so that the system can form a closer synergy between music and vibration to effectively achieve the target effect (such as relaxation, improved sleep, etc.) and enhance the user experience.
[0031] Furthermore, step S200 also includes step S240, introducing a vibration module and obtaining the underlying control logic of the vibration module, supervising the training of the second mutual inductance decision unit, wherein the first mutual inductance decision unit is the main one and the second mutual inductance decision unit assists, and the vibration module executes the vibration control of the massage bed; step S250, integrating the first mutual inductance decision unit and the second mutual inductance decision unit to generate the mutual inductance decision model.
[0032] Preferably, the vibration module is the key component responsible for physical feedback. It is typically connected to devices such as mattresses and cushions. The function of the vibration module is to adjust the intensity, frequency, and mode of vibration according to the needs of music rhythm teaching, thereby providing tactile feedback to the user and helping the user better perceive the rhythm of the music or achieve specific physiological and psychological goals. For example, in user rehabilitation training, the user's subjective rhythm under music rhythm, in conjunction with the assistance of the vibration module, enables the rhythmic posture to reach the standard and improves the training effect. The underlying control logic of the vibration module refers to how the vibration module is controlled through hardware and software. For example, the control logic may include vibration signal generation, frequency adjustment, intensity control, synchronization with the music rhythm, etc., which determines the specific form of vibration, whether it is synchronized with the music, and how to dynamically adjust according to user needs.
[0033] Preferably, a second mutual induction decision unit is obtained through supervised training and optimization of existing data and models. This second unit, within the framework of the first mutual induction decision unit, can accurately assist it in making more refined decisions. Specifically, the second mutual induction decision unit learns how to adjust the timing and intensity of vibration feedback according to the control logic of the vibration module, ensuring perfect coordination with the music and user needs. The first mutual induction decision unit is the primary unit, controlling the overall rhythm and tempo of the music and the coordinated operation of the vibrations. For example, if the goal is relaxation, the first mutual induction decision unit will decide what type of music to select (e.g., slow-paced music) and design a corresponding vibration mode. The second mutual induction decision unit assists in... The second mutual sensing decision unit is responsible for finely adjusting the vibration feedback. For example, when a user needs to enter a deeper state of relaxation, the second mutual sensing decision unit may further adjust the vibration frequency or intensity to enhance the relaxation effect. The vibration module is used to execute the vibration control of the massage bed. The massage bed (i.e., a mattress or cushion with vibration function) acts as the execution device and performs vibration control according to the output instructions. This mainly includes adjusting the intensity of the vibration to match the user's needs. For example, a stronger vibration may be needed for restorative goals, while a gentler vibration is needed for relaxation or improving sleep. It also adjusts different frequencies and modes. For example, a lower frequency vibration is suitable for relaxation and soothing, while a higher frequency vibration is suitable for activation or recovery.
[0034] Preferably, by integrating the first mutual induction decision unit and the second mutual induction decision unit, and combining them to form a mutual induction decision model, the model comprehensively processes information such as the user's somatosensory feedback, interest characteristics, and the control logic of the vibration module to make a final decision, guiding the generation of music and vibration feedback. Specifically, the mutual induction decision model not only includes the generation of music and vibration, but also involves real-time analysis and dynamic adjustment of user interests and somatosensory feedback, which can optimize decisions in real time, ensure perfect synchronization of music and vibration, and ultimately achieve the teaching goals set by the user (such as relaxation, improved sleep, etc.), providing users with a customized music and rhythm teaching experience.
[0035] Furthermore, step S250 also includes step S251, recording rhythm feedback information and storing it in a temporary database; step S252, based on a preset period, retrieving the periodic rhythm feedback information from the temporary database to update and learn the mutual sensing decision model.
[0036] Preferably, the rhythmic feedback information is the system's response data to the user's real-time music and vibration feedback. This may include the user's physiological data during the experience (such as heart rate, muscle response, etc.), the real-time adjustment of music and vibration (such as changes in music rhythm and vibration intensity), and whether the user has successfully reached the expected rhythmic state. This data is stored in a temporary database. Then, according to a set time period (such as minutes, hours, etc.), the stored rhythmic feedback information is periodically extracted from the temporary database to update and optimize the mutual sensing decision model. That is, by introducing new feedback data, the system can improve and optimize the decision model. For example, if the system finds that some feedback does not match the expected effect, the model can adjust the corresponding parameters (such as vibration intensity, music rhythm, etc.) to improve the future feedback effect, which helps to enhance the intelligence and adaptability of the system, making the entire experience more accurate and comfortable.
[0037] Step S300: Receive rhythmic teaching objectives, wherein the rhythmic teaching objectives are identified by a time interval.
[0038] Preferably, based on the learner's needs or the teaching plan, specific goals, namely rhythmic learning goals, are received and set. These goals are typically based on the user's personalized needs and can be determined by analyzing the user's behavioral data, health status, and emotional state. The system then generates corresponding music and rhythmic learning files based on these goals to help the user achieve the expected physiological or psychological improvements. Specifically, related to the learner's physical and mental state or specific needs, rhythmic learning goals can be aimed at improving physical and mental states. For example, through specific music and rhythmic movements and haptic feedback, users can relax their bodies, regulate their biological rhythms, and thus promote better sleep; through soothing music and rhythmic movements, users can relieve stress and anxiety, achieving a relaxing effect; and for people who have exercised or worked for long periods, specific rhythmic movements and vibrational feedback can help relax muscles and accelerate physical recovery. Each rhythmic learning objective is marked with a specific time interval. This time interval refers to the time range within which the objective is achieved, that is, the time period from the start to the end of the rhythmic learning activity. Based on the set time interval, the system can dynamically adjust parameters such as the rhythm and intensity of the music and rhythm according to the objective requirements within the specified time. For example, in the process of promoting recovery, the system can adjust the intensity of the rhythm in stages. At the beginning, it can help relax the muscles more gently, and then gradually increase the intensity to ensure that the effect of achieving the objective is maximized.
[0039] Step S400: The rhythm teaching objective is transmitted to the mutual induction decision model. By making objective decisions under the coordination of music and vibration, an objective generation file is output. The objective generation file contains transition compensation for instantaneous rhythmic transitions.
[0040] Preferably, the user's rhythmic learning goals (such as improving sleep, relaxation, and promoting recovery) are input into the interactive decision-making model. The rhythmic learning goals include specific time intervals, target effects (such as reducing anxiety and promoting sleep), and appropriate music rhythm and vibration intensity requirements. Specifically, the interactive decision-making model can make decisions on how to adjust the music and vibration feedback based on the user's needs by analyzing the user's somatosensory data and interest characteristics. For example, based on the user's response to rhythm and movement, the system can decide what type of music (such as fast or slow rhythm) and what vibration intensity (such as slight or strong vibration) to provide.
[0041] Preferably, goal-oriented decision-making is achieved through the synergy of music and vibration. Specifically, the synergy of music and vibration refers to the system's integration of vibration feedback with the musical rhythm during music and movement instruction. This enhances learners' rhythmic perception and experience through tactile feedback. This includes generating suitable musical content based on the goal (e.g., relaxation or recovery), typically incorporating elements such as harmony, rhythm, and melody. The musical characteristics change with time or goal requirements; for example, a gentle melody and slower rhythm are used for relaxation goals, while a strong rhythm is used for recovery goals to help activate the body. Based on the changes in music, corresponding tactile feedback is provided through the vibration module. The intensity, frequency, and duration of the vibration are coordinated with the rhythm of the music and the goal requirements. For instance, gentle vibrations are used for relaxation, while stronger and more regular vibrations can be used for recovery goals to help relax muscles or stimulate blood circulation. Goal-oriented decision-making achieves optimal teaching results by coordinating these elements (music and vibration).
[0042] Preferably, the final target generation file is output based on the decision results of the mutual induction decision model. This file contains music and rhythm teaching content that meets the user's needs, including a music file generated according to the user's target needs. This music file may include a combination of elements such as harmony, rhythm, and melody, which meets the target (e.g., relaxation, recovery). It also includes a vibration file, a vibration feedback file synchronized with the music file, specifying when and where to provide the user with vibration feedback of what intensity and frequency. The target generation file is used on the massage bed to guide the user's training state and help the user achieve the expected training effect according to the set rhythm teaching goals and the user's personalized needs. The target generation file includes transient rhythmic transition compensation. Transient rhythmic transitions refer to drastic or abrupt changes in rhythm or vibration patterns during music and tactile feedback, which may cause the user's perception to be unaccustomed or unable to immediately keep up with the changes. For example, a sudden switch from a slow rhythm to a fast rhythm, or a jump from weak vibration to strong vibration.
[0043] Preferably, transition compensation refers to a buffering transition mechanism added to the generated target file by the system when rhythmic transitions occur, to avoid discomfort or confusion for the user caused by such drastic changes. This typically includes rhythmic smoothing transitions, where the transition segment in the middle section is adjusted to make the rhythm change smoother. For example, when transitioning from a slower to a faster rhythm, gradually accelerating rhythm segments can be added to avoid abrupt transitions. Vibration smoothing transitions, similarly, when the vibration intensity jumps from low to high, the system can gradually increase the intensity and frequency of the vibration to make the user feel the change smoothly. The process will not feel abrupt or uncomfortable; dynamic adjustment, based on user feedback (such as monitoring the user's motion coordination or reaction speed through sensors), dynamically adjusts the smoothness of the transition, making the transition process as consistent as possible with the user's needs and adaptability; through transition compensation, the user's experience is enhanced, making them feel natural and comfortable during the rhythm teaching process, avoiding discomfort due to excessive jumps in rhythm or vibration, and ultimately achieving the user's personalized goals. For example, obtaining the rhythm intensity and transition time zone length before and after the transition node, and using sine interpolation to achieve a smooth transition, ensuring a smooth slope at the beginning and end, and avoiding abrupt changes in the body sensation.
[0044] Furthermore, step S400 also includes step S410, identifying the rhythm teaching objective, and generating a first file in conjunction with the first mutual induction decision unit, wherein the first file performs music rhythm management; step S420, based on the second mutual induction decision unit, performing user rhythm state compensation on the first file to determine a second file, wherein the second file performs vibration rhythm management; step S430, based on the synchronization timestamp, coordinating the first file and the second file, and performing instantaneous transition processing to obtain the target generated file.
[0045] Preferably, the rhythmic teaching objectives (such as relaxation, improved sleep, and promoted recovery) are identified, along with characteristics that may affect music and vibration (such as rhythm and vibration frequency). The rhythmic teaching objectives are then combined with the first mutual sensing decision unit to generate a first document. Specifically, the first mutual sensing decision unit generates a music rhythm control strategy, i.e., the first document, based on the teaching objectives and the user's somatosensory data. This determines which musical elements (such as rhythm and melody) can help the user achieve the predetermined goals, such as relaxing the user or inducing deep sleep. The first document executes music rhythm management, i.e., determines and adjusts the music content (such as selecting appropriate music, setting volume and rhythm). For example, during rehabilitation training, based on the user's subjective rhythmic state, differential rhythm compensation is performed according to the target rhythmic state, combined with the vibration module of the massage bed.
[0046] Preferably, the second mutual induction decision unit compensates for the first file based on the user's actual rhythmic state to generate a second file. This second file fine-tunes or supplements the music rhythm strategy in the first file, optimizing the coordination of music and vibration based on the user's current physical state, needs, and feedback. It identifies the user's physical feedback at specific moments (such as heart rate changes, body movements, etc.) and adjusts the rhythm, pitch, and volume of the music, or adds new music elements to adapt to the user's current state. The second file performs vibration rhythm management, controlling the vibration strategy generated by the vibration module (such as a massage bed, cushion, etc.). For example, based on the music content in the first file and the user's rhythmic state, it adjusts parameters such as the frequency, intensity, and duration of the vibration, so that the vibration and music work together to achieve the user's expected goals. The vibration module has a frequency range of 20Hz to 150Hz, a step of 1Hz, an amplitude range of 0 to 255, a linear PWM duty cycle, and waveform types that may be sine waves, square waves, or triangle waves. It is synchronized based on MIDI clock signals or timestamps with an accuracy of 10ms.
[0047] Preferably, a synchronization timestamp is used to coordinate the execution of the first and second files, ensuring that the first file (music rhythm management) and the second file (vibration rhythm management) work together at the same time and guaranteeing a consistent and comfortable user experience. The synchronization timestamp ensures the synchronization of music and vibration feedback; for example, a specific music beat may require a specific vibration frequency or intensity. The timestamp precisely marks the time of each frame and synchronizes the occurrence of music and vibration. Transition refers to the process of smoothly transitioning from one state to another. Instantaneous transition processing refers to smoothly transitioning to a new state when the rhythm or intensity of music or vibration changes, avoiding abrupt changes that cause discomfort and making the user experience smoother. For example, the music rhythm may change from fast to slow, and the vibration intensity may change from strong to weak; transition processing ensures a comfortable experience. Finally, a target file is generated, containing the final control strategy for music and vibration rhythms, and can be executed within a specific time period, ensuring a consistent and effective user experience throughout the process.
[0048] Furthermore, step S430 also includes step S431, coordinating the first file and the second file to obtain an initial generated file; step S432, identifying the initial generated file and locating rhythmic transition nodes, wherein the rhythmic transition nodes are file nodes in the neighborhood time phase that are greater than a preset rhythmic variable; step S433, identifying the rhythmic transition nodes, and using the preset rhythmic variable as a constraint, performing multi-step conversion of the rhythmic transition amount to determine the rhythmic transition time zone; step S434, based on the rhythmic transition time zone, replacing the rhythmic transition nodes of the initial generated file as the target generated file.
[0049] Preferably, the first file (music rhythm management file) and the second file (vibration rhythm management file) are coordinated to ensure the synchronization and coordination of music and vibration, forming an initial generated file containing the execution file of the preliminary music and vibration strategy. In the initial generated file, rhythm transition nodes are further identified and marked. Rhythm transition nodes refer to file nodes in the neighborhood with a rhythm transition node greater than the preset rhythm variable. That is, at certain moments when the rhythm of the music (such as rhythm, pitch, and speed) or the intensity and frequency of the vibration in the file changes significantly, the system identifies them as transition nodes. For example, during the transition from relaxation music to recovery music, the music rhythm may suddenly accelerate or the vibration frequency may change significantly, which is called a transition. For example, at the initial, termination, or intermediate moments, the instantaneous variables of music and vibration are large, and the music sound or rhythm suddenly becomes faster or the massage intensity increases. The preset rhythm variable is the standard value for judging the acceleration of the music rhythm.
[0050] Preferably, rhythmic transition nodes are identified, and the rhythmic transition amounts are processed. The transition amount refers to the magnitude of change from one state to another, such as an increase in rhythm or vibration intensity. Multi-step transitions of the rhythmic transition amount refer to a smooth transition process; that is, instead of directly jumping to a new rhythm or vibration intensity, adjustments are made gradually through multiple steps to achieve a smooth transition, avoiding sudden changes and ensuring consistency and comfort in the user experience. This determines the transition time interval, i.e., the rhythmic transition time zone, which refers to a time period in the file during which the rhythm of the music and vibration undergoes significant changes (i.e., transitions), such as a transition from a slow rhythm to a fast rhythm, or from low-intensity vibration to high-intensity vibration. Finally, the rhythmic transition nodes in the initially generated file are replaced. Based on the transition time zone and adjustment rules, the position and properties of the transition nodes are redefined and adjusted to obtain the target generated file. This allows for precise control of the rhythm, intensity, and other factors of the music and vibration to achieve the user's predetermined goals, providing a more comfortable and customized experience.
[0051] Furthermore, step S400 also includes step S440, controlling the music module and vibration module to respond to the target generated file; step S450, monitoring and determining the real-time status data of the target user based on the sensor array, wherein the sensor array is mounted on the massage bed and the sensor array has a communication connection with the mutual induction decision model; step S460, determining the rhythm state deviation of the real-time status data, locating the file node of the target generated file, and adjusting the rhythm feedback of the music and vibration.
[0052] Preferably, the music and vibration modules are controlled in response to a target-generated file. The music module controls the playback content, rhythm, volume, and other parameters of the music based on the target-generated file, while the vibration module controls the frequency, intensity, and duration of the vibration based on the target-generated file, ensuring coordination and consistency between the music and vibration. A sensor array is used to monitor the real-time physical state of the target user. This sensor array is typically installed on devices such as massage beds to capture and record the user's real-time physiological and somatosensory data, such as heart rate, body movement, muscle tension, body temperature, and information about the user's posture (sitting / lying position). The sensor array is connected to a mutual induction decision model, transmitting this monitoring data to the model in real-time. The mutual induction decision model combines sensor feedback and other input data (such as the user's interest characteristics and somatosensory relationships) to adjust the music and vibration strategies in real-time. Specifically, the massage bed embeds an 8×8 piezoelectric sensor array with a sampling frequency of 100Hz, integrating a triaxial accelerometer (±8g) and a photoplethysmography (PPG) heart rate sensor. Real-time status data is denoised using Kalman filtering and its deviation from the target rhythm state is determined. The deviation threshold is a rhythm deviation of >10% or an amplitude deviation of >15%. After the deviation is triggered, the corresponding timestamp node in the target generated file is located, and the mutual induction decision model is called to regenerate the rhythm strategy.
[0053] Preferably, the system analyzes real-time collected user data to determine rhythmic state deviation. Specifically, by comparing the difference between real-time state data and the expected music and vibration feedback, it determines whether the user needs adjustment. For example, the user may not be able to keep up with the rhythm of the music or get the expected vibration effect due to some reasons (such as fatigue, stress, improper posture, etc.). Rhythmic state deviation refers to the deviation between the user's current state and the system's expected state, which may be reflected in the user's physiological reaction (such as high heart rate) or physical feedback (such as physical discomfort). If a rhythmic state deviation is determined, the system locates the file nodes in the target generated file. The file nodes refer to the key time or event nodes in the target generated file. Each node corresponds to a specific music rhythm and vibration feedback, and adjustments are made according to the deviation, such as changing the music rhythm, volume, vibration intensity, etc., to help the user achieve the rhythmic teaching goal and ensure that the user can experience the best music and vibration feedback at every moment throughout the process.
[0054] In the above text, refer to Figure 1 This paper describes in detail a method for automatically generating rhythm files with music mutual sensing assistance according to an embodiment of the present invention. Next, we will refer to... Figure 2 This invention describes an automatic rhythm file generation system assisted by music mutual sensing according to an embodiment of the present invention.
[0055] The music-sensory interaction-assisted rhythmic file automatic generation system according to embodiments of the present invention addresses the technical problems existing in the prior art, such as the lack of effective tactile and somatosensory feedback mechanisms, the inability to accurately perceive rhythm and master it, and consequently the difficulty in effectively internalizing rhythm and tempo into muscle memory, resulting in insufficient user participation and unsatisfactory massage recovery effects. The system achieves the technical effects of enhancing user immersion and interactivity, increasing user participation, and improving therapeutic efficacy. The music-sensory interaction-assisted rhythmic file automatic generation system includes: a somatosensory relationship mining module 10, a sensory interaction decision model construction module 20, a rhythmic teaching goal receiving module 30, and a goal generation file output module 40.
[0056] The haptic relationship mining module 10 is used to acquire the target user's music haptic follow-up records and mine haptic relationships, wherein the haptic relationships represent the target user's haptic follow-up rhythm state in a music scene; the mutual sensing decision model construction module 20 is used to retrieve user interest features and combine them with the haptic relationships to assist the vibration module in constructing a mutual sensing decision model, wherein the mutual sensing decision model is trained based on the principle of adversarial networks and retains the generator construction after training; the rhythm teaching target receiving module 30 is used to receive rhythm teaching targets, wherein the rhythm teaching targets are identified by time intervals; the target generation file output module 40 is used to transmit the rhythm teaching targets to the mutual sensing decision model, and output a target generation file by performing target decision under the coordination of music and vibration, wherein the target generation file contains transition compensation for instantaneous rhythmic transitions.
[0057] The specific configuration of the mutual sensing decision-making model construction module 20 will be described in detail below. The mutual sensing decision-making model construction module 20 further includes: organizing and determining training samples based on the music-based haptic feedback recordings, wherein the training samples contain mapped sample teaching objectives and sample music strategies; performing alternating iterative training driven by the training samples based on the user interest features and the haptic feedback relationship to construct a generator-discriminator architecture; and splitting the generator of the generator-discriminator architecture to construct a first mutual sensing decision-making unit.
[0058] The specific configuration of the mutual inductance decision model construction module 20 will be described in detail below. The mutual inductance decision model construction module 20 may further include: introducing a vibration module and obtaining the underlying control logic of the vibration module; supervising the training of a second mutual inductance decision unit, wherein the first mutual inductance decision unit is the primary unit and the second mutual inductance decision unit assists, and the vibration module executes the vibration control of the massage bed; integrating the first mutual inductance decision unit and the second mutual inductance decision unit to generate the mutual inductance decision model.
[0059] The specific configuration of the target file generation output module 40 will be described in detail below. The target file generation output module 40 further includes: identifying the rhythm teaching objective; generating a first file in conjunction with the first mutual induction decision unit, wherein the first file performs music rhythm management; performing user rhythm state compensation on the first file based on the second mutual induction decision unit to determine a second file, wherein the second file performs vibration rhythm management; and obtaining the target generated file by coordinating the first and second files based on a synchronization timestamp and performing instantaneous transition processing.
[0060] The specific configuration of the target file generation output module 40 will be described in detail below. The target file generation output module 40 further includes: coordinating the first file and the second file to obtain an initial generated file; identifying the initial generated file and locating rhythmic transition nodes, wherein the rhythmic transition nodes are file nodes in the neighborhood time phase that are greater than a preset rhythmic variable; identifying the rhythmic transition nodes, and using the preset rhythmic variable as a constraint, performing a multi-step transformation of the rhythmic transition amount to determine the rhythmic transition time zone; and replacing the rhythmic transition nodes of the initial generated file with the rhythmic transition nodes based on the rhythmic transition time zone, using the target generated file as the target file.
[0061] The specific configuration of the target file generation output module 40 will be described in detail below. The target file generation output module 40 may further include: controlling a music module and a vibration module to respond to the target file; monitoring and determining the real-time status data of the target user based on a sensor array, wherein the sensor array is mounted on the massage bed and has a communication connection with a mutual induction decision model; determining the rhythmic state deviation of the real-time status data, locating the file node of the target file, and adjusting the rhythmic feedback of the music and vibration.
[0062] The specific configuration of the music-sensing-assisted rhythm file automatic generation method will be described in detail below. The music-sensing-assisted rhythm file automatic generation method further includes: recording rhythm feedback information and storing it in a temporary database; and, based on a preset period, retrieving the periodic rhythm feedback information from the temporary database to update and learn the mutual sensing decision model.
[0063] The music-assisted rhythm file automatic generation system provided in this embodiment of the invention can execute the music-assisted rhythm file automatic generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.
[0064] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.
[0065] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for automatically generating rhythm files with music-mediated interaction, characterized in that, The method includes: Acquire the music motion tracking records of the target user and mine the motion relationship, wherein the motion relationship represents the motion tracking rhythm state of the target user in the music scene; By retrieving user interest features and combining them with the somatosensory relationship, the vibration module is assisted in constructing a mutual sensing decision model. The mutual sensing decision model is trained based on the principle of adversarial networks, and the generator after training is retained. Receive rhythmic learning objectives, wherein the rhythmic learning objectives are identified by a time interval; The rhythmic teaching objectives are transmitted to the mutual induction decision model. By making objective decisions under the coordination of music and vibration, an objective generation file is output, wherein the objective generation file has transition compensation for instantaneous rhythmic transitions.
2. The method for automatically generating rhythm files with music mutual sensing assistance as described in claim 1, characterized in that, The construction of the mutual intuition decision-making model includes: Based on the music-sensory follow-up records, training samples are organized and determined, wherein the training samples include mapped sample teaching objectives and sample music strategies; Based on the user interest features and the somatosensory relationship, alternating iterative training driven by the training samples is performed to construct a generation-discrimination architecture; The generator of the aforementioned generator-discriminator architecture is split to construct the first mutual intuition decision unit.
3. The method for automatically generating rhythm files with music mutual sensing assistance as described in claim 2, characterized in that, The construction of the mutual intuition decision-making model includes: A vibration module is introduced, and the underlying control logic of the vibration module is obtained. The second mutual inductance decision unit is trained under supervision. The first mutual inductance decision unit is the main unit, and the second mutual inductance decision unit assists. The vibration module executes the vibration control of the massage bed. The mutual inductance decision unit and the second mutual inductance decision unit are integrated to generate the mutual inductance decision model.
4. The method for automatically generating rhythm files with music mutual sensing assistance as described in claim 3, characterized in that, Based on the mutual intuition decision-making model, the target generation file is output, including: Identify the rhythm teaching objectives, and generate a first file in conjunction with the first mutual induction decision unit, wherein the first file performs music rhythm management; Based on the second mutual inductance decision unit, the user rhythm state compensation is performed on the first file to determine the second file, and the second file performs vibration rhythm management. Based on the synchronization timestamp, the first file and the second file are coordinated, and an instantaneous transition process is performed to obtain the target generated file.
5. The method for automatically generating rhythm files with music mutual sensing assistance as described in claim 4, characterized in that, The instantaneous transition processing includes: Combine the first file and the second file to obtain the initial generated file; Identify the initially generated file and locate the rhythm transition node, wherein the rhythm transition node is a file node with a preset rhythm variable in the neighborhood temporal phase; Identify the rhythmic transition nodes, and use the preset rhythmic variables as constraints to perform multi-step conversion of the rhythmic transition amount to determine the rhythmic transition time zone; Based on the rhythmic transition time zone, the rhythmic transition nodes of the initially generated file are replaced as the target generated file.
6. The method for automatically generating rhythm files with music mutual sensing assistance as described in claim 1, characterized in that, After the output target file is generated, it also includes: Control the music module and vibration module to generate a file in response to the target; Based on a sensor array, the real-time status data of the target user is monitored and determined. The sensor array is mounted on the massage bed and has a communication connection with the mutual induction decision model. The real-time status data is used to determine the rhythm state deviation, locate the file node of the target generated file, and adjust the rhythm feedback of music and vibration.
7. The method for automatically generating rhythm files with music mutual sensing assistance as described in claim 1, characterized in that, The method further includes: Record rhythmic feedback information and store it in a temporary database; Based on a preset period, the rhythmic feedback information is retrieved periodically from the temporary database to update and learn the mutual sensing decision model.
8. A music-sensing-assisted rhythm file automatic generation system, characterized in that, The system is used to implement the automatic generation method for rhythm files with music mutual sensing assistance as described in any one of claims 1 to 7, the system comprising: The somatosensory relationship mining module is used to acquire the music somatosensory follow-up records of the target user and mine somatosensory relationships, wherein the somatosensory relationship represents the somatosensory follow-up rhythm state of the target user in the music scene; The mutual sensing decision model construction module is used to retrieve user interest features and combine them with the somatosensory relationship to assist the vibration module in constructing a mutual sensing decision model. The mutual sensing decision model is trained based on the principle of adversarial network and retains the generator after training. The rhythmic teaching objective receiving module is used to receive rhythmic teaching objectives, wherein the rhythmic teaching objectives are identified by a time interval; The target generation file output module is used to transmit the rhythm teaching target to the mutual induction decision model, and output the target generation file by making target decisions under the coordination of music and vibration. The target generation file has transition compensation for instantaneous rhythm transitions.