Multi-mode sound therapy method and electronic equipment thereof
By obtaining and processing multi-source physiological parameters, generating personalized physiological therapy audio, and driving the sound therapy intervention module, the problem that existing sound therapy methods cannot integrate multi-source physiological parameters and lack of AI personalized modeling is solved, and the intelligent and individualized effects of sound therapy are achieved.
Patent Information
- Application Number
- CN202510452339.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing multimodal sound therapy methods cannot integrate multi-source physiological parameters, lack the ability to personalize AI modeling, and cannot realize linkage intervention and closed-loop adjustment under audio-driven.
By obtaining the current physiological parameters of the target object, including heart sound data, blood pressure data and pulse data, the preset sound therapy processing model is used for processing, personalized physiotherapy audio data is generated, and the driver sound therapy intervention module is matched for sound therapy processing.
It realizes adaptive control between multimodal physiological parameters and personalized sound therapy, collects physiological data in real time, generates personalized audio, and drives the sound therapy intervention module, improving the intelligence and individualization of sound therapy.
Smart Images

Figure CN120204570A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of music therapy, and in particular to a multi-modal music therapy method, device, system, electronic device and its storage medium. Background Art
[0002] Music Therapy, as a non-drug intervention means, is widely used in various health management scenarios such as psychological regulation, blood pressure control, and sleep assistance. With the development of artificial intelligence and biosensing technologies, it is becoming increasingly popular to use AI models to judge users' physiological data.
[0003] However, traditional music therapy methods mainly rely on the playback of general audio content, without dynamically adapting to the individual's current physiological state. Moreover, most of them only use the audio playback method for one-way auditory stimulation, lacking associated physical intervention means and real-time perception and feedback of physiological changes after intervention, and unable to form an adaptive regulation mechanism. Therefore, the existing music therapy methods have problems such as being unable to integrate multi-source physiological parameters, utilize the AI personalized modeling ability, and achieve associated intervention and closed-loop regulation driven by audio. Summary of the Invention
[0004] Embodiments of the present invention provide a multi-modal music therapy method to solve the problems that the existing multi-modal music therapy methods are unable to integrate multi-source physiological parameters, utilize the AI personalized modeling ability, and achieve associated intervention and closed-loop regulation driven by audio.
[0005] In a first aspect, embodiments of the present invention provide a multi-modal music therapy method, and the method includes the following steps: Obtain the current physiological parameters of a target object, where the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data; Process the at least one of heart sound data, blood pressure data, and pulse data through a preset music therapy processing model to generate personalized physiotherapy audio data; Based on the personalized physiotherapy audio data, match and drive a music therapy intervention module to perform music therapy on the target object.
[0006] Optionally, the obtaining of the current physiological parameters of the target object, where the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data, includes: Detect the current physiological parameters of the target object through a preset multi-modal sensor to obtain first heart sound data, first blood pressure data, and first pulse data; Based on a preset detection interval, perform time series synchronization processing on the first heart sound data, first blood pressure data, and first pulse data to obtain second heart sound data, second blood pressure data, and second pulse data; Use the second heart sound data, second blood pressure data, and second pulse data as the physiological parameters of the current user.
[0007] Optionally, the method for performing time series synchronization processing on the first heart sound data, first blood pressure data, and first pulse data based on a preset detection interval to obtain second heart sound data, second blood pressure data, and second pulse data includes: Determine the beating interval time between two adjacent heart sound data and the pressure difference waveform data between two adjacent pulse data; Based on the beating interval time and the pressure difference waveform data, perform fusion processing on the first heart sound data, first blood pressure data, and first pulse data to obtain second heart sound data, second blood pressure data, and second pulse data.
[0008] Optionally, the method for processing at least one of the heart sound data, blood pressure data, and pulse data through a preset sound therapy processing model to generate personalized physiotherapy audio data includes: Extract the features of at least one of the heart sound data, blood pressure data, and pulse data to obtain the corresponding control feature vectors of at least one of the heart sound data, blood pressure data, and pulse data; Based on the corresponding control feature vectors of at least one, perform fusion processing to determine the corresponding control audio feature data; According to the control audio feature data, synthesize personalized physiotherapy audio data.
[0009] Optionally, for the method of performing fusion processing based on the corresponding control feature vectors of at least one to determine the corresponding control audio feature data, the method further includes: Extract the time series data of the corresponding control feature vectors of at least one; Perform time series alignment processing on the corresponding time series data, and fuse the corresponding control feature vectors of at least one to obtain fused feature data; Based on the fused feature data, generate control audio feature data including frequency, rhythm, and amplitude control parameters; During the generation process, collect the changes in the physiological parameters of the target object, and adjust the control audio feature data based on the collection results.
[0010] Optionally, for the method of matching and driving a sound therapy intervention module to perform sound therapy on the target object based on the personalized physiotherapy audio data, it includes: Real-time determine the control audio feature data in the current personalized physiotherapy audio data; Based on the controlled audio feature data, drive the corresponding audio therapy intervention module, so that the audio therapy intervention module performs rhythmic physical therapy intervention according to the controlled audio feature data, and the audio therapy intervention module is used to perform physical therapy intervention on the user in cooperation with personalized physical therapy audio data.
[0011] In a second aspect, an embodiment of the present invention further provides a multimodal audio therapy device, which includes: A first acquisition module, configured to acquire the current physiological parameters of a target object, where the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data; A first generation module, configured to process the at least one of heart sound data, blood pressure data, and pulse data through a preset audio therapy processing model to generate personalized physical therapy audio data; A first matching module, configured to match and drive an audio therapy intervention module to perform audio therapy on the target object based on the personalized physical therapy audio data.
[0012] In a third aspect, an embodiment of the present invention provides a multimodal audio therapy system, which includes: a multimodal audio therapy system device, a server, and an audio therapy intervention module.
[0013] In a fourth aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps in the multimodal audio therapy method provided by the embodiment of the present invention are implemented.
[0014] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the multimodal audio therapy method provided by the embodiment of the invention are implemented.
[0015] In the embodiment of the present invention, the current physiological parameters of a target object are acquired, where the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data; the at least one of heart sound data, blood pressure data, and pulse data is processed through a preset audio therapy processing model to generate personalized physical therapy audio data; based on the personalized physical therapy audio data, an audio therapy intervention module is matched and driven to perform audio therapy on the target object. Through the adaptive control between multimodal physiological parameters and personalized audio therapy, data such as the heart sound, blood pressure, and pulse of the target object are collected in real time, the control features are extracted by using a preset audio therapy processing model and personalized physical therapy audio is generated, and then the audio therapy intervention module is driven to realize the collaborative treatment of auditory stimulation and physical intervention, improving the intelligent and individualized level of non-drug intervention. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 is the system architecture diagram of a multi-modal sound therapy system provided by an embodiment of the present invention; Figure 2 is the flowchart of a multi-modal sound therapy method provided by an embodiment of the present invention; Figure 3 is the structural schematic diagram of another multi-modal sound therapy device provided in an embodiment of the present invention; Figure 4 is the structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0019] As Figure 1 shown, Figure 1 is the architecture diagram of a multi-modal sound therapy system 100 provided by an embodiment of the present invention. The multi-modal sound therapy system includes: a multi-modal sound therapy device 300, a server 101, and a sound therapy intervention module 102. Among them, the above multi-modal sound therapy device 300 further includes: a first acquisition module, which can be used to acquire the current physiological parameters of the target object, and the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data; a first generation module, which can be used to process at least one of heart sound data, blood pressure data, and pulse data through a preset sound therapy processing model to generate personalized physiotherapy audio data; a first matching module, which can be used to match and drive the sound therapy intervention module to perform sound therapy on the target object based on the personalized physiotherapy audio data.
[0020] Specifically, the above target object may refer to patients or sub-healthy people who need cardiovascular intervention, blood pressure management, stress relief, nerve regulation, or rehabilitation training. It can be understood that the target object can be in a resting, training, or clinical monitoring state, and the changes in its physiological parameters can be used as the basis for the response and adjustment of the multi-modal sound therapy system.
[0021] The above-mentioned current physiological parameters can refer to the key physiological information obtained in real time during the initiation or operation of audio therapy for the target object, which can be used to reflect the original physiological response parameters such as its current cardiovascular state or autonomic nervous system activity. Among them, it can include but is not limited to heart sound data, blood pressure data, and pulse data, and can also include other auxiliary physiological signals such as corresponding electroencephalogram, galvanic skin response, and respiratory rate.
[0022] More specifically, the above-mentioned heart sound data can refer to the heart sound signal of the target object collected by sensors (such as MEMS microphones, piezoelectric pickups), which is used to represent the mechanical vibration generated when the heart valve closes. It should be noted that after filtering, envelope extraction, rhythm analysis and other links, characteristic data such as S1 / S2 interval time, peak frequency, and waveform morphology can be obtained. This characteristic data can be used to evaluate heart rate, rhythm stability, and cardiac pumping function, providing a heart sound parameter basis for subsequent audio therapy intervention.
[0023] The above-mentioned blood pressure data can refer to the systolic blood pressure, diastolic blood pressure and their dynamic change information obtained by non-invasive methods (such as oscillometric method with a cuff). The above-mentioned blood pressure data can be used to judge vascular tone, sympathetic nerve activity, and overall circulatory load level, and serve as the blood pressure parameter basis for generating audio therapy control strategies.
[0024] The above-mentioned pulse data can refer to the pulse waveform signal obtained by photoplethysmography (PPG) or other non-invasive detection means. The above-mentioned pulse data can include but is not limited to parameters such as pulse frequency, waveform morphology, and pulse wave transit time (PWTT), and can be used to estimate blood pressure change trends, heart rate variability, and vascular compliance, serving as the pulse parameter basis for reflecting the state of peripheral circulation and autonomic nervous system.
[0025] In a possible embodiment, the above-mentioned multi-modal audio therapy system collects multiple physiological state data of the user in real time through a variety of sensor devices configured at specific parts of the user's body, and uses it as the parameter basis for subsequent audio intervention.
[0026] The above-mentioned preset audio therapy processing model can be a deep learning model or a multi-modal feature processing network that is pre-trained and deployed in the above-mentioned multi-modal audio therapy system. It is used to receive the current physiological parameters of the target object and generate audio control features that match its state. It can be understood that the above-mentioned preset audio therapy processing model can adopt a convolutional neural network (CNN), a Transformer structure, an attention fusion mechanism, or a combination thereof to achieve feature extraction, fusion calculation, and audio control mapping of multi-modal signals. And this model can be deployed on the AI algorithm processing unit to dynamically generate matching physiotherapy audio through blood pressure data. For example, when the systolic blood pressure is high, a low-frequency audio (such as 20-50 Hz) is generated to promote vasodilation. When the pulse is abnormal, a random frequency pulse waveform is used to enhance neuromodulation, and it supports the Bluetooth protocol to realize remote algorithm update.
[0027] The processing process through the above-mentioned preset audio therapy processing model can be a series of calculation steps such as data parsing, feature extraction, temporal modeling, feature fusion, and audio parameter generation performed after inputting the collected physiological parameters into the preset audio therapy processing model. For example, it can be the process of aligning the first type of data using a unified clock source or a software timestamp mechanism, including but not limited to: 1. Time window division: setting a reference time period based on heart sound data; 2. Interpolation / truncation / alignment: resampling or window alignment of the blood pressure band and the pulse band to ensure that the data can be jointly analyzed at the same moment; 3. Abnormal processing mechanism: If the data of a certain sensor is delayed beyond the allowable range (such as the blood pressure does not return a valid value), then the data of this cycle is invalidated or marked as invalid.
[0028] The above-mentioned personalized physiotherapy audio data can refer to audio control information that can have a positive intervention effect on the user's current physiological state and is output by the preset audio therapy processing model based on the current physiological state of the target object. Specifically, it can be generated through the intelligent selection mode of audio segments and the dynamic synthesis mode of audio segments. Among them, the above-mentioned intelligent selection mode of audio segments can be a generation mode in which the above-mentioned audio therapy audio data generation platform selects the most matching audio content from the preset audio therapy audio material library according to indicators such as rhythm, frequency band, and style output by the model and pushes it for playback. The above-mentioned dynamic synthesis of audio segments can be a process in which the above-mentioned audio therapy audio data generation platform uses an audio synthesis algorithm (such as a GAN-based music generator, WaveNet, etc.) to convert the structural parameters output by the model into a new, personalized physiotherapy audio. It should be noted that the above-mentioned personalized physiotherapy audio data can be used to guide audio synthesis to generate therapeutic audio content with specific rhythm, frequency, and sound pressure characteristics, which can reflect the mapping relationship between individual physiological needs and stimulation strategies.
[0029] In a possible embodiment, the above-mentioned multimodal sound therapy system corresponds and binds the control features in the generated personalized physiotherapy audio data to the execution mechanism in the sound therapy intervention module, for the purpose of matching the corresponding sound therapy intervention module through the personalized physiotherapy audio data and executing the corresponding sound therapy operations, including but not limited to synchronously mapping control parameters such as rhythm, frequency, and amplitude to the execution instructions of the sound therapy intervention module to ensure that the sound output and the physical intervention rhythm are coordinated in time and intensity.
[0030] The above-mentioned sound therapy intervention module may include, but is not limited to, an integrated module of a double-airbag pressure control module, a micro air pump, a pulse driver, an audio output unit, and a linkage control chip, etc., which are functional modules for driving rhythmic intervention behaviors according to the above-mentioned personalized physiotherapy audio data.
[0031] The above-mentioned intervention behaviors may include, but are not limited to, alternately or synchronously applying phased pressure to the upper limbs of the target object to simulate the ischemia-reperfusion response, so as to reduce the peripheral vascular resistance and regulate blood pressure; outputting electrical pulse signals or vibration signals synchronized with the audio rhythm to intervene in the pulse frequency and rhythm stability of the target object; realizing the real-time linkage between the audio control parameters and the execution actions of the intervention device, etc. It can be understood that the generation and execution of the above-mentioned intervention behaviors can be adapted according to the content of the above-mentioned personalized physiotherapy audio data and the corresponding sound therapy intervention module that is driven and matched. For example, when the above-mentioned personalized physiotherapy audio data contains a soothing rhythm component in the low frequency band (such as 20 Hz to 80 Hz), the above-mentioned multimodal sound therapy system can match and drive the cuff-type airbag module in the sound therapy intervention module to slowly inflate and maintain a constant pressure state, simulating diastolic vasodilation, so as to achieve the purpose of reducing diastolic blood pressure; When the rhythm in the above-mentioned personalized physiotherapy audio data speeds up or contains an intermediate frequency pulse segment (such as 100 Hz to 300 Hz), the above-mentioned multimodal sound therapy system can synchronously control the PWM pulse drive module to output electrical pulse signals or micro-vibration signals of corresponding frequencies to perform nerve activation intervention on the target object, for improving the pulse wave stability and autonomic nerve response; When there is a gradually increasing volume change (such as a rhythm structure with a gradually rising sound pressure) in the above-mentioned personalized physiotherapy audio data, the above-mentioned sound therapy intervention module can perform an ischemia-reperfusion cycle operation with phased pressure changes, such as gradually increasing the pressure to systolic blood pressure + 23 mmHg in 5 minutes and quickly exhausting the air in 5 minutes, so as to simulate the process of blood flow blockage and restoration and achieve the preconditioning stimulation effect; If the above-mentioned personalized physiotherapy audio data contains a specific brain wave rhythm induction frequency band (such as α wave 8 - 12 Hz or θ wave 4 - 8 Hz), it can trigger the above-mentioned multimodal sound therapy system to preferentially output the audio through headphones and reduce the physical stimulation intensity, making the intervention behavior more suitable for the resting state or the pre-sleep relaxation scenario; In addition, if abnormal changes in physiological parameters are detected during the intervention (such as a sudden increase in heart rate, systolic blood pressure ≥ 160 mmHg), the above multi-modal sound therapy system can automatically adjust the audio rhythm to the low-frequency band and pause or reduce the output intensity of the intervention module to achieve a protective response to high-risk states.
[0032] In a possible embodiment, the above multi-modal sound therapy system obtains the current physiological parameters of the target object and inputs them into a preset sound therapy processing model for calculation and processing to obtain personalized physiotherapy audio data, and then matches and drives the sound therapy intervention module to perform sound therapy on the target object.
[0033] Through the above method steps, multi-modal physiological parameters such as heart sound, blood pressure, and pulse of the target object can be collected, personalized physiotherapy audio data can be generated in combination with a preset sound therapy processing model, and the sound therapy intervention module can be driven to achieve synchronous linkage of sound control and physical intervention, improving the accuracy and individual adaptability of sound therapy intervention.
[0034] Such as Figure 2 shown, Figure 2 is a flowchart of a multi-modal sound therapy method provided by an embodiment of the present invention. The multi-modal sound therapy method includes the steps: 201. Obtain the current physiological parameters of the target object.
[0035] In an embodiment of the present invention, the above multi-modal sound therapy method can be applied to the above multi-modal sound therapy system. The above multi-modal sound therapy system has functions such as sound therapy audio data processing, sound therapy audio data transceiver, and sound therapy audio data memory storage, and can be built based on a server or a server cluster. The above server or server cluster can be an electronic device with sound therapy audio data processing capabilities.
[0036] The above target object can refer to patients or sub-healthy people who need cardiovascular intervention, blood pressure management, stress relief, nerve regulation, or rehabilitation training. It can be understood that the target object can be in a resting, training, or clinical monitoring state, and the changes in its physiological parameters can be used as the basis for the response and adjustment of the multi-modal sound therapy system.
[0037] The above current physiological parameters can refer to the key physiological information obtained in real time during the start or operation of sound therapy of the target object, which can be used to reflect its current cardiovascular state or original physiological response parameters such as the activity of the autonomic nervous system. Among them, it can include but is not limited to heart sound data, blood pressure data, and pulse data, and can also include other auxiliary physiological signals such as corresponding electroencephalogram, skin electrical response, and respiratory rate.
[0038] More specifically, the above-mentioned heart sound data can refer to the heart sound signal of the target object collected by sensors (such as MEMS microphones, piezoelectric pickups), which is used to represent the mechanical vibration generated when the heart valve closes. It should be noted that after filtering, envelope extraction, rhythm analysis and other processes, characteristic data such as S1 / S2 interval time, peak frequency, waveform morphology, etc. can be obtained. This characteristic data can be used to evaluate heart rate, rhythm stability and cardiac pumping function, providing a basis for heart sound parameters for subsequent audio therapy intervention.
[0039] The above-mentioned blood pressure data can refer to the systolic blood pressure, diastolic blood pressure and their dynamic change information obtained by non-invasive methods (such as oscillometric method with a cuff). The above-mentioned blood pressure data can be used to judge vascular tension, sympathetic nerve activity and overall circulatory load level, and serve as the blood pressure parameter basis for generating audio therapy control strategies.
[0040] The above-mentioned pulse data can refer to the pulse waveform signal obtained by photoplethysmography (PPG) or other non-invasive detection means. The above-mentioned pulse data can include but are not limited to parameters such as pulse frequency, waveform morphology, pulse wave transit time (PWTT), etc. It can be used to estimate the trend of blood pressure change, heart rate variability and vascular compliance, and is the pulse parameter basis for reflecting the state of the peripheral circulation and the autonomic nervous system.
[0041] In a possible embodiment, the above-mentioned multi-modal audio therapy system collects multiple physiological state data of the user in real time through a variety of sensor devices configured at specific parts of the user's body, and uses it as the parameter basis for subsequent audio intervention.
[0042] By deploying various types of physiological sensors at key parts of the user's body through the above method steps to achieve real-time data collection, the perception ability of the multi-modal audio therapy system for the user's current physiological state is enhanced, providing high-timeliness and high-resolution individual state parameter support for subsequent audio intervention strategies, which helps to improve the accuracy and dynamic adaptation ability of audio therapy output.
[0043] 202. Process at least one of the heart sound data, blood pressure data and pulse data through a preset audio therapy processing model to generate personalized physiotherapy audio data.
[0044] In an embodiment of the present invention, the above-mentioned personalized physiotherapy audio data may refer to audio control information that can have a positive intervention effect on the user's current physiological state, which is output by a preset audio therapy processing model based on the target object's current physiological state. Specifically, it can be generated through an audio segment intelligent selection mode and an audio segment dynamic synthesis mode. Among them, the above-mentioned audio segment intelligent selection mode may be a generation mode in which the audio therapy audio data generation platform screens the most matching audio content from a preset audio therapy audio material library according to indicators such as rhythm, frequency band, and style output by the model and pushes it for playback. The above-mentioned audio segment dynamic synthesis may be a process in which the audio therapy audio data generation platform uses an audio synthesis algorithm (such as a GAN-based music generator, WaveNet, etc.) to convert the structural parameters output by the model into a new and personalized audio therapy audio. It should be noted that the above-mentioned personalized physiotherapy audio data can be used to guide audio synthesis, generate therapeutic audio content with specific rhythm, frequency, and sound pressure characteristics, and can reflect the mapping relationship between individual physiological needs and stimulation strategies.
[0045] In a possible embodiment, the above-mentioned multimodal audio therapy system corresponds and binds the control features in the generated personalized physiotherapy audio data to the execution mechanism in the audio therapy intervention module, for the purpose of matching the corresponding audio therapy intervention module through the personalized physiotherapy audio data and executing the corresponding audio therapy operations, including but not limited to synchronously mapping control parameters such as rhythm, frequency, and amplitude to the execution instructions of the audio therapy intervention module to ensure that the sound output and the physical intervention rhythm are coordinated in time and intensity.
[0046] 203. Based on the personalized physiotherapy audio data, match and drive the audio therapy intervention module to perform audio therapy on the target object.
[0047] In an embodiment of the present invention, the above-mentioned audio therapy intervention module may include, but is not limited to, an integrated module of a double-airbag pressure control module, a micro air pump, a pulse driver, an audio output unit, and a linkage control chip, etc., which are used to perform drive rhythmic intervention behaviors according to the above-mentioned personalized physiotherapy audio data.
[0048] The above-mentioned intervention behaviors may include, but are not limited to, alternately or synchronously applying phased pressure to the upper limbs of the target object to simulate the ischemia-reperfusion response, so as to reduce the peripheral vascular resistance and regulate blood pressure; outputting electrical pulse signals or vibration signals synchronized with the audio rhythm to intervene in the pulse frequency and rhythm stability of the target object; realizing the real-time linkage between the audio control parameters and the actions executed by the intervention device, etc. It can be understood that the generation and execution of the above-mentioned intervention behaviors can be adapted according to the content of the above-mentioned personalized physiotherapy audio data and the corresponding matching-driven audio therapy intervention module. For example, when the above-mentioned personalized physiotherapy audio data contains soothing rhythm components in the low-frequency band (such as 20Hz to 80Hz), the above-mentioned multi-modal audio therapy system can match and drive the cuff-type airbag module in the audio therapy intervention module to slowly inflate and maintain a constant pressure state, simulating diastolic vasodilation, so as to achieve the purpose of reducing diastolic blood pressure.
[0049] In the embodiment of the present invention, the current physiological parameters of the target object are obtained, and the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data; at least one of heart sound data, blood pressure data, and pulse data is processed by a preset audio therapy processing model to generate personalized physiotherapy audio data; based on the personalized physiotherapy audio data, a matching-driven audio therapy intervention module is used to perform audio therapy on the target object. Through the adaptive control between multi-modal physiological parameters and personalized audio therapy, data such as the heart sound, blood pressure, and pulse of the target object are collected in real time, and a preset audio therapy processing model is used to extract control features and generate personalized physiotherapy audio, and then drive the audio therapy intervention module to realize the collaborative treatment of auditory stimulation and physical intervention, improving the intelligence and individualization level of non-drug intervention.
[0050] Optionally, in the step of obtaining the current physiological parameters of the target object, where the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data, the current physiological parameters of the target object can also be detected by a preset multi-modal sensor to obtain the first heart sound data, the first blood pressure data, and the first pulse data; based on a preset detection interval, the first heart sound data, the first blood pressure data, and the first pulse data are subjected to time series synchronization processing to obtain the second heart sound data, the second blood pressure data, and the second pulse data; the second heart sound data, the second blood pressure data, and the second pulse data are used as the physiological parameters of the current user.
[0051] In the embodiment of the present invention, the above-mentioned preset multi-modal sensor may include, but is not limited to, modular sensors such as MEMS heart sound sensors, non-invasive blood pressure cuffs, PPG pulse wave sensors, etc. for real-time collection of current physiological parameters. It should be noted that these sensors can all transmit the corresponding current physiological parameters to the above-mentioned multi-modal audio therapy system through transmission protocols such as wireless or Bluetooth.
[0052] In a possible embodiment, the above-mentioned multimodal sound therapy system non-invasively and real-time collects the current physiological state of a target object through a preset multimodal sensor configured on a specific part of the target object's body, including but not limited to operations such as signal sampling, amplification, filtering, and preliminary feature extraction, aiming to obtain the original heart sound, blood pressure, and pulse data reflecting the target object at a certain moment, which are used as the input basis for subsequent sound therapy model processing and intervention strategy generation.
[0053] The above-mentioned preset detection gap can be a unified sampling reference time window set for multi-channel physiological signal acquisition, used to coordinate the differences in the sampling time sequences of different sensors, usually in units of seconds or minutes (such as 5 seconds, 10 seconds). Within this time gap, various physiological data are collected and cached, so as to ensure the logical consistency of subsequent multimodal data in the time dimension for alignment and fusion processing.
[0054] The above-mentioned timing synchronization processing can refer to, for multimodal physiological data collected within the same detection gap but with time differences, through algorithmic means such as interpolation correction, timestamp alignment, and data window registration, to achieve the synchronous mapping of heart sound data, blood pressure data, and pulse data on the same time axis, ensuring the time consistency and relevance of the data when input into the sound therapy processing model, thereby improving the accuracy of multimodal feature fusion.
[0055] The above-mentioned first heart sound data, first blood pressure data, and first pulse data refer to the original physiological signals initially collected by the detection module within the preset detection gap, usually without timing alignment processing; while the second heart sound data, second blood pressure data, and second pulse data refer to the standardized and physiologically time-consistent data sets obtained through timing synchronization processing based on the above original data, which are the effective data benchmarks for subsequent input of the individual's current state into the sound therapy processing model.
[0056] In a possible embodiment, the above-mentioned multimodal sound therapy system executes the first round of detection process at a set sampling frequency and logical judgment conditions, respectively obtains the first heart sound data, first blood pressure data, and first pulse data at the current moment, and after alignment processing based on the preset detection gap, the processing results are respectively recorded as the second heart sound data, second blood pressure data, and second pulse data.
[0057] Through the above method steps, the preprocessing of standardization and alignment of the original multi-source physiological signals improves the accuracy of subsequent AI model modeling and the time synchronization of intervention instruction execution, which is conducive to realizing the stability, intelligence, and closed-loop control accuracy of the multimodal sound therapy system.
[0058] Optionally, in the step of performing time series synchronization processing on the first heart sound data, the first blood pressure data, and the first pulse data based on a preset detection interval to obtain the second heart sound data, the second blood pressure data, and the second pulse data, it further includes determining the beating interval time between two adjacent heart sound data and the pressure difference waveform data between two adjacent pulse data; based on the beating interval time and the pressure difference waveform data, performing fusion processing on the first heart sound data, the first blood pressure data, and the first pulse data to obtain the second heart sound data, the second blood pressure data, and the second pulse data.
[0059] In the embodiments of the present invention, the logical relationship between two data points with the same physiological characteristic type (such as two heart sound peaks, two pulse wave peaks) that appear immediately in the time axis during continuous physiological data acquisition can generally be defined as the relationship between basic units for physiological rhythm calculation within the same detection interval or continuous data frames, which is called "adjacent". Specifically, it can be set as the sequence relationship between two identified first heart sound (S1) peak points in the continuous heart sound signal within a time window with a preset detection interval of 10 seconds. Taking a sampling rate of 1000 Hz as an example, if 5 first heart sounds are identified within the detection interval, the system will regard the 1st and 2nd S1 as a group of adjacent heart sound events for calculating the beating interval time; similarly, in pulse wave acquisition, if 6 effective wave peaks are identified, the 3rd and 4th pulse wave peaks can also be used as a group of "adjacent" pulse data to further extract the pressure difference waveform data.
[0060] The above-mentioned beating interval time can refer to the time interval between two adjacent heart sound events in the heart sound data. Preferably, it corresponds to the time difference between the first heart sound (S1) and the next first heart sound, which is equivalent to the length of a complete cardiac cycle. It should be noted that this interval time can reflect the heart rate and heart rate variability of the target object, can be used to judge the stability of the cardiac rhythm, and can also be used as a reference anchor point for multi-modal data time series alignment.
[0061] The above-mentioned pressure difference waveform data can refer to the dynamic waveform curve formed by the pressure change trend between two adjacent pulse wave peaks in the pulse wave data, which can usually be derived by a PPG sensor in combination with the pulse wave transit time (PWTT). This data reflects the vascular pressure change and peripheral resistance fluctuation per unit time, and can be used to judge vascular compliance, peripheral circulation efficiency, and pressure response sensitivity.
[0062] In a possible embodiment, the above-mentioned multi-modal sound therapy system can perform multi-step fusion processing on the separately collected heart sound data, blood pressure data, and pulse data according to the time axis between the beating interval time and the pressure difference waveform data, feature normalization, signal unified coding, and other feature data, and construct a unified data input structure with physiological rhythm consistency and modal complementarity.
[0063] Optionally, in the step of generating personalized physiotherapy audio data by processing at least one of the heart sound data, blood pressure data, and pulse data through a preset sound therapy processing model, it further includes extracting features from at least one of the heart sound data, blood pressure data, and pulse data to obtain control feature vectors corresponding to at least one of the heart sound data, blood pressure data, and pulse data; performing fusion processing based on at least one corresponding control feature vector to determine corresponding control audio feature data; and synthesizing personalized physiotherapy audio data according to the control audio feature data.
[0064] In the embodiment of the present invention, corresponding signal features can be extracted from any one of the above-mentioned heart sound data, blood pressure data, and pulse data and other original physiological signal data to obtain corresponding control feature vectors. Specifically, it can include but is not limited to operations such as filtering, normalization, spectral transformation, waveform segmentation, key point recognition, etc., to convert complex time-domain or frequency-domain signals into a structured data expression form, and obtain corresponding control feature vectors. More specifically, the above-mentioned control feature vector can refer to a parameter vector output by feature extraction for describing an individual's current physiological state and can be directly used for audio control generation. The parameter vector can include but is not limited to fields such as heart sound rhythm interval, systolic blood pressure amplitude, and pulse wave conduction time, and is formed into a unified format through standardized coding as the basic input for driving audio rhythm, frequency, energy, and other control parameters.
[0065] The above-mentioned control audio feature data can refer to a set of control parameter sets output by fusion processing for guiding audio content generation, specifically including information such as audio rhythm (such as BPM), frequency range (such as low-frequency / high-frequency distribution), amplitude intensity (such as sound pressure level), and modulation mode (such as oscillation rhythm), etc. This data directly determines the performance form of the audio generated by subsequent audio synthesis and is the direct control signal source for personalized sound therapy content output.
[0066] In a possible embodiment, the above-mentioned multi-modal sound therapy system generates a unified input structure that can comprehensively express an individual's current overall physiological state, that is, control audio feature data, through the process of integrating multiple control feature vectors from different sources (such as vectors corresponding to heart sound, blood pressure, and pulse) in the feature dimension or time dimension. It can be understood that the above-mentioned fusion methods can include but are not limited to technical means such as feature splicing, weighted average, feature attention mechanism, or neural network fusion layer.
[0067] In another possible embodiment, the above multi-modal audio therapy system takes control audio feature data as input and converts it into a playable digital audio signal through an audio generator or synthesizer. This process may include methods such as spectrogram reconstruction, waveform prediction, parametric synthesis, or calling existing audio templates. Finally, a personalized physiotherapy audio file with therapeutic attributes, rhythm structure, and mood regulation effects is generated for implementing auditory stimulation and intervention linkage control on the target object.
[0068] Optionally, in the step of performing fusion processing based on at least one corresponding control feature vector to determine the corresponding control audio feature data, it further includes extracting the time series data of at least one corresponding control feature vector; performing time series alignment processing on the corresponding time series data, fusing at least one corresponding control feature vector to obtain fusion feature data; generating control audio feature data including frequency, rhythm, and amplitude control parameters based on the fusion feature data; during the generation process, collecting the changes in the physiological parameters of the target object and adjusting the control audio feature data based on the collection results.
[0069] In the embodiment of the present invention, the above time series data may refer to a physiological feature sequence continuously collected by a sensor and arranged in chronological order, such as the change in heart syllable rhythm, continuous blood pressure readings, or pulse waveform indicators within a certain time window. This time series data can be used to reflect the dynamic trend of the physiological parameters of the target object over time.
[0070] In a possible embodiment, for the situation where the sampling time, starting point, or time resolution of different modal control feature vectors are inconsistent, the above multi-modal audio therapy system makes multiple feature sequences aligned on a unified time axis by selecting interpolation, resampling, timestamp mapping, etc., to achieve the process of time series alignment processing.
[0071] The above fusion feature data refers to a high-dimensional feature representation generated by uniformly encoding, splicing, or non-linearly combining multiple control feature vectors (such as heart sound control vector, blood pressure control vector, pulse control vector) after time series alignment in the feature space, which can comprehensively describe the current overall physiological state of the target object.
[0072] The above changes in physiological parameters may refer to the dynamic changes in the key physiological indicators (such as blood pressure, heart rate, pulse wave, etc.) of the target object during the execution of audio therapy intervention, which can be collected in real time by the preset audio therapy sensor supporting the above multi-modal audio therapy system and the trend can be judged.
[0073] In another possible embodiment, the above multi-modal audio therapy system modifies the parameters or dynamically optimizes the generated control audio feature data according to the change result of the current physiological parameters of the target object, and strategies including but not limited to rhythm slowdown, frequency downshift, volume reduction, etc. can be used to ensure that the audio therapy output is synchronized with the actual state of the individual, and improve the personalization and safety of the treatment.
[0074] Optionally, in the step of performing audio therapy on the target object by the matching-driven audio therapy intervention module based on the personalized physiotherapy audio data, it further includes determining the control audio feature data in the current personalized physiotherapy audio data in real time; driving the corresponding audio therapy intervention module based on the control audio feature data, so that the audio therapy intervention module performs rhythmic physiotherapy intervention according to the control audio feature data, and the audio therapy intervention module is used to perform physical physiotherapy intervention on the user in cooperation with the personalized physiotherapy audio data.
[0075] In the embodiment of the present invention, the above control audio feature data may include core information carriers such as audio rhythm (beats per minute, BPM), frequency distribution (such as 20–100 Hz), sound pressure intensity (Sound Pressure Level, SPL), and their change trends for implementing rhythmic control logic.
[0076] The above rhythmic physiotherapy intervention may include but not limited to controlling the inflation and deflation of the cuff airbag according to the audio rhythm beats (such as 60 times per minute); adjusting the output frequency and amplitude of the electrical pulse according to the high and low audio frequencies; synchronously applying vibration stimulation at the rhythm peak for triggering vasodilation and nerve reflexes and other audio therapy function processes performed by the audio therapy intervention module.
[0077] In a possible embodiment, while generating the personalized physiotherapy audio data, the above multi-modal audio therapy system also synchronously extracts and determines in real time the control audio feature data contained in the audio data, and further drives the configured audio therapy intervention module to perform rhythm control operations according to the above control audio feature data. The intervention module and the audio playback module work together to make the auditory stimulation (audio output) and the physical stimulation (intervention output) highly consistent in rhythm, and construct a "sound-pressure" or "sound-electricity" linkage intervention mechanism.
[0078] This mechanism can be automatically controlled and updated by the above multi-modal audio therapy system during the audio therapy execution process to ensure that the intervention behavior highly matches the current physiological rhythm of the target object. It should be noted that the above audio therapy intervention module can communicate with the audio therapy main control chip through low-power Bluetooth and refine the execution instructions through PWM (pulse width modulation) method to ensure high-precision and high-responsiveness of rhythm control, and can perform offline control.
[0080] Through the above method steps, not only is the precise matching between the audio control parameters and the physical intervention behaviors achieved, but also the consistency and dynamic linkage ability of the multimodal intervention are effectively improved; compared with the traditional playback-based music therapy, it can achieve the coordinated stimulation of the auditory and physical dual channels, the same frequency band, and the same rhythm, enhance the intervention depth and the physiological response effect, and improve the effectiveness, personalized adaptation ability, and closed-loop regulation ability of the music therapy treatment.
[0081] As Figure 3 shown, an embodiment of the present invention further provides a multimodal music therapy device 300, and the multimodal music therapy device 300 includes: A first acquisition module 301, configured to acquire the current physiological parameters of a target object, where the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data; A first generation module 302, configured to process the at least one of heart sound data, blood pressure data, and pulse data through a preset music therapy processing model to generate personalized physiotherapy audio data; A first matching module 303, configured to match and drive a music therapy intervention module to perform music therapy processing on the target object based on the personalized physiotherapy audio data.
[0082] Optionally, the above first acquisition module 301 includes: A first detection sub-module, configured to detect the current physiological parameters of the target object through a preset multimodal sensor to obtain first heart sound data, first blood pressure data, and first pulse data; A first processing sub-module, configured to perform timing synchronization processing on the first heart sound data, first blood pressure data, and first pulse data based on a preset detection interval to obtain second heart sound data, second blood pressure data, and second pulse data; A second processing sub-module, configured to use the second heart sound data, second blood pressure data, and second pulse data as the physiological parameters of the current user.
[0083] Optionally, the above first processing sub-module includes: A first determination unit, configured to determine the beating interval time between two adjacent heart sound data and the pressure difference waveform data between two adjacent pulse data; A first fusion unit, configured to fuse the first heart sound data, first blood pressure data, and first pulse data based on the beating interval time and the pressure difference waveform data to obtain second heart sound data, second blood pressure data, and second pulse data.
[0084] Optionally, the above first generation module 302 includes: The first extraction sub-module is used to extract features from the at least one heart sound data, blood pressure data, and pulse data to obtain a control feature vector corresponding to the at least one heart sound data, blood pressure data, and pulse data; The first determination sub-module is used to perform fusion processing based on the at least one corresponding control feature vector to determine corresponding control audio feature data; The first synthesis sub-module is used to synthesize personalized physiotherapy audio data according to the control audio feature data.
[0085] Optionally, the above device further includes: An extraction module is used to extract the time series data of the at least one corresponding control feature vector; A fusion module is used to perform time series alignment processing on the corresponding time series data, fuse the at least one corresponding control feature vector, and obtain fused feature data; A generation module is used to generate control audio feature data including frequency, rhythm, and amplitude control parameters based on the fused feature data; A feedback module is used to collect changes in the physiological parameters of the target object during the generation process and adjust the control audio feature data based on the collection results.
[0086] Optionally, the above first matching module 303 includes: A second determination sub-module is used to determine the control audio feature data in the current personalized physiotherapy audio data in real time; A driving sub-module is used to drive the corresponding audio therapy intervention module based on the control audio feature data, so that the audio therapy intervention module performs rhythmic physiotherapy intervention according to the control audio feature data, and the audio therapy intervention module is used to perform physical physiotherapy intervention on the user in cooperation with the personalized physiotherapy audio data.
[0087] As Figure 4 shown, an embodiment of the present invention further provides an electronic device 400, including a processor, and the above processor can execute any one of the above multi-modal audio therapy methods.
[0088] Specifically, it includes a processor 401 and a memory 402, and a computer program for executing the multi-modal audio therapy method stored on the memory 402 and capable of running on the processor 401, where: The processor 401 runs the calculator program of the multi-modal audio therapy method stored in the memory 402 and executes the following steps: Obtain the current physiological parameters of the target object, where the physiological parameters include at least one heart sound data, blood pressure data, and pulse data; Process the at least one heart sound data, blood pressure data, and pulse data through a preset audio therapy processing model to generate personalized physiotherapy audio data; Based on the personalized physiotherapy audio data, the matching driving audio therapy intervention module performs audio therapy processing on the target object.
[0089] Optionally, the processor 401 executes the obtaining of the current physiological parameters of the target object, where the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data, and includes: Detect the current physiological parameters of the target object through a preset multi-modal sensor to obtain first heart sound data, first blood pressure data, and first pulse data; Based on a preset detection interval, perform time series synchronization processing on the first heart sound data, first blood pressure data, and first pulse data to obtain second heart sound data, second blood pressure data, and second pulse data; Use the second heart sound data, second blood pressure data, and second pulse data as the physiological parameters of the current user.
[0090] Optionally, the processor 401 executes the time series synchronization processing on the first heart sound data, first blood pressure data, and first pulse data based on a preset detection interval to obtain second heart sound data, second blood pressure data, and second pulse data, and includes: Determine the beating interval time between two adjacent heart sound data and the pressure difference waveform data between two adjacent pulse data; Based on the beating interval time and the pressure difference waveform data, perform fusion processing on the first heart sound data, first blood pressure data, and first pulse data to obtain second heart sound data, second blood pressure data, and second pulse data.
[0091] Optionally, the processor 401 executes the processing of the at least one of heart sound data, blood pressure data, and pulse data through a preset audio therapy processing model to generate personalized physiotherapy audio data, and includes: Extract features from the at least one of heart sound data, blood pressure data, and pulse data to obtain control feature vectors corresponding to the at least one of heart sound data, blood pressure data, and pulse data; Based on the fusion processing of the at least one corresponding control feature vector, determine the corresponding control audio feature data; According to the control audio feature data, synthesize personalized physiotherapy audio data.
[0092] Optionally, the processor 401 executes the fusion processing based on the at least one corresponding control feature vector to determine the corresponding control audio feature data, and the method further includes: Extract the time series data of the at least one corresponding control feature vector; Perform temporal alignment processing on the corresponding temporal data, and fuse the at least one corresponding control feature vector to obtain fused feature data; Based on the fused feature data, generate control audio feature data including frequency, rhythm, and amplitude control parameters; During the generation process, collect the physiological parameter changes of the target object, and adjust the control audio feature data based on the collection results.
[0093] Optionally, the processor 401 also executes the matching drive audio therapy intervention module to perform audio therapy on the target object based on the personalized physical therapy audio data, including: Determine the control audio feature data in the current personalized physical therapy audio data in real time; Based on the control audio feature data, drive the corresponding audio therapy intervention module, so that the audio therapy intervention module performs rhythmic physical therapy intervention according to the control audio feature data. The audio therapy intervention module is used to cooperate with the personalized physical therapy audio data to perform physical therapy intervention on the user.
[0094] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it realizes each process of the multi-modal audio therapy method or the application-side multi-modal audio therapy method provided by the embodiment of the present invention, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.
[0095] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0096] The above-disclosed are only the preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present invention still fall within the scope covered by the present invention.
Claims
1. A multimodal sound therapy method, characterized in that: include: Acquiring current physiological parameters of the target object, wherein the physiological parameters include at least one of heart sound data, blood pressure data, and pulse data; Processing the at least one heart sound data, blood pressure data and pulse data by a preset sound therapy processing model to generate personalized physical therapy audio data; Based on the personalized therapy audio data, a sound therapy intervention module is matched and driven to perform sound therapy treatment on the target object.
2. The multimodal sound therapy method according to claim 1, characterized in that: The step of acquiring the current physiological parameters of the target object, wherein the physiological parameters include at least one of heart sound data, blood pressure data and pulse data, comprises: Detecting the current physiological parameters of the target object by a preset multimodal sensor to obtain first heart sound data, first blood pressure data, and first pulse data; Based on a preset detection interval, performing time-series synchronization processing on the first heart sound data, the first blood pressure data, and the first pulse data to obtain second heart sound data, second blood pressure data, and second pulse data; The second heart sound data, the second blood pressure data and the second pulse data are used as physiological parameters of the current user.
3. The multimodal sound therapy method according to claim 2, characterized in that: The method of performing time-series synchronization processing on the first heart sound data, the first blood pressure data, and the first pulse data based on the preset detection interval to obtain second heart sound data, second blood pressure data, and second pulse data includes: Determine the beat interval time between two adjacent heart sound data and the pressure difference waveform data between two adjacent pulse data; Based on the beat interval time and the pressure difference waveform data, the first heart sound data, the first blood pressure data and the first pulse data are fused to obtain second heart sound data, second blood pressure data and second pulse data.
4. The multimodal sound therapy method according to claim 1, characterized in that: The processing of the at least one heart sound data, blood pressure data and pulse data by a preset sound therapy processing model to generate personalized physical therapy audio data includes: Extracting features of the at least one heart sound data, blood pressure data and pulse data to obtain a control feature vector corresponding to the at least one heart sound data, blood pressure data and pulse data; Perform fusion processing based on the at least one corresponding control feature vector to determine corresponding control audio feature data; According to the control audio feature data, personalized physical therapy audio data is synthesized.
5. The multimodal sound therapy method according to claim 4, characterized in that: The method further comprises: performing fusion processing based on the at least one corresponding control feature vector to determine corresponding control audio feature data; Extracting time series data of the at least one corresponding control feature vector; Performing time series alignment processing on the corresponding time series data, fusing the at least one corresponding control feature vector to obtain fused feature data; Based on the fused feature data, generating control audio feature data including frequency, rhythm and amplitude control parameters; During the generation process, changes in physiological parameters of the target object are collected, and the control audio feature data is adjusted based on the collected results.
6. The multimodal sound therapy method according to claim 1, characterized in that: The matching and driving of the sound therapy intervention module to perform sound therapy on the target object based on the personalized physical therapy audio data includes: Determine in real time the control audio feature data in the current personalized physical therapy audio data; Based on the control audio feature data, the corresponding sound therapy intervention module is driven so that the sound therapy intervention module performs rhythmic physical therapy intervention according to the control audio feature data. The sound therapy intervention module is used to perform physical therapy intervention on the user in conjunction with personalized physical therapy audio data.
7. A multimodal sound therapy device, characterized in that: include: A first acquisition module, used to acquire current physiological parameters of the target object, wherein the physiological parameters include at least one of heart sound data, blood pressure data and pulse data; A first generating module, configured to process the at least one heart sound data, blood pressure data and pulse data by using a preset sound therapy processing model to generate personalized physical therapy audio data; The first matching module is used to match and drive the sound therapy intervention module to perform sound therapy treatment on the target object based on the personalized physical therapy audio data.
8. A multimodal sound therapy system, characterized in that: The multimodal sound therapy system comprises: a multimodal sound therapy device; The multimodal sound therapy device implements a multimodal sound therapy method described in claim 1.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps in the multimodal sound therapy method as described in any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the multimodal sound therapy method as described in any one of claims 1 to 6.