A personalized sound medicine formula generation and evaluation method based on voice emotion recognition

By using voice emotion recognition and deep neural networks, personalized sound-based medication formulas are generated, solving the problem of inaccurate prediction of effects in existing music therapy programs. This achieves personalized and refined sound-based medication formula matching, thereby improving treatment outcomes.

CN122290642APending Publication Date: 2026-06-26SHANDONG SHANGYI HEALTHCARE TRADITIONAL CHINESE MEDICINE TECHNOLOGY DEVELOPMENT CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG SHANGYI HEALTHCARE TRADITIONAL CHINESE MEDICINE TECHNOLOGY DEVELOPMENT CO LTD
Filing Date
2025-12-02
Publication Date
2026-06-26

Smart Images

  • Figure CN122290642A_ABST
    Figure CN122290642A_ABST
Patent Text Reader

Abstract

This invention relates to a method for generating and evaluating personalized sound therapy formulas based on voice emotion recognition, belonging to the field of voice analysis, and more specifically to the field of digital music development and production. The method includes: using an AI sound therapy effect judgment model to intelligently determine the increase in the target user's happiness level after using each five-tone sound therapy formula based on the target user's current five-element attribute data and current five-organ attribute data, and the digital content of each five-tone sound therapy formula; selecting the five-tone sound therapy formula with the largest increase value as the target user's personalized sound therapy formula. This invention addresses the technical problems of difficulty in providing the most suitable personalized music therapy plan for different users and the insufficient precision of music therapy plans. It employs an AI-based traversal method to intelligently judge each sound therapy formula for the target user, and uses five-tone five-element sound therapy formulas to improve the precision of the sound therapy formulas, thereby solving the aforementioned technical problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The speech analysis of this invention relates more specifically to the field of digital music development and production, and in particular to a method for generating and evaluating personalized sound drug formulas based on speech emotion recognition. Background Technology

[0002] Health-related information obtained through voice analysis can be used for appropriate health management. Voice analysis is a key process in digital music development and production. Feedback from voice analysis provides crucial reference data for optimizing digital music development and production. When developed digital music is used as a therapeutic formula for patients, voice analysis feedback can also be used to optimize the therapeutic formula. More specifically, by analyzing different voice segments of patients before and after treatment, their emotional state before and after treatment can be identified, providing crucial reference data for the therapeutic effect of the developed digital music.

[0003] For example, Chinese invention patent publication CN102188773A discloses a digital music therapy device, which includes a music playback module and a pulse therapy module. The music playback module has two output terminals: one for direct playback and one connected to the pulse therapy module. The pulse therapy module includes a music amplification module for extracting and amplifying music signals; a pulse generation module for converting the amplified music signals into pulse waveforms; and a voltage boosting module for amplifying the pulse waveforms into therapeutic electrical signals that act on acupoints. It can play music while simultaneously generating therapeutic electrical signals synchronized with the music's rhythm, enhancing the integration of music and electrotherapy. It also features a blood pressure measurement module, an emotion analysis module, and a music selection module, allowing for targeted treatment based on the user's emotions, selecting music more suitable for the user, and achieving better therapeutic effects.

[0004] For example, Chinese invention patent publication CN111870877A discloses a music conducting and music therapy system. The system includes a treatment room body, which comprises a patient room and a doctor's room, separated by a wall. One-way glass is nested within the wall body. A treatment mechanism is installed within the patient room body, and the treatment mechanism includes a support platform fixed to the bottom surface of the patient room. A display screen is fixedly connected to the top of the support platform, and an infrared receiver is installed inside the display screen. A conductive device is connected to the surface of the display screen via wires. In this invention, under the action of the treatment mechanism, music conducting instructional videos can be played on the display screen, allowing the patient to perform music conducting movements by holding a handheld device. Through the infrared generator, PLC control board, and infrared receiver, the patient can judge whether the music conducting movements are correct, improving the patient's motor coordination and increasing the efficiency of treatment.

[0005] However, existing music therapy solutions either require following the user's emotional state to execute subsequent music arrangements, or simply rely on the user's reaction to music for music-assisted therapy. The former may lead to deviations in the music therapy effect because the user's emotions precede the music therapy, while the latter requires the user to move with the music and is urgently used to improve the user's motor coordination. Neither of these treatment solutions can accurately predict the therapeutic effect of music therapy, lacks feedback data on the effect of music therapy, and therefore cannot provide the most suitable personalized music therapy solution for different users. At the same time, the music therapy solutions provided lack sufficient precision. Summary of the Invention

[0006] To address technical issues in related fields, this invention provides a method for generating and evaluating personalized sound-based medication formulas based on voice emotion recognition. By employing basic data including voice recognition results and a customized artificial intelligence model for intelligently judging the therapeutic effect of sound-based medication formulas, the method intelligently judges various sound-based medication formulas for the target user through a traversal approach to obtain the optimal personalized sound-based medication formula matching the target user. Simultaneously, it utilizes a five-tone, five-element sound-based medication formula to enhance the precision of the formula, thereby ensuring the therapeutic effect of sound-based medications for different users through a feedback mechanism.

[0007] According to the present invention, a method for generating and evaluating personalized sound-based drug prescriptions based on voice emotion recognition is provided, the method comprising: Collect voice samples of a set time length from the target user to serve as the target user's current voice sample. Perform uniform segmentation on the target user's current voice sample to obtain multiple current voice sub-samples. Convert each current voice sub-sample into activation value and valence value in a two-dimensional emotion space to serve as the target user's voice emotion feature value in that current voice sub-sample. An intelligent emotion recognition model is used to intelligently identify the target user's current Five Elements attribute data and the target user's current Five Organs attribute data based on the numerical values ​​of multiple voice emotion features corresponding to multiple current voice sub-samples. By traversing various frequencies, rhythms, and sequences of the five tones, we can obtain various five-tone medicine recipes. The duration of each five-tone medicine recipe is equal to the set time length. The AI-powered sound medicine effect judgment model uses a set time length, multiple user parameters of the target user, the target user's current five elements attribute data and the target user's current five internal organs attribute data, and the WMA format digital content of each five-tone sound medicine formula to intelligently judge the increase in the target user's happy mood level after using the five-tone sound medicine formula. The five-tone sound medicine formula with the largest corresponding increase value will be matched to the target user as the personalized sound medicine formula; The AI-powered sound drug effect judgment model is a deep neural network that has been trained multiple times, and the number of training sessions is positively correlated with the number of current speech sub-samples.

[0008] Compared with the prior art, the present invention has at least the following main inventive points: Invention Point A: For target users of sound-based drug treatment, a voice sample of the target user is collected; features are extracted from the voice sample to obtain voice emotion feature values; the voice emotion feature values ​​are input into an intelligent emotion recognition model after multiple trainings to identify the target user's current five-element attribute data and current five-organ attribute data; based on the recognition results and each five-tone sound-based drug formula traversed, the improvement of each five-tone sound-based drug formula for the target user's happy emotion level is intelligently judged; the five-tone sound-based drug formula with the highest improvement is used as the personalized sound-based drug formula to match the target user, thereby completing the generation and evaluation of personalized sound-based drug formulas based on voice emotion recognition, ensuring the sound-based drug treatment effect for different users; Invention Point B: To intelligently determine the extent to which each of the five-tone sound medicine formulas improves the target user's level of happiness, a customized AI sound medicine effect judgment model is used. The AI ​​sound medicine effect judgment model is a deep neural network that has been trained multiple times, and the number of training times is positively correlated with the number of current voice sub-samples. The deep neural network includes multiple hidden layers, and the number of hidden layers is positively correlated with the set time length, thereby ensuring the reliability and stability of the intelligent judgment results. Invention Point C: To intelligently determine the extent to which each Five-Tone Sound Medicine Formula improves the target user's level of happiness, multiple basic data are specifically selected. These include the set duration, multiple user parameters of the target user, the target user's current Five Elements attribute data and current Five Organs attribute data, and the WMA format digital content of each Five-Tone Sound Medicine Formula. The multiple user parameters of the target user include the target user's gender identifier, age information, disease type code, and disease duration. The target user's current Five Elements attribute data and current Five Organs attribute data are derived from the intelligent recognition results of the currently collected voice samples of the target user. Furthermore, the formulas are obtained by traversing various frequencies, rhythms, and sequences of the Five Tones. The duration of each Five-Tone Sound Medicine Formula is equal to the set duration. Through the targeted selection of the above multiple basic data, the reliability and stability of the intelligent judgment results are further guaranteed. Invention Point D: In each training iteration of the deep neural network, the increase in a user's happiness level after using a known Five-Tone Sound Medicine formula is used as the output of the deep neural network. The set time length, multiple user parameters of the user, the user's current Five Elements attribute data and the user's current Five Organs attribute data, and the WMA format digital content of the Five-Tone Sound Medicine formula are used as the input of the deep neural network to complete the training, thereby ensuring the training effect of the deep neural network in each iteration. Invention Point E: In the specific sound medicine formula traversal operation, traversing the various frequencies of the five tones refers to traversing the frequency of occurrence of each note in the five tones within a set time length; traversing the various sequences of the five tones refers to traversing the display sequence of each note in the five tones during each display within a set time length; and traversing the various rhythms of the five tones refers to traversing the display rhythm of each note in the five tones during each display within a set time length. The above three traversals are carried out simultaneously to obtain each set of five-tone sound medicine formulas, thereby completing the traversal of a massive number of five-tone sound medicine formulas and providing basic data for the intelligent analysis and judgment of the therapeutic effects of the subsequent massive number of five-tone sound medicine formulas. Invention Point F: The happiness level of the target user corresponding to each speech subsample is determined by the activation value and valence value in the speech emotion feature value of the target user corresponding to each speech subsample. The arithmetic mean of the multiple happiness levels corresponding to multiple current speech subsamples of the target user's current speech sample is taken as the happiness level of the target user corresponding to the current speech sample, thereby providing a targeted analysis mechanism for the analysis of the happiness levels of different users at different times. Attached Figure Description

[0009] The embodiments of the present invention will now be described with reference to the accompanying drawings, wherein: Figure 1 This is a schematic diagram of the working scenario of the personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to the present invention.

[0010] Figure 2 This is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 1 of the present invention.

[0011] Figure 3 The following is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 2 of the present invention.

[0012] Figure 4 This is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 3 of the present invention.

[0013] Figure 5 This is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 4 of the present invention.

[0014] Figure 6 This is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 5 of the present invention.

[0015] Attached image labels: 1. Microphone; 2. Five-tone sound-absorbing formula; 3. Earphone Detailed Implementation

[0016] like Figure 1 The diagram illustrates a working scenario for a personalized voice-based medication formulation generation and evaluation method based on voice emotion recognition, according to the present invention. The voice analysis of this invention specifically relates to the field of digital music development and production.

[0017] exist Figure 1 In this process, the target user listens to the Five-Tone Sound Medicine Formula 2 through headphones 3 to receive music therapy, and the target user collects a voice sample of a set duration through microphone 1 as the target user's current voice sample, which is used to obtain the target user's voice emotion feature value through voice recognition. The specific treatment scenario can be a closed sound therapy venue, such as a room enclosed by soundproofing materials.

[0018] The specific technical process of this invention is as follows: Technical Process 1: For the target users of the five-tone sound medicine formula treatment, traverse the massive number of available formulas. Specifically, in the sound medicine formula traversal operation, each five-tone sound medicine formula is obtained by traversing various frequencies, rhythms and sequences of the five tones. The duration of each five-tone sound medicine formula is equal to the set time length. More specifically, traversing the various frequencies of the five notes refers to traversing the frequency of occurrence of each note in the five notes within a set time length; traversing the various sequences of the five notes refers to traversing the order in which each note in the five notes is displayed in each performance within a set time length; and traversing the various rhythms of the five notes refers to traversing the rhythm in which each note in the five notes is displayed in each performance within a set time length. The above three traversals are performed simultaneously to obtain each five-note medicine formula. In this way, the massive number of five-tone medicine formulas are traversed, providing basic data for the intelligent analysis and judgment of the treatment effects of the massive number of five-tone medicine formulas. Technical Process 2: For the target users of the sound-based medicine formula treatment, a customized AI sound-based medicine effect judgment model is designed to intelligently determine the extent to which each of the five-tone sound-based medicine formulas improves the target user's level of happiness. Specifically, the customized structural design of the AI ​​sound-based drug efficacy judgment model is mainly reflected in the following aspects: Firstly, the AI-powered sound drug effect judgment model used is a deep neural network that has been trained multiple times, and the number of training times is positively correlated with the number of current speech sub-samples. The second aspect: AI sound drug effect judgment model. The deep neural network architecture used in the AI ​​sound drug effect judgment model includes multiple hidden layers, a single input layer and a single output layer. The number of hidden layers is positively correlated with the set time length. Thirdly: In each training session of the deep neural network, the increase in a user's happiness level after using a known Five-Tone Sound Medicine formula is used as the output of the deep neural network. The set time length, multiple user parameters of the user, the user's current Five Elements attribute data and the user's current Five Organs attribute data, and the WMA format digital content of the Five-Tone Sound Medicine formula are used as the input of the deep neural network to complete the training, thereby ensuring the training effect of the deep neural network in each training session. In this way, through the customized structural design of the above-mentioned aspects, the reliability and stability of the intelligent judgment results of the treatment effect of each five-tone medicine prescription are guaranteed. Technical Process 3: For the target users of the sound-based medicine formula treatment, in order to intelligently determine the extent to which each sound-based medicine formula improves the target user's level of happiness, multiple basic data are selected accordingly. Specifically, the multiple basic data include the set time length, multiple user parameters of the target user, the target user's current five elements attribute data and the target user's current five internal organs attribute data, and the WMA format digital content of each five-tone medicine formula; More specifically, the target user's multiple user parameters include the target user's gender identifier, age information, disease type code, and duration of disease; More specifically, the target user's current Five Elements attribute data and the target user's current Five Organs attribute data are derived from the intelligent recognition results of the target user's current voice sample collected at the moment. Specifically, a voice sample of the target user is collected as the current voice sample, and features are extracted from the current voice sample to obtain voice emotion feature values. The voice emotion feature values ​​are then input into the intelligent emotion recognition model after multiple training sessions to identify the target user's current Five Elements attribute data and current Five Organs attribute data. More specifically, by traversing various frequencies, rhythms, and sequences of the five tones, each five-tone medicine formula is obtained, and the duration of each five-tone medicine formula is equal to the set time length. In this way, by selectively analyzing the aforementioned fundamental data, the reliability and stability of the intelligent assessment results of the therapeutic effects of each five-tone herbal medicine prescription are further enhanced. Technical Process 4: The AI ​​sound medicine effect judgment model adopts the customized structure design of Technical Process 2. Based on multiple basic data selected in Technical Process 3, it intelligently judges the degree of improvement in the target user's happy mood level after each five-tone sound medicine formula goes through the technical process. Specifically, the increase in the target user's happiness level is directly given by the artificial intelligence model. In each training process of the deep neural network, when it is used as the output data of the deep neural network, the following steps are required to calculate the target user's happiness level before or after use: In the voice samples of the target user collected before or after use, the activation value and valence value of the voice emotion feature value of the target user corresponding to each voice sub-sample are used to determine the happiness level of the target user corresponding to each voice sub-sample. The arithmetic mean of the multiple happiness levels corresponding to the multiple current voice sub-samples of the target user's voice sample is taken as the happiness level of the target user corresponding to the voice sample of the target user. Technical Process 5: For the target users of the sound medicine formula treatment, obtain a large number of five-tone sound medicine formulas and the improvement of each happy mood level. The five-tone sound medicine formula with the highest improvement is used as the personalized sound medicine formula to match the target user. It can be seen that through the coordinated operation of the above five technical processes, the generation and evaluation of personalized sound-based drug prescriptions based on voice emotion recognition were completed. While improving the accuracy of sound-based drug prescription production, the use of a feedback mode based on the basic data of voice recognition results ensured the therapeutic effect of sound-based drugs for different users.

[0019] The key points of this invention are: intelligent judgment of various sound medicine formulas for target users based on artificial intelligence traversal method, use of five-tone and five-element sound medicine formulas to improve the precision of sound medicine formulas, customized structural design of artificial intelligence model, customized analysis of user's five-element attribute data and five-organ attribute data based on speech recognition, and targeted identification of user's happy mood level at different times.

[0020] The following will describe in detail the method for generating and evaluating personalized voice-based drug prescriptions based on voice emotion recognition of the present invention through examples.

[0021] Example 1 Figure 2 This is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 1 of the present invention.

[0022] like Figure 2 As shown, the personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition includes the following specific steps: Step S21: Collect voice samples of a set time length from the target user as the target user's current voice sample, perform uniform segmentation on the target user's current voice sample to obtain multiple current voice sub-samples, and convert each current voice sub-sample into activation value and valence value in a two-dimensional emotional space as the target user's voice emotion feature value in that current voice sub-sample. For example, a set time period of voice samples is collected from the target user as the target user's current voice sample. The target user's current voice sample is uniformly segmented to obtain multiple current voice sub-samples. Each current voice sub-sample is converted into activation value and valence value in a two-dimensional emotional space as the target user's voice emotional feature value in that current voice sub-sample. The target user is a user to be treated with a sound-based medicine formula. The treatment scenario can be a closed sound therapy venue, such as a room enclosed by soundproofing materials. For example, collecting voice samples of a set duration from the target user as the target user's current voice sample, performing uniform segmentation on the target user's current voice sample to obtain multiple current voice sub-samples, and converting each current voice sub-sample into activation and valence values ​​in a two-dimensional emotional space as the target user's voice emotion feature values ​​in that current voice sub-sample also includes: the set duration can be 5 minutes, performing uniform segmentation on the target user's current voice sample to obtain 30 current voice sub-samples, and the duration of each current voice sub-sample is 10 seconds; Step S22: Using an intelligent emotion recognition model, the target user's current Five Elements attribute data and the target user's current Five Organs attribute data are intelligently identified based on the multiple voice emotion feature values ​​corresponding to multiple current voice sub-samples; Specifically, the intelligent emotion recognition model intelligently identifies the target user's current five elements attribute data and the target user's current five internal organs attribute data based on multiple voice emotion feature values ​​corresponding to multiple current voice sub-samples. This includes: the intelligent recognition process can be tested and simulated using a numerical simulation mode. Step S23: Traverse the various frequencies, rhythms and sequences of the five tones to obtain each five-tone medicine formula. The duration of each five-tone medicine formula is equal to the set time length. For example, by iterating through various frequencies, rhythms, and sequences of the five tones, we can obtain various five-tone medicine recipes. The duration of each five-tone medicine recipe is equal to the set time length, including: the duration of each five-tone medicine recipe is 5 minutes. Step S24: Using the AI ​​sound medicine effect judgment model, based on the set time length, multiple user parameters of the target user, the target user's current five elements attribute data and the target user's current five internal organs attribute data, and the WMA format digital content of each five-tone sound medicine formula, the AI ​​sound medicine effect judgment model will intelligently determine the increase in the target user's happy mood level after using the five-tone sound medicine formula. Specifically, the analog audio signal corresponding to each Wuyin Yinyao formula is converted from analog to digital to obtain the WMA format digital content of each Wuyin Yinyao formula. Step S25: Match the five-tone sound medicine formula with the largest corresponding increase value as the personalized sound medicine formula for the target user; Therefore, by following the steps described above, different personalized sound medicine formulas can be matched for different users. The AI ​​sound drug effect judgment model is a deep neural network that has been trained multiple times, and the number of training times is positively correlated with the number of current voice sub-samples. For example, the positive correlation between the number of training iterations and the number of current speech sub-samples includes: 30 current speech sub-samples correspond to 500 training iterations, 40 current speech sub-samples correspond to 600 training iterations, 50 current speech sub-samples correspond to 700 training iterations, 60 current speech sub-samples correspond to 800 training iterations, and so on. The process of converting each current speech subsample into activation and valence values ​​in a two-dimensional emotional space to serve as the speech emotion feature values ​​of the target user in that current speech subsample includes: the two-dimensional emotional space is a two-dimensional activation-valence space, and the user's emotion is described as a point in the two-dimensional emotional space. Each dimension in the two-dimensional emotional space corresponds to a psychological attribute of the emotion, namely activation or valence. Activation (arousal) represents the intensity of the emotion, and valence represents the positive or negative degree of the emotion. The two-dimensional emotional space is a two-dimensional activation-valence space, in which the user's emotions are described as points in the two-dimensional emotional space. Each dimension of the two-dimensional emotional space corresponds to a psychological attribute of the emotion, namely activation or valence. Activation represents the intensity of the emotion, and valence represents the positive or negative degree of the emotion. It is represented by the following: the user's happy emotion level is positively correlated with the user's activation value and positively correlated with the user's valence value; the user's sad emotion level is negatively correlated with the user's activation value and negatively correlated with the user's valence value. In this way, the activation value and valence value of each user can be used to analyze the level of happiness or sadness of each user in different time intervals. Specifically, the happiness level of the target user corresponding to each voice subsample is determined by the activation value and valence value in the voice emotion feature value of the target user corresponding to each voice subsample. The arithmetic mean of the multiple happiness levels corresponding to multiple current voice subsamples of the target user's current voice sample is taken as the happiness level of the target user corresponding to the current voice sample. Similarly, the level of sadness of the target user corresponding to each voice subsample can be determined by the activation value and valence value in the voice emotion feature value of the target user corresponding to each voice subsample. The arithmetic mean of the multiple sadness levels corresponding to multiple current voice subsamples of the target user's current voice sample is taken as the level of sadness of the target user corresponding to the current voice sample. The process involves traversing various frequencies, rhythms, and sequences of the five tones to obtain each five-tone medicine formula. The duration of each five-tone medicine formula is equal to a set time length, which includes: traversing various frequencies of the five tones refers to traversing the frequency of occurrence of each note in the five tones within the set time length; traversing various sequences of the five tones refers to traversing the display order of each note in the five tones during each performance within the set time length; and traversing various rhythms of the five tones refers to traversing the display rhythm of each note in the five tones during each performance within the set time length. The above three traversals are performed simultaneously to obtain each five-tone medicine formula. The method of using an intelligent emotion recognition model to intelligently identify the target user's current five elements attribute data and the target user's current five internal organs attribute data based on multiple voice emotion feature values ​​corresponding to multiple current voice sub-samples includes: the intelligent emotion recognition model is a BP neural network that has been trained multiple times, and the number of training times is proportional to the set time length; For example, the intelligent emotion recognition model is a BP neural network that has been trained multiple times, and the number of training times is proportional to the set time length, including: a set time length of 5 minutes and 200 training times, a set time length of 6 minutes and 500 training times, a set time length of 7 minutes and 700 training times, and so on. In each training iteration of the deep neural network, the increase in a user's happiness level after using a known Five-Tone Sound Medicine formula is used as the output of the deep neural network. The set time length, multiple user parameters of the user, the user's current Five Elements attribute data and the user's current Five Organs attribute data, and the WMA format digital content of the Five-Tone Sound Medicine formula are used as the input of the deep neural network to complete the training. Specifically, in each training iteration of the deep neural network, the increase in a user's happiness level after using a known Five-Tone Sound Medicine formula is used as the output of the deep neural network. The input of the deep neural network includes a set time length, multiple user parameters of the user, the user's current Five Elements attribute data and the user's current Five Organs attribute data, and the WMA format digital content of the Five-Tone Sound Medicine formula. The training process includes: optionally using the MATLAB toolbox to test and simulate each training iteration of the deep neural network. Among them, the Five Elements attribute data consists of the digital values ​​of the five elements: metal, wood, water, fire, and earth. The value of each attribute ranges from 0 to 1. The closer the value of each attribute is to 0, the further away the user's Five Elements attribute is from that attribute. The Five Organs attribute data consists of the digital values ​​of the five organs: heart, liver, spleen, lungs, and kidneys. The value of each organ attribute ranges from 0 to 1. The closer the value of each organ attribute is to 0, the further away the user's Five Organs attribute is from that organ attribute. Among these, the target user's multiple user parameters include the target user's gender identifier, age information, disease type corresponding to the type code, and the duration of the disease.

[0023] Example 2 Figure 3 The following is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 2 of the present invention.

[0024] like Figure 3 As shown, with Figure 2 Unlike the previous embodiment, in the personalized sound-based medicine formula generation and evaluation method based on voice emotion recognition, after using an AI sound-based medicine effect judgment model to intelligently determine the increase in the target user's happy emotion level after using the five-tone sound-based medicine formula based on a set time length, multiple user parameters of the target user, the target user's current five-element attribute data and the target user's current five-organ attribute data, and the WMA format digital content of each five-tone sound-based medicine formula, that is, after step S24, the method further includes: Step S26: Perform multiple training operations on the deep neural network to obtain a deep neural network after multiple training operations and use it as the output of the AI ​​sound drug effect judgment model; For example, the model representation of the AI ​​sound drug effect judgment model is completed by using the various model parameters of the AI ​​sound drug effect judgment model.

[0025] Example 3 Figure 4 This is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 3 of the present invention.

[0026] like Figure 4 As shown, with Figure 2 Unlike the previous embodiment, in the personalized sound drug formula generation and evaluation method based on voice emotion recognition, after matching the five-tone sound drug formula with the largest corresponding increase value as the personalized sound drug formula for the target user, that is, after step S25, the method further includes: Step S27: Receive the target user's user ID and the target user's personalized sound drug formula, input the target user's user ID and the target user's personalized sound drug formula into the network data packet, and wirelessly transmit the obtained network data packet to the remote sound drug management server through the wireless communication link. For example, receiving the target user's user ID and the target user's personalized sound drug formula, inserting the target user's user ID and the target user's personalized sound drug formula into a network data packet, and wirelessly transmitting the obtained network data packet to a remote sound drug management server via a wireless communication link includes: receiving the target user's user ID and the target user's personalized sound drug formula, inserting the target user's user ID and the target user's personalized sound drug formula into an IP data packet, and wirelessly transmitting the obtained network data packet to a remote sound drug management server via a wireless communication link.

[0027] Example 4 Figure 5 This is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 4 of the present invention.

[0028] like Figure 5 As shown, with Figure 4 Unlike the previous embodiment, in the personalized sound drug formula generation and evaluation method based on voice emotion recognition, after matching the five-tone sound drug formula with the largest corresponding increase value as the personalized sound drug formula for the target user, that is, after step S25, the method further includes: Step S28: Receive the target user's user ID and the target user's personalized sound drug formula, and display the target user's user ID and each formula parameter of the target user's personalized sound drug formula simultaneously; For example, an LCD display array can be used to receive the user ID of the target user and the personalized sound drug formula of the target user, and to display the user ID of the target user and the formula parameters of each set of personalized sound drug formula of the target user simultaneously. This includes receiving the target user's user ID and the target user's personalized sound drug formula, and simultaneously displaying the target user's user ID and each formula parameter of the target user's personalized sound drug formula, including the frequency parameters, rhythm parameters, and sequence parameters of the five tones of the target user's personalized sound drug formula.

[0029] Example 5 Figure 6 This is a flowchart illustrating the steps of a personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition according to Embodiment 5 of the present invention.

[0030] like Figure 6 As shown, with Figure 2 Unlike the previous embodiment, the personalized voice-based drug formulation generation and evaluation method based on voice emotion recognition, after collecting voice samples of a set time length from the target user as the target user's current voice sample, uniformly segmenting the target user's current voice sample to obtain multiple current voice sub-samples, and converting each current voice sub-sample into activation and valence values ​​in a two-dimensional emotion space as the target user's voice emotion feature values ​​in that current voice sub-sample, i.e., after step S21, the method further includes: Step S29: Perform multiple training operations on the BP neural network to obtain a BP neural network after multiple training operations, and output the BP neural network after multiple training operations as the intelligent emotion recognition model. For example, a programmable logic device can be selected to perform multiple training operations on the BP neural network to obtain a BP neural network after multiple training operations, and the BP neural network after multiple training operations can be output as an intelligent emotion recognition model.

[0031] Next, the various method embodiments of the present invention will be described in detail.

[0032] In a personalized sound-based drug formulation generation and evaluation method based on voice emotion recognition according to various embodiments of the present invention: Traversing the five tones at various frequencies refers to traversing the frequency of occurrence of each note in the five tones within a set time period. Traversing the five tones in various orders refers to traversing the order in which each note in the five tones is displayed in each performance within a set time period. Traversing the five tones in various rhythms refers to traversing the rhythm in which each note in the five tones is displayed in each performance within a set time period. The above three traversals are performed simultaneously to obtain each five-tone sound medicine formula, which includes the five tones being the five notes: Gong, Shang, Jiao, Zhi, and Yu. It can be seen that among the various frequencies, sequences, and rhythms of the five tones, as long as one of the traversal methods is different, different formulas for the five-tone medicine can be obtained. Generally, in the specific traversal process, we can first determine the frequency of the five notes, that is, the number of times each note of the five notes appears in the current five-note medicine formula and the duration of each note when it appears. Then we can determine the order of each note of the five notes when it appears each time, that is, the specific relative position, and then determine the specific rhythm used by each note of the five notes when it appears each time, that is, the specific beat. The AI ​​sound drug effect judgment model is a deep neural network that has been trained multiple times, and the number of training times is positively correlated with the number of current speech sub-samples, including: using a content conversion formula to represent the content conversion relationship that is positively correlated with the number of training times and the number of current speech sub-samples; The content conversion formula, which expresses the positive correlation between the number of training iterations and the number of current speech sub-samples, includes the following: In the content conversion formula, the content conversion formula is the input content, and the number of training iterations is the output content.

[0033] In a personalized sound-based drug formulation generation and evaluation method based on voice emotion recognition according to various embodiments of the present invention: The AI-powered sound-based medicine effect assessment model intelligently determines the increase in the target user's happiness level after using a specific sound-based medicine formula based on a set time duration, multiple user parameters of the target user, the target user's current Five Elements attribute data, the target user's current Five Organs attribute data, and the WMA format digital content of each Five Elements sound-based medicine formula. This includes performing binary value conversion processing on the set time duration, multiple user parameters of the target user, the target user's current Five Elements attribute data, the target user's current Five Organs attribute data, and the WMA format digital content of each Five Elements sound-based medicine formula before synchronously inputting them into the AI ​​sound-based medicine effect assessment model. For example, an ASIC chip can be used to implement binary value conversion and synchronous input; The AI-powered sound medicine effect judgment model intelligently determines the increase in the target user's happiness level after using the five-tone sound medicine formula based on the set time length, multiple user parameters of the target user, the target user's current five-element attribute data and the target user's current five-organ attribute data, and the WMA format digital content of each five-tone sound medicine formula. It also includes running the AI-powered sound medicine effect judgment model to obtain the increase in the target user's happiness level after using the five-tone sound medicine formula, as output by the AI-powered sound medicine effect judgment model.

[0034] And in a personalized sound-based drug formulation generation and evaluation method based on voice emotion recognition according to various embodiments of the present invention: Matching the five-tone sound medicine formula with the largest corresponding increase value as the personalized sound medicine formula for the target user includes: obtaining the increase values ​​corresponding to each five-tone sound medicine formula, and using the five-tone sound medicine formula corresponding to the largest increase value among the increase values ​​as the personalized sound medicine formula for the target user. For example, a programmable logic device can be selected to implement the logic process of obtaining the amplification values ​​corresponding to each of the five-tone sound medicine formulas, and using the five-tone sound medicine formula corresponding to the amplification value with the largest value among the amplification values ​​as the personalized sound medicine formula for the target user. Among them, matching the five-tone sound medicine formula with the largest corresponding increase value as the target user's personalized sound medicine formula also includes: the increase value corresponding to each five-tone sound medicine formula is positive, negative or zero.

[0035] In addition, the present invention may also refer to the following technical contents to highlight the significant technical progress of the present invention: The AI ​​sound drug effect judgment model is a deep neural network that has been trained multiple times, and the number of training times is positively correlated with the number of current speech sub-samples. The deep neural network includes multiple hidden layers, and the number of hidden layers is positively correlated with a set time length. The AI ​​sound drug effect judgment model is a deep neural network that has been trained multiple times, and the number of training times is positively correlated with the number of current speech sub-samples. It also includes the following: the deep neural network includes multiple hidden layers, a single input layer and a single output layer, with the multiple hidden layers located between the single input layer and the single output layer. For example, the deep neural network includes multiple hidden layers, and the number of hidden layers is positively correlated with a set time length, which includes: using a numerical mapping function to represent the numerical mapping relationship between the number of hidden layers and the set time length, where the set time length is the input parameter of the numerical mapping function and the number of hidden layers is the output parameter of the numerical mapping function.

[0036] While only exemplary embodiments of the invention have been described in detail above, those skilled in the art will immediately understand that many variations can be made in these exemplary embodiments without substantially departing from the novel principles and advantages of the invention. Accordingly, all such variations are to be included within the scope of the invention as defined in the following claims. In the claims, the means and function clauses are intended to include the structures described herein as implementing the stated functions, and include not only structural equivalents but also equivalent structures.

Claims

1. A method for generating and evaluating personalized voice-based drug prescriptions based on voice emotion recognition, characterized in that, The method includes: Collect voice samples of a set time length from the target user to serve as the target user's current voice sample. Perform uniform segmentation on the target user's current voice sample to obtain multiple current voice sub-samples. Convert each current voice sub-sample into activation value and valence value in a two-dimensional emotion space to serve as the target user's voice emotion feature value in that current voice sub-sample. An intelligent emotion recognition model is used to intelligently identify the target user's current Five Elements attribute data and the target user's current Five Organs attribute data based on the numerical values ​​of multiple voice emotion features corresponding to multiple current voice sub-samples. By traversing various frequencies, rhythms, and sequences of the five tones, we can obtain various five-tone medicine recipes. The duration of each five-tone medicine recipe is equal to the set time length. The AI-powered sound medicine effect judgment model uses a set time length, multiple user parameters of the target user, the target user's current five elements attribute data and the target user's current five internal organs attribute data, and the WMA format digital content of each five-tone sound medicine formula to intelligently judge the increase in the target user's happy mood level after using the five-tone sound medicine formula. The five-tone sound medicine formula with the largest corresponding increase value will be matched to the target user as the personalized sound medicine formula; The AI-powered sound drug effect judgment model is a deep neural network that has been trained multiple times, and the number of training sessions is positively correlated with the number of current speech sub-samples.

2. The method for generating and evaluating personalized sound-based drug prescriptions based on voice emotion recognition as described in claim 1, characterized in that: Converting each current speech subsample into activation and valence values ​​in a two-dimensional emotion space to serve as the speech emotion feature values ​​of the target user in that current speech subsample includes: the two-dimensional emotion space is a two-dimensional activation-valence space, and the user's emotion is described as a point in the two-dimensional emotion space. Each dimension in the two-dimensional emotion space corresponds to a psychological attribute of the emotion, namely activation or valence. Activation represents the intensity of the emotion, and valence represents the positive or negative degree of the emotion. The two-dimensional emotional space is a two-dimensional activation-valence space, in which the user's emotions are described as points in the two-dimensional emotional space. Each dimension of the two-dimensional emotional space corresponds to a psychological attribute of the emotion, namely activation or valence. Activation represents the intensity of the emotion, and valence represents the positive or negative degree of the emotion. It is represented by the following: the user's happy emotion level is positively correlated with the user's activation value and positively correlated with the user's valence value; the user's sad emotion level is negatively correlated with the user's activation value and negatively correlated with the user's valence value. Specifically, the happiness level of the target user corresponding to each voice subsample is determined by the activation value and valence value in the voice emotion feature value of the target user corresponding to each voice subsample. The arithmetic mean of the multiple happiness levels corresponding to multiple current voice subsamples of the target user's current voice sample is taken as the happiness level of the target user corresponding to the current voice sample. The process involves traversing various frequencies, rhythms, and sequences of the five tones to obtain each five-tone medicine formula. The duration of each five-tone medicine formula is equal to a set time length, which includes: traversing various frequencies of the five tones (the frequency of each note in the five tones within the set time length), traversing various sequences of the five tones (the order in which each note in the five tones is displayed in each performance within the set time length), and traversing various rhythms of the five tones (the rhythm in which each note in the five tones is displayed in each performance within the set time length). The above three traversals are performed simultaneously to obtain each five-tone medicine formula.

3. The method for generating and evaluating personalized sound-based drug prescriptions based on voice emotion recognition as described in claim 2, characterized in that: The intelligent emotion recognition model intelligently identifies the target user's current five elements attribute data and the target user's current five internal organs attribute data based on multiple voice emotion feature values ​​corresponding to multiple current voice sub-samples. The intelligent emotion recognition model is a BP neural network that has been trained multiple times, and the number of training times is proportional to the set time length. In each training iteration of the deep neural network, the increase in a user's happiness level after using a known Five-Tone Sound Medicine formula is used as the output of the deep neural network. The set time length, multiple user parameters of the user, the user's current Five Elements attribute data and the user's current Five Organs attribute data, and the WMA format digital content of the Five-Tone Sound Medicine formula are used as the input of the deep neural network to complete the training. Among them, the Five Elements attribute data consists of the digital values ​​of the five elements: metal, wood, water, fire, and earth. The value of each attribute ranges from 0 to 1. The closer the value of each attribute is to 0, the further away the user's Five Elements attribute is from that attribute. The Five Organs attribute data consists of the digital values ​​of the five organs: heart, liver, spleen, lungs, and kidneys. The value of each organ attribute ranges from 0 to 1. The closer the value of each organ attribute is to 0, the further away the user's Five Organs attribute is from that organ attribute. Among them, the target user's multiple user parameters include the target user's gender identifier, age information, disease type corresponding type code, and disease duration.

4. The method for generating and evaluating personalized voice-based drug prescriptions based on voice emotion recognition as described in claim 3, characterized in that, After employing an AI-powered sound-based medicine effect assessment model to intelligently determine the increase in the target user's level of happiness after using a given sound-based medicine formula, based on a set time duration, multiple user parameters of the target user, the target user's current Five Elements attribute data, the target user's current Five Organs attribute data, and the WMA format digital content of each Five-Tone Sound-Based Medicine Formula, the method further includes: The deep neural network is trained multiple times to obtain a deep neural network after multiple training sessions, which is then used as the output of the AI ​​sound drug effect judgment model.

5. The method for generating and evaluating personalized voice-based drug prescriptions based on voice emotion recognition as described in claim 3, characterized in that, After matching the five-tone sound medicine formula with the largest corresponding increase value as the personalized sound medicine formula for the target user, the method further includes: The system receives the target user's user ID and personalized sound drug formula, inserts the user ID and personalized sound drug formula into a network data packet, and wirelessly transmits the obtained network data packet to a remote sound drug management server via a wireless communication link.

6. The method for generating and evaluating personalized sound-based drug prescriptions based on voice emotion recognition as described in claim 3, characterized in that, After matching the five-tone sound medicine formula with the largest corresponding increase value as the personalized sound medicine formula for the target user, the method further includes: Receive the target user's user ID and the target user's personalized sound medicine formula, and display the target user's user ID and each formula parameter of the target user's personalized sound medicine formula simultaneously; This includes receiving the target user's user ID and the target user's personalized sound drug formula, and simultaneously displaying the target user's user ID and each formula parameter of the target user's personalized sound drug formula, including the frequency parameters, rhythm parameters, and sequence parameters of the five tones of the target user's personalized sound drug formula.

7. The method for generating and evaluating personalized sound-based drug prescriptions based on voice emotion recognition as described in claim 3, characterized in that, After collecting voice samples of a set time length from the target user as the target user's current voice sample, uniformly segmenting the target user's current voice sample to obtain multiple current voice sub-samples, and converting each current voice sub-sample into activation and valence values ​​in a two-dimensional emotion space as the target user's voice emotion feature values ​​in that current voice sub-sample, the method further includes: The BP neural network is trained multiple times to obtain a BP neural network after multiple training sessions, and the BP neural network after multiple training sessions is used as the output of the intelligent emotion recognition model.

8. The method for generating and evaluating personalized sound-based drug prescriptions based on voice emotion recognition as described in any one of claims 3-7, characterized in that: Traversing the five tones at various frequencies refers to traversing the frequency of occurrence of each note in the five tones within a set time period. Traversing the five tones in various orders refers to traversing the order in which each note in the five tones is displayed in each performance within a set time period. Traversing the five tones in various rhythms refers to traversing the rhythm in which each note in the five tones is displayed in each performance within a set time period. The above three traversals are performed simultaneously to obtain each five-tone sound medicine formula, which includes the five tones being the five notes: Gong, Shang, Jiao, Zhi, and Yu. The AI ​​sound drug effect judgment model is a deep neural network that has been trained multiple times, and the number of training times is positively correlated with the number of current speech sub-samples, including: using a content conversion formula to represent the content conversion relationship that is positively correlated with the number of training times and the number of current speech sub-samples; The content conversion formula, which expresses the positive correlation between the number of training iterations and the number of current speech sub-samples, includes the following: In the content conversion formula, the content conversion formula is the input content, and the number of training iterations is the output content.

9. The method for generating and evaluating personalized sound-based drug prescriptions based on voice emotion recognition as described in any one of claims 3-7, characterized in that: The AI-powered sound-based medicine effect assessment model intelligently determines the increase in the target user's happiness level after using a specific sound-based medicine formula based on a set time duration, multiple user parameters of the target user, the target user's current Five Elements attribute data, the target user's current Five Organs attribute data, and the WMA format digital content of each Five Elements sound-based medicine formula. This includes performing binary value conversion processing on the set time duration, multiple user parameters of the target user, the target user's current Five Elements attribute data, the target user's current Five Organs attribute data, and the WMA format digital content of each Five Elements sound-based medicine formula before synchronously inputting them into the AI ​​sound-based medicine effect assessment model. The AI-powered sound medicine effect judgment model intelligently determines the increase in the target user's happiness level after using the five-tone sound medicine formula based on the set time length, multiple user parameters of the target user, the target user's current five-element attribute data and the target user's current five-organ attribute data, and the WMA format digital content of each five-tone sound medicine formula. It also includes running the AI-powered sound medicine effect judgment model to obtain the increase in the target user's happiness level after using the five-tone sound medicine formula, as output by the AI-powered sound medicine effect judgment model.

10. The method for generating and evaluating personalized sound-based drug prescriptions based on voice emotion recognition as described in any one of claims 3-7, characterized in that: Matching the five-tone sound medicine formula with the largest corresponding increase value as the personalized sound medicine formula for the target user includes: obtaining the increase values ​​corresponding to each five-tone sound medicine formula, and using the five-tone sound medicine formula corresponding to the largest increase value among the increase values ​​as the personalized sound medicine formula for the target user. Among them, matching the five-tone sound medicine formula with the largest corresponding increase value as the target user's personalized sound medicine formula also includes: the increase value corresponding to each five-tone sound medicine formula is positive, negative or zero.

Citation Information

Patent Citations

  • Digital music therapy instrument

    CN102188773A

  • Music command and music treatment system

    CN111870877A