Mental health interaction method and system based on holographic projection

By collecting and analyzing voiceprint, micro-expression, and heart rate information, a multimodal emotion fusion model is constructed, realizing the functional integration and dynamic feedback of traditional mental health interactive devices, improving user experience and privacy protection, and solving the functional fragmentation and privacy issues of traditional devices.

CN121331384APending Publication Date: 2026-01-13JIANGSU ZHUODUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511422990.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Traditional interactive devices for mental health suffer from problems such as fragmented functional modules, insufficient interactive experience, lack of dynamic feedback, and weak privacy protection.

Method used

Voiceprint, micro-expression, and heart rate information are collected through microphone arrays, infrared cameras, and heart rate sensors. Feature extraction and classification are performed using MFCC, GRU, CNN, and LSTM to construct a multimodal emotion fusion model. Weighted fusion is performed based on the attention mechanism to output the user's stress index. Dynamic art therapy or shouting catharsis therapy is then performed. Privacy protection is achieved by combining laser holography and blockchain hash storage.

Benefits of technology

It achieves closed-loop analysis of user psychological state, improves interactive experience and privacy protection, increases the efficiency of function integration, enhances the accuracy of dynamic feedback and privacy security, supports multiple emotional modal recognition and real-time healing content linkage, and reduces the risk of voiceprint leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121331384A_ABST
    Figure CN121331384A_ABST
Patent Text Reader

Abstract

The invention provides a psychological health interaction method and system based on holographic projection, and relates to the technical field of psychological health, and the method comprises the steps: respectively collecting voiceprint information, micro-expression information and heart rate information of a user side through a microphone array, an infrared camera and a heart rate sensor; performing MFCC feature extraction and GRU neural network classification on the voiceprint information to generate voiceprint features, extracting micro-expression features based on CNN, and extracting heart rate variability features based on LSTM; constructing a multi-modal emotion fusion model, and performing weighted fusion on all features based on an attention mechanism to output a user pressure index; an anxiety score is obtained based on the user pressure index, a first anxiety score threshold value is set, if the anxiety score is smaller than or equal to the first anxiety score threshold value, dynamic art healing is conducted on the user side, otherwise, yelling catharsis healing is conducted, and therefore the device function integration efficiency can be improved, the user interaction experience can be improved, and the dynamic feedback precision can be ensured; and the privacy protection intensity is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of mental health services, and particularly relates to a mental health interaction method and system based on holographic projection. BACKGROUND

[0002] In the field of mental health services, intelligent holographic interaction devices, namely mental holographic cabins, have been gradually applied to psychological assessment, art therapy, emotional release and other scenes as emerging intervention tools.

[0003] In the prior art, the traditional mental health interaction device has significant technical defects. First, the functional modules are in a fragmented state, and the core functions of psychological assessment, art therapy, and emotional release are independently run, and the data generated by each module cannot be interchanged, making it difficult to form a closed-loop analysis of the user's psychological state. Second, the interaction experience is insufficient, the voice wake-up function only supports fixed instruction recognition, and cannot adaptively recognize emotional expressions such as calls with crying tones, reducing user convenience and interaction continuity. Third, dynamic feedback is missing, and the art therapy content is static and pre-set, and cannot dynamically adjust the dance movements, picture rhythm and other therapy content parameters according to the user's real-time heart rate, micro-expression and other physiological and emotional signals, affecting the effect of art therapy. Finally, privacy protection is weak, and the original audio data generated by the emotional release is stored in plaintext form, which has the risk of voiceprint information leakage, and cannot guarantee the privacy and security of the user.

[0004] Therefore, it is necessary to provide a mental health interaction method and system based on holographic projection to solve the above technical problems. SUMMARY

[0005] To solve the above technical problems, the present application provides a mental health interaction method and system based on holographic projection, which is used to solve the problems of functional fragmentation, insufficient interaction experience, missing dynamic feedback and weak privacy protection of traditional mental health interaction devices.

[0006] The mental health interaction method based on holographic projection provided by the present application comprises: The voiceprint information, micro-expression information and heart rate information of the user end are collected by a microphone array, an infrared camera and a heart rate sensor respectively; The voiceprint information is subjected to MFCC feature extraction and GRU neural network classification to generate voiceprint features, the micro-expression features in the micro-expression information are extracted based on CNN, and the heart rate variability features corresponding to the heart rate information are extracted based on LSTM; A multi-modal emotion fusion model is constructed, the voiceprint features, the micro-expression features and the heart rate variability features are weighted and fused based on an attention mechanism, and a user stress index is outputted; Anxiety score reflecting the user psychological assessment result is obtained based on the user stress index, a first anxiety score threshold is set, if the anxiety score is less than or equal to the first anxiety score threshold, dynamic art healing is performed on the user terminal, and if the anxiety score is greater than the first anxiety score threshold, shouting catharsis healing is performed on the user terminal.

[0007] Preferably, the MFCC feature extraction and GRU neural network classification are performed on the voiceprint information to generate voiceprint features, the micro-expression features in the micro-expression information are extracted based on CNN, and the heart rate variability features corresponding to the heart rate information are extracted based on LSTM, and specifically comprising: The voiceprint information is segmented into continuous audio frames according to a preset segmentation parameter, and background noise and invalid signals are filtered to obtain valid voiceprint segments; The MFCC algorithm is used to perform Fourier transform, Mel filter bank mapping and cepstrum calculation on the valid voiceprint segments in sequence to extract MFCC features; The MFCC features are input into the input layer of the GRU neural network; the reset gate and update gate mechanisms of the 128-dimensional hidden layer of the GRU neural network capture the time sequence dependence of the MFCC features, filter invalid noise features and strengthen key voiceprint attributes to generate initial voiceprint features; the classification layer of the GRU neural network classifies the initial voiceprint features through a trained classification model to classify emotion types and language features, and maps the high-dimensional initial voiceprint features to low-dimensional voiceprint features containing emotion labels and language labels; The face dynamic features in the micro-expression information are extracted based on CNN, at least including frown frequency and mouth corner radian, and the micro-expression features are output; The heart rate information is time series modeled based on LSTM to capture the fluctuation law of the heart rate signal to obtain the heart rate variability features.

[0008] Preferably, before constructing the multi-modal emotion fusion model, weighting and fusing the voiceprint features, the micro-expression features and the heart rate variability features based on the attention mechanism, and outputting the user stress index, the multi-modal emotion fusion trigger threshold is dynamically adjusted according to the emotion label of the voiceprint features, when the emotion label is excited, the adjusted multi-modal emotion fusion trigger threshold is 70% of the initial multi-modal emotion fusion trigger threshold; when the emotion label is calm, the multi-modal emotion fusion trigger threshold remains unchanged.

[0009] Preferably, the multi-modal emotion fusion model is constructed, the voiceprint features, the micro-expression features and the heart rate variability features are weighted and fused based on the attention mechanism, and the user stress index is output, and the corresponding calculation formula is as follows: YHYL = SW * 0.6 + WBQ * a + XLBYX * b a + b = 0.4 In the formula, YHYL represents a user stress index; SW represents a voiceprint feature; WBQ represents a micro-expression feature; XLBYX represents a heart rate variability feature; a represents a weight coefficient corresponding to the micro-expression feature; and b represents a weight coefficient corresponding to the heart rate variability feature.

[0010] Preferably, the dynamic art therapy for the user terminal specifically includes: setting a second anxiety score threshold, and if the anxiety score is less than or equal to the second anxiety score threshold, directly matching a dance therapy mode for the user terminal; if the anxiety score is greater than the second anxiety score threshold, preferentially matching a breathing training guide mode for the user terminal, and after the breathing training guide mode ends, matching the dance therapy mode for the user terminal.

[0011] Preferably, the matching process of the dance therapy mode is as follows: adopting a kinematics inverse solution algorithm, retrieving and matching a dance therapy action from a 3D action library based on the user stress index, and the action parameters of the dance therapy action including action speed, action amplitude and music rhythm; wherein the user stress index and the action parameters of the dance therapy action satisfy the following mapping relationship: f(YHYL) = 0.5 * DZSD + 0.3 * DZFD + 0.2 * YYJZ In the formula, f represents a mapping function between the user stress index and the action parameters of the dance therapy action; DZSD represents the action speed of the dance therapy action; DZFD represents the action amplitude of the dance therapy action; and YYJZ represents the music rhythm of the dance therapy action. adopting a laser holographic technology, generating a holographic projection picture based on the matched dance therapy action and displaying the holographic projection picture to the user terminal; capturing action capture data of the user terminal through a camera, and dynamically adjusting the action speed based on the action capture data according to an OpenPose algorithm.

[0012] Preferably, the scream catharsis therapy for the user terminal specifically includes: capturing a scream catharsis audio of the user terminal; performing Fourier transform processing on the scream catharsis audio, i.e., converting time domain information in the scream catharsis audio into frequency domain information, performing voiceprint desensitization processing on the scream catharsis audio through a frequency domain desensitization feature mask generation algorithm, and filtering out a key voiceprint frequency band of 300 to 3400 Hz in the scream catharsis audio; The scream relief audio after the voiceprint desensitization processing is stored by AES-256 encryption, and the generation process of the encryption key needs to embed the data hash value of the scream relief audio; The data hash value of the scream relief audio is block chain hash notarized at a fixed interval of 10 seconds per block.

[0013] Preferably, after the user end is subjected to scream relief healing, the average sound intensity of the scream relief audio is obtained; When the average sound intensity is less than or equal to 90 decibels, a music relaxation list is automatically pushed, wherein alpha wave music is matched according to the voiceprint features and the music relaxation list is formed; When the average sound intensity is greater than 90 decibels, psychological counseling AI is automatically triggered to perform psychological counseling on the user end.

[0014] A holographic projection-based mental health interaction system, the system comprising: An information acquisition module for acquiring voiceprint information, micro-expression information and heart rate information of a user end through a microphone array, an infrared camera and a heart rate sensor respectively; A feature extraction module for performing MFCC feature extraction and GRU neural network classification on the voiceprint information to generate voiceprint features, extracting micro-expression features in the micro-expression information based on CNN, and extracting heart rate variability features corresponding to the heart rate information based on LSTM; A multi-modal fusion module for constructing a multi-modal emotion fusion model, performing weighted fusion on the voiceprint features, the micro-expression features and the heart rate variability features based on an attention mechanism, and outputting a user stress index; A user healing module for obtaining an anxiety score reflecting a user psychological evaluation result based on the user stress index, setting a first anxiety score threshold, and if the anxiety score is less than or equal to the first anxiety score threshold, performing dynamic art healing on the user end, and if the anxiety score is greater than the first anxiety score threshold, performing scream relief healing on the user end.

[0015] Compared with the related art, the holographic projection-based mental health interaction method and system provided by the present application has the following beneficial effects: The application collects voiceprint information, micro-expression information and heart rate information through a microphone array, an infrared camera and a heart rate sensor respectively; performs MFCC (Mel Frequency Cepstral Coefficients) feature extraction and GRU (Gate Recurrent Unit) neural network classification on the voiceprint information to generate voiceprint features, extracts micro-expression features in the micro-expression information based on a CNN (Convolutional Neural Network), and extracts heart rate variability features corresponding to the heart rate information based on an LSTM (Long Short-Term Memory); a multi-modal emotion fusion model is constructed, the voiceprint features, the micro-expression features and the heart rate variability features are weighted and fused based on an attention mechanism, and a user stress index is output; an anxiety score reflecting a user psychological evaluation result is obtained based on the user stress index, a first anxiety score threshold is set, if the anxiety score is less than or equal to the first anxiety score threshold, dynamic art healing is performed on the user end, if the anxiety score is greater than the first anxiety score threshold, shouting catharsis healing is performed on the user end, so that the device function integration efficiency can be improved, the user interaction experience can be improved, the dynamic feedback accuracy can be ensured, the privacy protection strength can be enhanced, and the problems of function fragmentation, insufficient interaction experience, missing dynamic feedback and weak privacy protection of traditional mental health interaction devices can be solved.

[0016] The present application can break the functional fragmentation of traditional mental health interaction devices through multi-module cooperation, increase the data intercommunication rate from 30% of traditional devices to 98%, realize closed-loop analysis of user psychological state, and shorten the process cycle of "evaluation-intervention-feedback" to within 5 minutes; the present application can perform MFCC feature extraction and GRU neural network classification on voiceprint information to generate voiceprint features, and dynamically adjust the multi-modal emotion fusion trigger threshold, break through the bottleneck of traditional voice wake-up supporting only fixed instructions, increase the success rate of emotional voice wake-up from 60% to 95%, and support the recognition of more than 20 emotional modalities such as crying tone and trembling sound, support adaptive recognition of Sichuan Mandarin, Cantonese and other dialectal variants, and greatly improve the convenience and adaptability of user interaction; the present application can obtain a user stress index based on attention mechanism weighted fusion of voiceprint, micro-expression and heart rate variability features, retrieve matching dance healing actions from a 3D action library through kinematics inverse solution algorithm, dynamically adjust the action speed according to the motion capture data relying on the OpenPose algorithm, so that the dance healing action retrieval delay is less than 10 milliseconds, the adjustment delay is less than 80 milliseconds, and the matching degree of heart rate and dance rhythm is 0.85, realizing real-time linkage of healing content and user emotions, and improving the effect of art healing; the present application performs Fourier transform, voiceprint feature desensitization and AES-256 encryption storage, and cooperates with blockchain hash evidence, so that the information entropy of desensitized voiceprint data is reduced by 70%, effectively avoiding the risk of voiceprint leakage, and protecting the privacy and security of users. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 A flowchart of a holographic projection-based mental health interaction method is provided for the embodiments of the present application. Figure 2 A flowchart of triggering multi-modal emotion fusion is provided for the embodiments of the present application. Figure 3 A flowchart of shout relief audio privacy protection is provided for the embodiments of the present application. Figure 4 A system block diagram of a holographic projection-based mental health interaction system is provided for the embodiments of the present application. Figure 5 A hardware structure schematic diagram of an electronic device is provided for the embodiments of the present application. DETAILED DESCRIPTION

[0018] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0019] As Figure 1 shown is a flowchart of a holographic projection-based mental health interaction method provided by an embodiment of the present application, Figure 1 The execution subject of the method shown can be a software and / or hardware device. The execution subject of the present application can include but is not limited to at least one of the following: a user device, a network device, etc. Among them, the user device can include but is not limited to a computer, a smart phone, a personal digital assistant (Personal Digital Assistant, PDA for short) and the above-mentioned electronic devices, etc. The network device can include but is not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing, wherein cloud computing is a kind of distributed computing, which is a super virtual computer composed of a group of loosely coupled computers. The present embodiment does not make any limitation. Including steps S1 to S4, as follows: S1, collecting voiceprint information, micro-expression information and heart rate information of the user end through a microphone array, an infrared camera and a heart rate sensor respectively; Among them, the microphone array is a multi-channel array structure supporting 360° sound source positioning, which can realize full-space coverage positioning of the user's sound source through the cooperative work of multiple microphone units, and accurately collect the voiceprint information of the user in the scene of shouting and venting, etc., to ensure the all-around and accuracy of voiceprint information collection. The infrared camera has high frame rate image capture capability, which can clearly capture the micro-expression information formed by the subtle muscle movements of the user's face, covering the subtle lifting or drooping of the corners of the mouth, the frequency of frowning and other expressions that are not easily detected by the naked eye, providing data support for subsequent construction of multi-modal emotion fusion model. The heart rate sensor provides two collection modes for selection, i.e. contact mode and non-contact mode. The contact mode can realize real-time and high-precision collection of heart rate signals by directly contacting the user's skin, such as wrist and fingertip contact. The non-contact mode indirectly obtains heart rate information by measuring the reflection or absorption degree of light emitted by the LED light source through the user's skin. The two modes adapt to different application scenarios, ensuring the flexibility and applicability of heart rate information collection.

[0020] In the above manner, the physiological and behavioral information of the user, including voiceprint, micro-expression and heart rate information, can be synchronously collected by the multi-modal sensor group, laying a foundation for subsequent multi-dimensional data fusion analysis.

[0021] S2, performing MFCC feature extraction and GRU neural network classification on the voiceprint information to generate voiceprint features, extracting micro-expression features in the micro-expression information based on CNN, and extracting heart rate variability features corresponding to the heart rate information based on LSTM; The MFCC feature extraction method conforms to the design of human ear hearing characteristics, which can effectively capture the spectral features closely related to the voiceprint characteristics in the speech signal, including the frequency distribution, energy change and other key information of the speech, by converting the time domain signal of the voiceprint information into mel frequency domain cepstral coefficient, providing a basis for subsequent accurate classification of voiceprint features. GRU is a gated recurrent unit, which is an improved recurrent neural network structure. It can adaptively capture the time sequence dependence in the voiceprint information by setting the reset gate and update gate, and effectively classify and learn the features extracted by MFCC to generate voiceprint features with clear category attributes, effectively solving the gradient disappearance or gradient explosion problem of traditional recurrent neural network when processing long sequence data. CNN is a convolutional neural network, which is a deep learning model with local receptive field and weight sharing characteristics. Through the synergistic effect of convolutional layer and pooling layer structure, it can automatically extract the local features of the key facial regions in the micro-expression information, and realize the transformation from low-level features to high-level semantic features through multi-layer feature mapping. LSTM is a long short-term memory network, which is a variant of recurrent neural network specially used for processing time series data. Through the gating mechanism of input gate, forget gate and output gate, it can effectively remember the key features of long time sequence in heart rate information, while filtering irrelevant noise, and accurately extract the heart rate variability features reflecting the heart rate fluctuation law from the heart rate signal. This feature can effectively represent the activity state of the autonomic nervous system, providing an important basis for emotion analysis and other applications.

[0022] The MFCC feature extraction and GRU neural network classification of the voiceprint information generate voiceprint features, the micro-expression features in the micro-expression information are extracted based on CNN, and the heart rate variability features corresponding to the heart rate information are extracted based on LSTM. Specifically, it includes: The voiceprint information is segmented into continuous audio frames according to the preset segmentation parameters, and background noise and invalid signals are filtered out to obtain valid voiceprint segments; The MFCC algorithm is used to perform Fourier transform, mel filter bank mapping and cepstrum calculation on the valid voiceprint segments in turn to extract MFCC features; The MFCC features are input into the input layer of the GRU neural network; the reset gate and update gate mechanism of the 128-dimensional hidden layer of the GRU neural network capture the time sequence dependence of the MFCC features, filter out invalid noise features and strengthen key voiceprint attributes to generate initial voiceprint features; the classification layer of the GRU neural network classifies the initial voiceprint features into emotion types and language features through a trained classification model, and maps the high-dimensional initial voiceprint features into low-dimensional voiceprint features containing emotion labels and language labels; The face dynamic features in the micro-expression information are extracted based on CNN, including at least frown frequency and mouth corner arc, and the micro-expression features are output. The heart rate information is time series modeled based on the LSTM, and fluctuation rules of the heart rate signal are captured to obtain the heart rate variability feature.

[0023] It can be understood that the setting of the preset segmentation parameter needs to be combined with the acquisition scene of the voiceprint information and the speech signal characteristics to ensure that the continuous audio frames obtained by segmentation can not only completely retain the emotion and language features in the voiceprint, but also avoid frame information redundancy or loss. When filtering out background noise and invalid signals, an improved spectral subtraction combined with a speech activity detection algorithm is used to accurately identify and retain the valid voiceprint segments containing speech components, and to eliminate environmental noise and silent segment interference, thereby ensuring the accuracy of subsequent voiceprint feature extraction.

[0024] In the MFCC algorithm extraction process, Fourier transform is used to convert the time domain signal of the valid voiceprint segment into a frequency domain signal, clearly presenting the frequency composition characteristics of the voiceprint. The Mel filter bank mapping is based on the human ear's perception rules of different frequency sounds, and filters the frequency domain signal to highlight the key frequency components related to voiceprint recognition. The cepstrum calculation extracts the MFCC features that can represent the uniqueness of the voiceprint by performing mathematical transformation on the filtered frequency domain signal, providing high-quality feature input for subsequent GRU neural network classification.

[0025] In the 128-dimensional hidden layer of the GRU neural network, the reset gate is used to control the retention weight of historical information, and the update gate is used to adjust the fusion ratio of current input and historical state, so as to effectively capture the long-time dependence relationship in the MFCC feature, filter out invalid noise features and strengthen key voiceprint attributes, generate initial voiceprint features, and the recognition accuracy of the GRU neural network reaches 92%. The classification layer of the GRU neural network classifies the initial voiceprint features through the trained classification model to output the probability distribution of emotion types and language features, such as calm, excited, etc. The language features include dialect variants such as Sichuan Mandarin and Cantonese, so as to map the high-dimensional initial voiceprint features to low-dimensional voiceprint features containing emotion labels and language labels. For new language adaptation of language features, the training parameters of the GRU neural network are transferred to the new language dataset through transfer learning, and only the output layer parameters of the network need to be fine-tuned to quickly adapt, reducing the sample size and time cost of retraining.

[0026] When extracting micro-expression features based on CNN, the micro-expression information is first preprocessed to locate key facial regions, such as eyebrows and lips. Then, through CNN operations such as convolution and pooling, facial dynamic features are extracted layer by layer, including at least the frequency of frowning and the curvature of the corners of the mouth. Among them, the frequency of frowning reflects the movement frequency of the eyebrow region, and the curvature of the corners of the mouth reflects the morphological changes of the lip region. Both types of features are closely related to the user's emotional state and can effectively represent the user's subtle emotional changes.

[0027] When extracting heart rate variability features using LSTM, the raw heart rate information is first preprocessed. Low-pass filtering removes high-frequency noise interference, ensuring the purity of the heart rate data. Then, interpolation converts non-uniformly spaced heart rate data into equally spaced sequences, achieving temporal consistency. The LSTM network, through the coordinated operation of input, forget, and output gates, accurately captures the temporal variation patterns of heart rate information and extracts heart rate variability features that characterize heart rate fluctuations.

[0028] like Figure 2 As shown, before constructing a multimodal emotion fusion model and weighting and fusing the voiceprint features, micro-expression features, and heart rate variability features based on an attention mechanism to output a user stress index, the multimodal emotion fusion trigger threshold is dynamically adjusted according to the emotion label of the voiceprint features. When the emotion label is excitement, the adjusted multimodal emotion fusion trigger threshold is 70% of the initial multimodal emotion fusion trigger threshold; when the emotion label is calm, the multimodal emotion fusion trigger threshold remains unchanged.

[0029] The initial multimodal emotion fusion trigger threshold is a preset baseline threshold, the value of which is determined based on statistical analysis of historical sample data, and is used to determine whether to trigger the multimodal emotion fusion operation.

[0030] When the voiceprint feature's emotion label is "excited," it indicates that the user may be experiencing strong emotional fluctuations. In this case, the trigger threshold is lowered to 70% of the initial threshold to reduce the starting barrier for multimodal emotion fusion operations, ensuring timely response and accurate capture of high emotional intensity states, and avoiding feature fusion delays or omissions caused by excessively high thresholds. When the voiceprint feature's emotion label is "calm," it indicates that the user's emotional state is relatively stable. The trigger threshold remains unchanged from its initial value to reduce unnecessary feature fusion calculations, improving system efficiency while ensuring analytical accuracy.

[0031] By the above mode, the multi-modal emotion fusion trigger threshold can be dynamically adjusted, the threshold is reduced to 70% of the initial value under excited emotion, the problem of insufficient response of traditional fixed threshold to emotional expression is solved, the emotion adaptation is more accurate; and the feature fusion trigger condition can be adjusted in combination with the emotion label, the interference of non-adaptation scene is reduced, the voiceprint, micro-expression, heart rate variability feature fusion is more suitable for the real-time emotion state of the user, and the output user stress index is more accurate; and the feature fusion trigger delay caused by emotional fluctuation can be avoided, especially for excited emotion, the overall psychological health service efficiency is improved.

[0032] S3, constructing a multi-modal emotion fusion model, weighting and fusing the voiceprint feature, the micro-expression feature and the heart rate variability feature based on an attention mechanism, and outputting a user stress index; The multi-modal emotion fusion model is constructed, the voiceprint feature, the micro-expression feature and the heart rate variability feature are weighted and fused based on an attention mechanism, and a user stress index is output, and the corresponding calculation formula is as follows: YHYL=SW*0.6+WBQ*α+XLBYX*β α+β=0.4 In the formula, YHYL represents the user stress index; SW represents the voiceprint feature; WBQ represents the micro-expression feature; XLBYX represents the heart rate variability feature; α represents the weight coefficient corresponding to the micro-expression feature; and β represents the weight coefficient corresponding to the heart rate variability feature.

[0033] The multi-modal emotion fusion model takes the voiceprint feature, the micro-expression feature and the heart rate variability feature as core inputs, dynamically allocates weights based on an attention mechanism and completes weighted fusion. The voiceprint feature maintains a fixed weight of 0.6 to ensure its core reference value in stress evaluation. When the micro-expression feature contains more significant stress-related facial action patterns, such as high frown frequency and low curvature value of the mouth corner arc, the attention mechanism will give α a larger value. When the heart rate variability feature shows more obvious stress physiological response, such as increased low-frequency component proportion and reduced heart rate fluctuation amplitude, the value of β will increase accordingly to ensure that the user stress index calculated by the final fusion can more accurately reflect the actual stress level of the user, solving the problem of fixed weight that is difficult to adapt to individual differences and dynamic environment. The dynamic allocation of α and β realizes the complementary effect of micro-expression and heart rate variability features under different stress states, and improves the robustness of the multi-modal emotion fusion model.

[0034] S4, obtaining an anxiety score reflecting the result of the user's psychological assessment based on the user stress index, and setting a first anxiety score threshold, if the anxiety score is less than or equal to the first anxiety score threshold, performing dynamic art healing on the user terminal, if the anxiety score is greater than the first anxiety score threshold, performing a shout relief healing on the user terminal.

[0035] Since there is a mapping relationship between the anxiety score and the user stress index, the user stress index can be converted into a quantitative indicator directly reflecting the user's anxiety level, i.e. the anxiety score, through a preset conversion algorithm. The score can represent the anxiety level in the user's current psychological state.

[0036] The first anxiety score threshold is the switching standard of the dynamic art healing mode and the shout relief healing mode. When the anxiety score is less than or equal to the first anxiety score threshold, it indicates that the user's anxiety level is relatively low. At this time, dynamic art healing is carried out for the user, i.e. through the guidance of art form to help the user relax. When the anxiety score is greater than the first anxiety score threshold, it indicates that the user's anxiety level is high. At this time, shout relief healing is carried out for the user to help the user quickly relieve the accumulated anxiety in the heart.

[0037] For example, the range of the user stress index is 0-100 points. Through a preset conversion algorithm, i.e. equal proportion mapping, the user stress index can be converted into an anxiety score of the same dimension, which is positively correlated, i.e. the higher the stress index, the higher the anxiety score, which directly represents the user's anxiety level. Set the first anxiety score threshold to 50 points. If the user stress index is 42 points, it is converted into an anxiety score of 42 points through equal proportion mapping, which is less than 50 points. At this time, it is determined that the user is in a mild anxiety state, and the dynamic art healing mode is triggered. If the user stress index is 68 points, it is converted into an anxiety score of 68 points, which is greater than 50 points. It is determined that the user is in a moderate or severe anxiety state, and the shout relief healing mode is started.

[0038] The dynamic art healing for the user terminal specifically includes: Setting a second anxiety score threshold, if the anxiety score is less than or equal to the second anxiety score threshold, directly matching a dance healing mode for the user terminal; If the anxiety score is greater than the second anxiety score threshold, the user terminal is preferentially matched with a breathing training guidance mode, and after the breathing training guidance mode is ended, the user terminal is matched with the dance healing mode.

[0039] The second anxiety score threshold is the determination standard for selecting different modes within the dynamic art healing, and its value is lower than the first anxiety score threshold, which is used to divide the applicable situations of the dance healing mode and the breathing training guidance mode plus the dance healing mode.

[0040] If the anxiety score is less than or equal to the second anxiety score threshold, it means that the user's anxiety level is in a lower level within the applicable range of dynamic art healing, and the dance healing mode can be directly matched. Through the stretching and rhythm guidance of dance movements, the user's emotions can be relaxed and regulated. If the anxiety score is greater than the second anxiety score threshold and less than or equal to the first anxiety score threshold, it means that the user's anxiety level is relatively high within the applicable range of dynamic art healing. At this time, the breathing training guidance mode is preferentially matched. Through regular breathing guidance, the user's emotions can be calmed and the state can be stabilized. After the breathing training guidance mode ends, the dance healing mode is matched to further deepen the emotional relief effect and realize a step-by-step healing process.

[0041] The setting of the above threshold and the selection of the healing mode are both based on the current anxiety state of the user. Through differentiated and progressive healing methods, the user's psychological state can be effectively intervened and regulated.

[0042] The matching process of the dance healing mode is as follows: A kinematics inverse solution algorithm is used to retrieve and match dance healing movements from a 3D action library based on the user stress index, and the action parameters of the dance healing movements include action speed, action amplitude, and music rhythm. Wherein, the mapping relationship between the user stress index and the action parameters of the dance healing movements satisfies: f(YHYL)=0.5*DZSD+0.3*DZFD+0.2*YYJZ In the formula, f represents the mapping function between the user stress index and the action parameters of the dance healing movements; DZSD represents the action speed of the dance healing movements; DZFD represents the action amplitude of the dance healing movements; YYJZ represents the music rhythm of the dance healing movements; A laser holographic technology is used to generate a holographic projection picture based on the matched dance healing movements and display it to the user terminal. Action capture data of the user terminal is collected through a camera, and the action speed is dynamically adjusted based on the OpenPose algorithm according to the action capture data.

[0043] It can be understood that the kinematics inverse solution algorithm can perform reverse retrieval in the 3D action library according to the user stress index, filter out action parameters that are suitable for the current user state from the pre-stored multiple sets of dance healing movements by analyzing the correlation between the stress index and the dance healing movements, and ensure that the retrieved dance healing movements can specifically relieve the stress state of the user.

[0044] The mapping function is used to establish a quantitative correlation between the user stress index and the dance healing action parameters, wherein the action speed, action amplitude and music rhythm are involved in the mapping calculation according to their respective weights in emotion regulation. The action speed reflects the execution frequency of the dance action, the action amplitude represents the stretching range of the action, and the music rhythm corresponds to the tempo of the accompaniment music. The three factors work together to form a linkage relationship with the user stress index through the mapping function, so that the overall style of the dance healing action matches the current stress level of the user.

[0045] The laser holographic technology is used to model and project the matched dance healing action in three dimensions, achieving a 0.1mm-level action accuracy restoration of the 3D action character, and generating a holographic projection picture with spatial stereoscopic effect. This picture can intuitively show the details and rhythm of the dance action, providing an immersive action reference for the user and enhancing the interactivity and experience of the artistic healing process.

[0046] The action capture data collected by the camera is used to reflect the user's following of the dance healing action. Through the OpenPose algorithm, the action capture data can be analyzed to identify the execution speed and standard of the user's action, and compared with the preset dance healing action parameters. When it is detected that the user's action lags behind the projected action or the action completion degree is low, the algorithm automatically adjusts the speed parameter of the dance healing action to adapt to the user's action ability and rhythm following, ensuring that the user can smoothly complete the dance healing process and improving the effectiveness of the artistic healing effect.

[0047] In the above manner, the dance healing action can be accurately adjusted according to the user's stress state and real-time performance, fully realizing the role of dance healing in emotion regulation.

[0048] As shown in Figure 3 , the method for performing scream relief healing on the user terminal specifically includes: Collecting the scream relief audio of the user terminal; Performing Fourier transform processing on the scream relief audio, i.e., converting the time domain information in the scream relief audio into frequency domain information, and performing voiceprint desensitization processing on the scream relief audio through a feature mask generation algorithm for frequency domain desensitization, to filter out the key voiceprint frequency band of 300 to 3400 Hz in the scream relief audio; AES-256 encrypting and storing the scream relief audio after voiceprint desensitization processing, and the generation process of the encryption key needs to embed the data hash value of the scream relief audio; According to a fixed interval of 10 seconds per block, performing blockchain hash notarization on the data hash value of the scream relief audio.

[0049] The core purpose of the Fourier transform processing is to convert the time domain signal of the shouting relief audio into a frequency domain signal that can be analyzed, providing frequency domain features for subsequent voiceprint desensitization processing. The feature mask generation algorithm of the frequency domain desensitization constructs a targeted mask model to accurately locate and filter out the key voiceprint frequency band of 300 to 3400 Hz, which is the core frequency band of voiceprint recognition. After filtering, the user's voiceprint information can be effectively avoided from being restored or misused. In addition, the volume and pitch change characteristics related to emotional relief in the shouting audio can be retained to ensure that the subsequent audio-based emotional feedback analysis is not affected.

[0050] Further, the shouting relief audio after voiceprint desensitization is encrypted using the AES-256 algorithm. During the encryption process, the data hash value of the shouting relief audio is embedded in the generation logic of the encryption key, so that the key and the audio data form a unique binding relationship. If the audio data is tampered with, the data hash value will change, causing the key to be invalid and unable to normally decrypt the audio. In this way, the integrity and security of the audio data during storage are ensured, and the data is prevented from being illegally tampered with or stolen.

[0051] Next, the data hash value of the shouting relief audio is segmented and processed at a fixed interval of 10 seconds per block. Each block corresponds to the hash value of a segment of audio, and the hash values of each block are sequentially uploaded to the blockchain system for notarization. The decentralized and tamper-proof nature of the blockchain ensures that the notarization record of each audio hash value cannot be modified unilaterally, and the notarization information can be verified by the blockchain nodes in a distributed manner, providing a credible basis for the subsequent tracing and integrity verification of the audio data. In this way, while ensuring the emotional relief effect of the user, the personal privacy and data security of the user are maximized.

[0052] After the user end is relieved through shouting, the average sound intensity of the shouting relief audio is obtained; When the average sound intensity is less than or equal to 90 decibels, a music relaxation list is automatically pushed, wherein the alpha wave music is matched according to the voiceprint features to form the music relaxation list; When the average sound intensity is greater than 90 decibels, a psychological counseling AI is automatically triggered to perform psychological counseling on the user end.

[0053] After the shouting relief therapy is completed, the intensity features of the voiceprint desensitization audio can be extracted and the average value can be calculated through the audio signal analysis algorithm. The average value is used to quantify the emotional release degree during the user's shouting relief process, providing a basis for the selection of subsequent intervention strategies.

[0054] When the average sound intensity is less than or equal to 90 decibels, it indicates that the degree of emotional release of the user after shouting is relatively mild, and at this time, the music relaxation list is automatically pushed. Specifically, based on the voiceprint features extracted in advance, the alpha wave music that matches the voiceprint emotional attributes of the user is screened through a feature matching algorithm. The frequency characteristics of the alpha wave music are coordinated with the brain wave frequency of the human body in a relaxed state, and the matching mechanism of the voiceprint features can ensure that the pushed music accurately matches the current emotional state of the user, further strengthening the emotional relaxation effect. In addition, the music relaxation list needs to be sorted from high to low according to the matching degree of the music and the user's emotion, so as to facilitate the user to preferentially select the music content that best meets their own needs.

[0055] When the average sound intensity is greater than 90 decibels, it indicates that the emotional intensity released by the user in the process of shouting is high, and there may still be anxiety or stress that has not been completely calmed down, at this time, the psychological counseling AI is automatically triggered. Specifically, the psychological counseling AI first calls the user's stress index, anxiety score and shouting audio related features stored in advance to build a comprehensive evaluation model of the user's current psychological state, and then generates personalized psychological counseling rhetoric and guidance logic based on the model, and transmits the counseling content to the user through the interactive interface of the user terminal. At the same time, feedback information of the user is received in real time during the counseling process, and the counseling strategy is dynamically adjusted to ensure the pertinence and effectiveness of the psychological counseling.

[0056] As shown in Figure 4 Fig. 1 is a system block diagram of a holographic projection psychological health interaction system provided by an embodiment of the present application, the system comprises: An information acquisition module is configured to acquire voiceprint information, micro-expression information and heart rate information of a user terminal through a microphone array, an infrared camera and a heart rate sensor, respectively. A feature extraction module is configured to perform MFCC feature extraction and GRU neural network classification on the voiceprint information to generate voiceprint features, extract micro-expression features in the micro-expression information based on a CNN, and extract heart rate variability features corresponding to the heart rate information based on an LSTM. A multi-modal fusion module is configured to build a multi-modal emotion fusion model, and perform weighted fusion on the voiceprint features, the micro-expression features and the heart rate variability features based on an attention mechanism to output a user stress index. A user healing module is configured to obtain an anxiety score reflecting a user psychological evaluation result based on the user stress index, and set a first anxiety score threshold. If the anxiety score is less than or equal to the first anxiety score threshold, dynamic art healing is performed on the user terminal, and if the anxiety score is greater than the first anxiety score threshold, shouting catharsis healing is performed on the user terminal.

[0057] Figure 4 The device of the embodiment shown in Fig. 1 can be used to perform the method shown inFigure 1 The steps in the method embodiments shown, the implementation principles and technical effects are similar, and will not be repeated here.

[0058] An electronic device includes a memory and a processor, the memory stores a computer program, when the processor runs the computer program stored in the memory, the processor executes the steps of any one of the holographic projection-based mental health interaction method.

[0059] As Figure 5 shown, is a hardware structure schematic diagram of an electronic device provided by an embodiment of the application, the electronic device 50 includes: a processor 51, a memory 52 and a computer program; wherein The memory 52 is used for storing the computer program, and the memory can also be a flash memory. The computer program is, for example, an application program, a functional module and the like for implementing the above method.

[0060] The processor 51 is used for executing the computer program stored in the memory to realize each step of the device in the above method. For details, please refer to the related description in the above method embodiment.

[0061] Optionally, the memory 52 can be independent or integrated with the processor 51.

[0062] When the memory 52 is a device independent of the processor 51, the device can further include: A bus 53 is used for connecting the memory 52 and the processor 51.

[0063] A readable storage medium, the readable storage medium stores a computer program, the computer program is executed by a processor to realize the steps of any one of the holographic projection-based mental health interaction method.

[0064] The readable storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates transfer of a computer program from one place to another. A computer storage medium can be any available medium that can be accessed by a general purpose or special purpose computer. For example, the readable storage medium can be coupled to the processor, such that the processor can read information from, and write information to, the readable storage medium. Of course, the readable storage medium can be a component of the processor. Accordingly, the processor and the readable storage medium can be considered to be a machine-readable storage medium that can be used to store instructions to perform any one or more of the methods described herein. The machine-readable storage medium can be a non-transitory machine-readable storage medium. Accordingly, any one or more of the methods described herein can be embodied in a non-transitory machine-readable storage medium. The machine-readable storage medium can be a non-transitory machine-readable storage medium. In the above embodiments of the device, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), or the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the method disclosed in the present application can be directly embodied in the hardware processor, or be executed by the combination of hardware and software modules in the processor.

[0065] The present application also provides a program product including execution instructions stored in a readable storage medium. At least one processor of the device can read the execution instructions from the readable storage medium, and the at least one processor executes the execution instructions to enable the device to implement the method provided by the various embodiments described above.

[0066] In the above embodiments of the device, it should be understood that the processor can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), or the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor or the like. The steps of the method disclosed in the present application can be directly embodied in the hardware processor, or be executed by the combination of hardware and software modules in the processor.

[0067] Through the introduction of the above embodiments, the present application can collect voiceprint information, micro-expression information and heart rate information through a microphone array, an infrared camera and a heart rate sensor respectively through the holographic projection mental health interaction method and system based on the present application; the voiceprint information is subjected to MFCC feature extraction and GRU neural network classification to generate voiceprint features, the micro-expression features in the micro-expression information are extracted based on CNN, and the heart rate variability features corresponding to the heart rate information are extracted based on LSTM; a multi-modal emotion fusion model is constructed, the voiceprint features, micro-expression features and heart rate variability features are weighted and fused based on an attention mechanism, and a user stress index is output; an anxiety score reflecting the user psychological evaluation result is obtained based on the user stress index, a first anxiety score threshold is set, if the anxiety score is less than or equal to the first anxiety score threshold, dynamic art healing is performed on the user end, if the anxiety score is greater than the first anxiety score threshold, a shout relief healing is performed on the user end, so that the device function integration efficiency can be improved, the user interaction experience can be improved, the dynamic feedback accuracy can be ensured, the privacy protection strength can be enhanced, and the problems of functional fragmentation, insufficient interaction experience, missing dynamic feedback and weak privacy protection of traditional mental health interaction devices can be solved.

[0068] The present application can break through the functional fragmentation of traditional mental health interaction devices, improve the data intercommunication rate from 30% of traditional devices to 98%, realize closed-loop analysis of the user's psychological state, shorten the process cycle of "evaluation-intervention-feedback" to within 5 minutes, extract MFCC features from voiceprint information and generate voiceprint features through GRU neural network classification, and dynamically adjust the multi-modal emotion fusion trigger threshold, breaking through the bottleneck of traditional voice wake-up that only supports fixed instructions, making the success rate of emotional voice wake-up jump from 60% to 95%, and supporting the recognition of more than 20 emotional modalities such as crying and trembling, supporting adaptive recognition of dialectal variants such as Sichuan Mandarin and Cantonese, and greatly improving the user interaction convenience and adaptability; the present application can obtain a user stress index based on the weighted fusion of voiceprint, micro-expression and heart rate variability features through an attention mechanism, retrieve matching dance healing actions from a 3D action library through a kinematics inverse solution algorithm, dynamically adjust the action speed based on the motion capture data through the OpenPose algorithm, so that the dance healing action retrieval delay is less than 10 milliseconds, the adjustment delay is less than 80 milliseconds, the matching degree of heart rate and dance rhythm is 0.85, the healing content and user emotions are linked in real time, and the art healing effect is improved; the present application can reduce the information entropy of desensitized voiceprint data by 70% through Fourier transform, voiceprint feature desensitization and AES-256 encryption storage, and cooperate with blockchain hash evidence, effectively avoiding the risk of voiceprint leakage, and protecting the privacy and security of users.

[0069] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A psychological health interaction method based on holographic projection, characterized in that, The method includes: Voiceprint information, micro-expression information, and heart rate information from the user's device are collected through a microphone array, an infrared camera, and a heart rate sensor, respectively. The voiceprint information is subjected to MFCC feature extraction and GRU neural network classification to generate voiceprint features. Micro-expression features are extracted from the micro-expression information based on CNN. Heart rate variability features corresponding to the heart rate information are extracted based on LSTM. A multimodal emotion fusion model is constructed, and the voiceprint features, micro-expression features and heart rate variability features are weighted and fused based on the attention mechanism to output the user stress index; An anxiety score reflecting the user's psychological assessment results is obtained based on the user stress index, and a first anxiety score threshold is set. If the anxiety score is less than or equal to the first anxiety score threshold, dynamic art therapy is performed on the user. If the anxiety score is greater than the first anxiety score threshold, shouting and catharsis therapy is performed on the user.

2. The method for psychological health interaction based on holographic projection according to claim 1, characterized in that, The process of extracting MFCC features and classifying them using a GRU neural network to generate voiceprint features, extracting micro-expression features from the micro-expression information based on CNN, and extracting heart rate variability features corresponding to the heart rate information based on LSTM specifically includes: The voiceprint information is segmented into continuous audio frames according to preset segmentation parameters, and background noise and invalid signals are filtered out to obtain effective voiceprint segments. The MFCC algorithm is used to sequentially perform Fourier transform, Mel filter bank mapping, and cepstral calculation on the effective voiceprint segments to extract MFCC features; The MFCC features are input into the input layer of the GRU neural network; the temporal dependencies of the MFCC features are captured through the reset and update gate mechanisms of the gated recurrent units in the 128-dimensional hidden layer of the GRU neural network, invalid noise features are filtered out and key voiceprint attributes are enhanced to generate initial voiceprint features; the classification layer of the GRU neural network classifies the initial voiceprint features into emotion type and language features through the trained classification model, mapping the high-dimensional initial voiceprint features into low-dimensional voiceprint features containing emotion labels and language labels; Based on CNN, facial dynamic features are extracted from the micro-expression information, including at least the frequency of frowning and the curvature of the corners of the mouth, and the micro-expression features are output. The heart rate information is modeled in time series using LSTM to capture the fluctuation pattern of the heart rate signal and obtain the heart rate variability characteristics.

3. The method for psychological health interaction based on holographic projection according to claim 2, characterized in that, Before constructing a multimodal emotion fusion model and weighting and fusing the voiceprint features, micro-expression features, and heart rate variability features based on an attention mechanism to output a user stress index, the multimodal emotion fusion trigger threshold is dynamically adjusted according to the emotion label of the voiceprint features. When the emotion label is excitement, the adjusted multimodal emotion fusion trigger threshold is 70% of the initial multimodal emotion fusion trigger threshold; when the emotion label is calm, the multimodal emotion fusion trigger threshold remains unchanged.

4. The method for psychological health interaction based on holographic projection according to claim 1, characterized in that, The multimodal emotion fusion model is constructed by weighting and fusing the voiceprint features, micro-expression features, and heart rate variability features based on an attention mechanism, and outputting a user stress index. The corresponding calculation formula is as follows: YHYL=SW*0.6+WBQ*α+XLBYX*β α+β=0.4 In the formula, YHYL represents the user stress index; SW represents voiceprint characteristics; WBQ represents micro-expression features; XLBYX represents heart rate variability features; α represents the weighting coefficient corresponding to the micro-expression features; β represents the weighting coefficient corresponding to the heart rate variability features.

5. The method for psychological health interaction based on holographic projection according to claim 1, characterized in that, The aforementioned dynamic art therapy for users specifically includes: A second anxiety score threshold is set. If the anxiety score is less than or equal to the second anxiety score threshold, then the user terminal is directly matched with the dance therapy mode. If the anxiety score is greater than the second anxiety score threshold, the user terminal will be preferentially matched with the breathing training guidance mode. After the breathing training guidance mode is completed, the user terminal will be matched with the dance therapy mode.

6. The method for psychological health interaction based on holographic projection according to claim 5, characterized in that, The matching process for the dance therapy pattern is as follows: Using the inverse kinematics algorithm, dance therapy movements are retrieved and matched from the 3D motion library based on the user stress index, and the motion parameters of the dance therapy movements include motion speed, motion amplitude and music rhythm. The user stress index and the movement parameters of the dance therapy movements satisfy the following mapping relationship: f(YHYL)=0.5*DZSD+0.3*DZFD+0.2*YYJZ In the formula, f represents the mapping function between the user stress index and the movement parameters of the dance therapy movement; DZSD represents the movement speed of the dance therapy movement; DZFD represents the movement amplitude of the dance therapy movement; YYJZ represents the musical rhythm of the dance therapy movement. Using laser holographic technology, a holographic projection image is generated based on the matched dance therapy movements and displayed to the user terminal; The user's motion capture data is collected by a camera, and the motion speed is dynamically adjusted based on the OpenPose algorithm.

7. The method for psychological health interaction based on holographic projection according to claim 1, characterized in that, The aforementioned cathartic and therapeutic shouting exercise for the user terminal specifically includes: Collect the audio of the user's shouts and venting; The shouting and venting audio is subjected to Fourier transform processing, that is, the time domain information in the shouting and venting audio is converted into frequency domain information. The shouting and venting audio is then subjected to voiceprint desensitization processing by a frequency domain desensitization feature mask generation algorithm, filtering out the key voiceprint frequency band of 300 to 3400 Hz in the shouting and venting audio. The shouting and venting audio after voiceprint desensitization is encrypted and stored using AES-256, and the encryption key generation process needs to embed the data hash value of the shouting and venting audio. The data hash value of the shouting and venting audio is stored on the blockchain at a fixed interval of 10 seconds per block.

8. The method for psychological health interaction based on holographic projection according to claim 7, characterized in that, After the user terminal performs a cathartic shouting session, the average sound intensity of the cathartic shouting audio is obtained. When the average sound intensity is less than or equal to 90 decibels, a music relaxation list is automatically pushed, wherein alpha wave music is matched according to the voiceprint characteristics to form the music relaxation list; When the average sound intensity is greater than 90 decibels, the AI ​​for psychological counseling is automatically triggered to provide psychological counseling to the user.

9. A holographic projection-based psychological health interaction system, applied to the holographic projection-based psychological health interaction method as described in any one of claims 1-8, characterized in that, The system includes: The information acquisition module is used to collect voiceprint information, micro-expression information and heart rate information from the user terminal through a microphone array, an infrared camera and a heart rate sensor, respectively. The feature extraction module is used to perform MFCC feature extraction and GRU neural network classification on the voiceprint information to generate voiceprint features, extract micro-expression features from the micro-expression information based on CNN, and extract heart rate variability features corresponding to the heart rate information based on LSTM. The multimodal fusion module is used to construct a multimodal emotion fusion model. Based on the attention mechanism, it performs weighted fusion of the voiceprint features, the micro-expression features, and the heart rate variability features to output a user stress index. The user healing module is used to obtain an anxiety score reflecting the user's psychological assessment results based on the user stress index, and to set a first anxiety score threshold. If the anxiety score is less than or equal to the first anxiety score threshold, dynamic art healing is performed on the user. If the anxiety score is greater than the first anxiety score threshold, shouting and catharsis healing is performed on the user.