Artificial intelligence psychological assessment method and device based on multiple modes

Through multimodal data processing technology, including facial micro-expression segmentation, speech tone change and text emotional correlation, the feature distortion problem caused by anonymization in artificial intelligence psychological assessment is solved, and high-sensitivity emotional state interpretation and error reduction are achieved.

CN120600318AActive Publication Date: 2025-09-05HANGZHOU XINWA TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511089381.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-05
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

In the artificial intelligence psychological assessment, the structural distortion of biometrics and local desensitization of text semantics destroys the contextual emotional correlation characteristics due to anonymization, resulting in an increase in detection error rate and the subtle changes in user psychological state cannot be accurately captured.

Method used

By collecting user's facial micro-expression videos, regional pixel segmentation is performed to generate privacy interference levels; collecting voice pattern variation processing is performed; extracting emotional keywords in text semantic data to build cross-stage emotional association chains; adjusting the fusion weights between video and voice in combination with privacy interference levels, generating phased evaluation parameters and timeline labels, and building a psychological state evolution map.

Benefits of technology

It realizes accurate identification of facial micro-expressions and speech emotions under the premise of privacy protection, repair semantic breaks of multimodal data, improves the sensitivity and accuracy of emotional recognition, reduces the risk of misjudgment, and can intuitively present the spatial and temporal laws of emotional evolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600318A_ABST
    Figure CN120600318A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence psychological assessment method and equipment based on multiple modes, particularly relates to the technical field of psychological assessment analysis, and aims to solve the problem of high detection error rate caused by biological characteristic distortion due to privacy desensitization in existing psychological assessment. The method comprises the following steps of: dynamically generating an evaluation instruction matched with a user behavior, collecting a facial micro-expression video, segmenting and quantifying a privacy interference level based on region pixels, and fuzzifying a non-sensitive region while keeping key micro-expression region characteristics; synchronously acquiring a voice spectrum waveform and reconstructing voiceprint tone modification and fundamental frequency waveform; extracting text sentiment keywords to construct a cross-stage sentiment association chain to repair semantic fracture caused by local desensitization; and according to the privacy interference level, dynamically adjusting the fusion weight of the video and the voice to generate a stage evaluation parameter and a time sequence label, and finally constructing a psychological state evolution graph and outputting an evaluation report, so that the emotional feature fidelity and the psychological state analysis precision are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of psychological assessment and analysis, and more specifically, to a multimodal artificial intelligence psychological assessment method and device. Background Art

[0002] In artificial intelligence psychological assessments, existing technologies generally adopt multi-source data collection methods such as video, voice, and text, and construct user psychological portraits through a model that combines biometric recognition with semantic analysis; under medical compliance requirements, strict privacy protection standards must be followed during the data collection process. For example, the original data containing personal identity information is desensitized through technologies such as facial blurring and voiceprint modulation to meet restrictions on the storage and use of biometric data.

[0003] However, existing anonymization processing will lead to structural distortion of key biometric features, and the local desensitization operation of text semantics will destroy the contextual emotional association features. This multimodal feature degradation phenomenon caused by privacy constraints makes it impossible for artificial intelligence models to accurately capture subtle changes in the user's psychological state, resulting in a significant increase in the detection error rate. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the embodiments of the present invention provide a multimodal artificial intelligence psychological assessment method and device to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions: A multimodal artificial intelligence psychological assessment method includes the following steps: S1. Generate dynamic evaluation instructions based on user behavior history and send them to users for execution; S2. Collect the user's facial micro-expression video and perform pixel segmentation on the region, and generate a privacy interference level based on the pixel activity intensity of the segmented region; S3, collecting the user's voice spectrum waveform and performing voiceprint pitch shifting processing to generate a pitch-shifted voice spectrum waveform; S4, extract emotional keywords from the user's text semantic data and build a cross-stage emotional association chain based on the task sequence; S5. Combining the pitch-shifted speech spectrum waveform reconstructed from the fundamental frequency waveform with the cross-stage emotional association chain, and adjusting the fusion weight of the video and speech according to the privacy interference level to generate stage-by-stage evaluation parameters and time series labels. S6. Construct a psychological state evolution map based on the stage-by-stage evaluation parameters and time series labels and generate a psychological assessment report.

[0006] In a preferred embodiment, generating a dynamic evaluation instruction based on the user's behavior history and issuing it to the user for execution includes: Extract interaction frequency, task completion time, and emotional labels from user historical interaction records as user behavior history; Match the difficulty level of cognitive assessment tasks based on interaction frequency, divide the trigger threshold of emotion-inducing tasks based on task completion time, and generate dynamic assessment instructions containing difficulty level and trigger threshold based on emotion tags; The dynamic evaluation instructions are sent to the user terminal through the visual operation interface, and the task execution countdown module is loaded on the user terminal to synchronously start the evaluation process.

[0007] In a preferred embodiment, a user's facial micro-expression video is collected and segmented into regional pixels, and a privacy interference level is generated based on the pixel activity intensity of the segmented region, including: The camera collects the user's facial micro-expression video when executing the dynamic assessment command, and captures the video clips including the eye area, mouth corner area and forehead area; The eye area and mouth corner area are processed by edge detection algorithm to segment the micro-expression muscle motor unit area, retaining the original pixel value of the micro-expression muscle motor unit area and performing Gaussian blur encryption on other facial areas; Calculate the pixel grayscale change frequency of the segmented area within the preset time window as the pixel activity intensity. When the pixel grayscale change frequency exceeds the preset grayscale threshold, it is marked as a high-activity area. Based on the proportion of high-activity areas and the average grayscale change amplitude, a privacy interference level value is generated through a linear mapping function.

[0008] In a preferred embodiment, collecting the user's speech spectrum waveform and performing voiceprint pitch shifting processing to generate a pitch-shifted speech spectrum waveform includes: The microphone array is used to collect the user's voice signal when executing the dynamic assessment command, and the original voice spectrum waveform is generated through Fourier transform; Perform voiceprint pitch shifting on the original speech spectrum waveform, adjust the fundamental frequency to the preset gender-neutral frequency band, and randomly perturb the formant distribution to eliminate voiceprint characteristics; The fundamental frequency waveform of the pitch-shifted speech spectrum waveform is reconstructed, the pitch-shifted speech spectrum waveform is retained, and the pitch-shifted speech spectrum waveform is generated and stored in the buffer queue.

[0009] In a preferred embodiment, emotional keywords are extracted from the user's text semantic data, and a cross-stage emotional association chain is constructed based on the task sequence, including: Perform word segmentation on the text semantic data when the user executes the dynamic evaluation instruction, match the emotional keywords based on the emotional dictionary and generate an emotional keyword list; Based on the task timing timestamps of the dynamic evaluation instructions, the list of emotional keywords is divided into multiple subsets according to the task stages; The co-occurrence frequency and position offset of emotional keywords between adjacent task stage subsets are counted, and a cross-stage emotional association chain is established when the co-occurrence frequency exceeds a preset threshold.

[0010] In a preferred embodiment, the pitch-shifted speech spectrum waveform reconstructed from the fundamental frequency waveform is combined with the cross-stage emotional association chain, and the fusion weight of the video and speech is adjusted according to the privacy interference level to generate stage-by-stage evaluation parameters and time series labels, including: Apply dynamic compensation to the video feature vector after regional pixel segmentation based on the privacy interference level value; Semantic consistency parameters are generated based on the co-occurrence frequency and position offset in the cross-stage emotional association chain, and the intonation feature vector of the pitch-shifted speech spectrum waveform is contextually calibrated. The compensated video feature vector and the calibrated intonation feature vector are aligned with each other according to the task timing timestamp. Stage-by-stage evaluation parameters are generated based on the continuity of the emotional feature evolution in the aligned events, and task timing association labels are marked according to the distribution density of position offsets in the cross-stage emotional association chain.

[0011] In a preferred embodiment, the compensation intensity of the dynamic compensation is inversely correlated with the privacy interference level value. When the interference level is high, the contribution of the features of the periocular area is enhanced to offset the blurring loss. The calibration direction of contextual calibration is consistent with the keyword emotion polarity in the cross-stage emotion association chain.

[0012] In a preferred embodiment, a psychological state evolution map is constructed based on the staged evaluation parameters and the time series labels, and a psychological assessment report is generated, including: Integrate the phased evaluation parameters into the timeline according to the task sequence timestamps, and generate the node weights of the mental state evolution graph based on the continuity of the emotional feature evolution; Generate the connection strength between nodes based on the distribution density and standard deviation of the position offset in the cross-stage emotional association chain; Generate a psychological assessment report based on node weights and inter-node connection strengths; the psychological assessment report includes a depression tendency fluctuation curve and an anxiety state heat map.

[0013] In a preferred embodiment, the connection strength between nodes is dynamically adjusted by the ratio of distribution density to standard deviation, where the connection strength between nodes is enhanced when the distribution density increases, and the connection strength between nodes is weakened when the standard deviation increases.

[0014] In another aspect, the present invention provides a multimodal artificial intelligence psychological assessment device, comprising: Behavior instruction generation module: generates dynamic evaluation instructions based on user behavior history and sends them to users for execution; Facial Micro-Privacy Grading Module: This module collects user facial micro-expression videos and performs pixel segmentation, generating a privacy interference level based on the pixel activity intensity of the segmented regions. Voiceprint spectrum modulation module: collects the user's voice spectrum waveform and performs voiceprint pitch shifting to generate a pitch-shifted voice spectrum waveform; Cross-stage emotion chain construction module: extracts emotional keywords from users' text semantic data and constructs cross-stage emotion association chains based on task sequence; Multi-modal fusion weight adjustment module: This module combines the pitch-shifted speech spectrum waveform reconstructed from the fundamental frequency waveform with the cross-stage emotional association chain. It adjusts the fusion weight of video and speech based on the privacy interference level to generate stage-by-stage evaluation parameters and time series labels. Mental map evolution report module: Constructs a mental state evolution map based on stage evaluation parameters and time series labels and generates a psychological assessment report.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. Through regional pixel segmentation, the key muscle movement units of facial micro-expressions are accurately identified. The core biometric features are retained while locally encrypting non-critical areas. Combined with the fundamental frequency waveform reconstruction technology after voiceprint pitch shifting, the identity features are eliminated while maintaining the regularity of voice intonation and emotion. This achieves an effective balance between privacy desensitization and feature preservation, avoiding the feature distortion caused by traditional global blurring. This enables the emotion recognition model to capture subtle changes in expression and voice intonation, significantly improving the sensitivity of emotional state judgment. 2. By constructing cross-stage emotional association chains and dynamically adjusting fusion weights, the problems of semantic discontinuity and credibility imbalance in multimodal data are resolved. Emotional keyword co-occurrence analysis and position offset statistics based on task time series are used to repair contextual association breaks caused by text desensitization and ensure the continuity of emotional context. At the same time, the fusion weights of video and voice are adaptively adjusted according to the level of privacy interference, prioritizing the feature contributions of high-credibility data sources, effectively suppressing the error accumulation of traditional unimodal analysis. Combined with the dynamic modeling capabilities of the psychological state evolution map, it can intuitively present the spatiotemporal laws of emotional evolution and significantly reduce the risk of misjudgment. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a flow chart of a multimodal artificial intelligence psychological assessment method of the present invention; Figure 2 This is a structural diagram of a multimodal artificial intelligence psychological assessment device of the present invention. DETAILED DESCRIPTION

[0017] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0018] Example 1: Figure 1 The present invention provides a multimodal artificial intelligence psychological assessment method, which includes the following steps: S1. Generate dynamic evaluation instructions based on user behavior history and send them to users for execution; S2. Collect the user's facial micro-expression video and perform pixel segmentation on the region, and generate a privacy interference level based on the pixel activity intensity of the segmented region; S3, collecting the user's voice spectrum waveform and performing voiceprint pitch shifting processing to generate a pitch-shifted voice spectrum waveform; S4, extract emotional keywords from the user's text semantic data and build a cross-stage emotional association chain based on the task sequence; S5. Combining the pitch-shifted speech spectrum waveform reconstructed from the fundamental frequency waveform with the cross-stage emotional association chain, and adjusting the fusion weight of the video and speech according to the privacy interference level to generate stage-by-stage evaluation parameters and time series labels. S6. Construct a psychological state evolution map based on the stage-by-stage evaluation parameters and time series labels and generate a psychological assessment report.

[0019] S1. Generate dynamic evaluation instructions based on user behavior history and send them to users for execution, including: Extract interaction frequency, task completion time, and emotional labels from user historical interaction records as user behavior history; Match the difficulty level of cognitive assessment tasks based on interaction frequency, divide the trigger threshold of emotion-inducing tasks based on task completion time, and generate dynamic assessment instructions containing difficulty level and trigger threshold based on emotion tags; The dynamic evaluation instructions are sent to the user terminal through the visual operation interface, and the task execution countdown module is loaded on the user terminal to synchronously start the evaluation process.

[0020] User behavior history is extracted by recording user interaction data in historical assessment tasks. This interaction data includes, but is not limited to, the number of times a user triggers a assessment task, the duration from the start of a single task to its submission, and the user's self-labeled emotional state tags. Interaction frequency is calculated by counting the number of times a user has actively initiated assessment tasks in the past week. For example, if a user completes assessment tasks more than a preset threshold in the past week, the interaction frequency is marked as high; if it falls below the threshold, it is marked as low.

[0021] Task completion time is calculated by recording the difference in timestamps from the time the user completes the task interface loading to the time they submit their results. For example, in a cognitive assessment task, the time between starting to answer and clicking the Submit button is 3 minutes and 20 seconds. Emotion labels can be obtained by selecting the emotional state option through the interface when performing historical tasks, or by using AI models to analyze the user's historical speech and text data. For example, a user may select "anxiety" in a task, or the model may identify rapid intonation in the speech spectrum and label it as anxious.

[0022] The difficulty level matching rules of cognitive assessment tasks are dynamically adjusted according to the interaction frequency. When the interaction frequency is high, cognitive assessment tasks containing high-order logical reasoning questions are generated, such as graphic pattern recognition or number sequence deduction questions; when the interaction frequency is low, cognitive assessment tasks containing basic cognitive test questions are generated, such as simple arithmetic or vocabulary memory questions.

[0023] The trigger thresholds for emotion-inducing tasks are determined based on task completion time statistics. When task completion time is less than 20% of the average for similar tasks, a high-intensity emotion-inducing task is triggered, such as a rapid decision-making stress test or an emergency scenario simulation. When task completion time is more than 20% of the average for similar tasks, a low-intensity emotion-inducing task is triggered, such as a relaxation scenario or soothing music. Dynamic assessment instructions are generated by combining the matched difficulty level with the trigger threshold and associating it with the dominant emotion type in the user's emotional tags. For example, if anxiety tends to account for more than 50% of the user's emotional tags, the dynamic assessment instructions will include guided tasks to alleviate anxiety.

[0024] The design of the visual operation interface includes a task instruction display area, a countdown display area and an operation button area. The task instruction display area presents the specific content of the dynamic assessment instructions in the form of graphics and text, such as using a flowchart to illustrate the answering steps of a cognitive assessment task, or demonstrating the situational setting of an emotion-inducing task through animation.

[0025] The task execution countdown module is loaded through synchronization with the user's terminal's system clock. After the user clicks the Start button, the countdown module displays the remaining time in seconds, automatically resetting the countdown duration when switching between different task phases. For example, the countdown for the cognitive assessment task phase is set to 5 minutes, and the countdown for the emotion induction task phase is set to 3 minutes. The synchronous start of the assessment process is achieved through network time protocol calibration between the user terminal and the server, ensuring consistency in the assessment progress timeline for users on multiple terminals. For example, if the time difference between the user terminal and the server exceeds 500 milliseconds, the countdown module is automatically corrected based on the server time.

[0026] S2. Collect the user's facial micro-expression video and perform pixel segmentation. Generate a privacy interference level based on the pixel activity intensity of the segmented area, including: The camera collects the user's facial micro-expression video when executing the dynamic assessment command, and captures the video clips including the eye area, mouth corner area and forehead area; The eye area and mouth corner area are processed by edge detection algorithm to segment the micro-expression muscle motor unit area, retaining the original pixel value of the micro-expression muscle motor unit area and performing Gaussian blur encryption on other facial areas; Calculate the pixel grayscale change frequency of the segmented area within the preset time window as the pixel activity intensity. When the pixel grayscale change frequency exceeds the preset grayscale threshold, it is marked as a high-activity area. Based on the proportion of high-activity areas and the average grayscale change amplitude, a privacy interference level value is generated through a linear mapping function.

[0027] The collection of facial micro-expression videos is achieved by using the user terminal camera to capture the dynamic facial images of the user in real time while executing the dynamic assessment instructions. The video collection range covers the preset angles of the front and side of the user's face. For example, the camera automatically focuses on the line connecting the user's eyes as the horizontal reference line to ensure that the area around the eyes, the corners of the mouth, and the forehead are completely captured. The video clips are captured by selecting the key frame sequence based on the task type of the dynamic assessment instruction. For example, in the emotion induction task stage, the video clips are captured from 3 seconds after the start of the task to 1 second before the end of the task to ensure the integrity of the micro-expression features. The area around the eyes is defined as a rectangular area with the pupils of both eyes as the center and extending outward by 2 cm. The area around the mouth is defined as a rectangular area with the midline of the lips as the reference and extending 3 cm to both sides. The forehead area is defined as a rectangular area below the hairline to above the brow bone.

[0028] The edge detection algorithm is applied to video frames around the eyes and mouth. The video frames are first converted to grayscale images. The boundary contours of micro-expression muscle motor units are then identified by detecting the locations of grayscale value mutations, such as the contraction boundary of the orbicularis oculi or the deformation boundary of the orbicularis oris. Micro-expression muscle motor unit regions are segmented by retaining pixels within the boundary contours. Gaussian blur encryption of other facial regions uses a blur radius with a standard deviation of 15 pixels to blur non-segmented areas, such as pixels in the forehead and cheeks to eliminate identifying features. The original pixel values ​​of the segmented regions are retained within the entire RGB color space, for example, retaining the red channel value of the eye area for vascular dilation analysis.

[0029] The calculation of the frequency of pixel grayscale changes is achieved by counting the grayscale value differences of the segmented area in continuous video frames. The preset time window is set to 30 frames of video data within 1 second. For example, the grayscale average value of each frame in the eye area is counted within 1 second and the difference between adjacent frames is calculated. When the difference exceeds the preset grayscale threshold and the number of changes reaches 30 times per second, the area is marked as a high-activity area. For example, the user blinks frequently, causing the grayscale change frequency of the eye area to reach 35 times per second. The average grayscale change amplitude is calculated by accumulating all differences that exceed the preset grayscale threshold and taking the average value. For example, if a segmented area has 20 differences exceeding the preset grayscale threshold within 1 second and the accumulated value is 200, the average amplitude is 10.

[0030] The privacy interference level is generated using a linear mapping function. This weighted sum of the percentage of high-activity areas and the average grayscale variation is then mapped to a range of 0 to 100. For example, when the percentage of high-activity areas is 30% and the average grayscale variation is 50, the privacy interference level is (30% × 0.6 + 50 × 0.4) × 100 = 70. The weighting coefficients of the linear mapping function are set based on the biological characteristics of micro-expression muscle motor units. For example, the weighting coefficient for the eye area is higher than that for the mouth corners to reflect their greater sensitivity to emotional representation.

[0031] S3. Collecting the user's voice spectrum waveform and performing voiceprint pitch shifting processing to generate a pitch-shifted voice spectrum waveform, including: The microphone array is used to collect the user's voice signal when executing the dynamic assessment command, and the original voice spectrum waveform is generated through Fourier transform; Perform voiceprint pitch shifting on the original speech spectrum waveform, adjust the fundamental frequency to the preset gender-neutral frequency band, and randomly perturb the formant distribution to eliminate voiceprint characteristics; The fundamental frequency waveform of the pitch-shifted speech spectrum waveform is reconstructed, the pitch-shifted speech spectrum waveform is retained, and the pitch-shifted speech spectrum waveform is generated and stored in the buffer queue.

[0032] Voice signals are collected using a microphone array deployed on the user terminal. The microphone array adopts a circular symmetrical layout, ensuring coverage of the user's sound field from 0 to 180 degrees when executing dynamic assessment commands. For example, when the user faces the terminal screen, the microphone array prioritizes voice signals within a 120-degree range directly in front of them. The voice signal sampling rate is set to 16kHz to cover the fundamental frequency and formant frequency range of human speech. For example, the fundamental frequency range for adult males is typically 85Hz to 180Hz, and for adult females, it is typically 165Hz to 255Hz. Fourier transform processing uses a Hanning window function to frame the voice signal, with each frame length of 25 milliseconds and a frame shift of 10 milliseconds. This converts the time-domain voice signal into a frequency-domain voice spectrum waveform. For example, the time-domain waveform of a user reading a text in an emotion-eliciting task is converted into a spectrum graph containing features such as fundamental frequency and formants.

[0033] The fundamental frequency of the voiceprint pitch shifting is adjusted using a linear scaling algorithm, mapping the original fundamental frequency to a preset gender-neutral frequency band of 180Hz to 220Hz, which covers the intersection of the fundamental frequencies of male and female voices. For example, if the user's original fundamental frequency is 200Hz, it remains unchanged; if the original fundamental frequency is 150Hz, it is increased to 190Hz; and if the original fundamental frequency is 250Hz, it is decreased to 210Hz. Random perturbations of the formant distribution are achieved by shifting the formant peak frequency by ±50Hz within the frequency range of 1kHz to 3kHz. For example, the original formant peak frequency is randomly shifted from 2500Hz to 2450Hz or 2550Hz. The offset is determined based on speech intelligibility experiments to ensure semantic integrity. Voiceprint features are eliminated by adjusting both the fundamental frequency and the formants, making the pitch-shifted speech unattainable through voiceprint recognition algorithms. For example, the pitch-shifted speech spectrum waveforms generated by the same user in different tasks cannot be matched to the same identity in the voiceprint database.

[0034] The fundamental frequency waveform reconstruction process extracts the fundamental frequency contour from the pitch-shifted speech spectrum waveform. This contour is then smoothed using a cubic spline interpolation algorithm to eliminate jagged fluctuations introduced by pitch shifting. For example, the fundamental frequency fluctuations corresponding to a rising pitch when a user reads a question are preserved, while the unnatural jumps caused by pitch shifting are smoothed out. Preserving the characteristics of pitch variation is achieved by comparing the relative rate of change of the fundamental frequency contour before and after reconstruction. For example, when the fundamental frequency of the user's speech increases from 200Hz to 240Hz within 1 second, the reconstructed fundamental frequency waveform retains a 25% relative increase. Storage of the pitch-shifted speech spectrum waveform is managed through a buffer queue, caching at least 5 seconds of speech data in timestamp order to ensure temporal alignment with the facial micro-expression video. For example, if the video data lags by 200 milliseconds due to network latency, the buffer queue provides speech data for the corresponding time window for simultaneous analysis. The buffer duration is set based on typical network latency statistics.

[0035] S4. Extract emotional keywords from the user's text semantic data and build a cross-stage emotional association chain based on the task sequence, including: Perform word segmentation on the text semantic data when the user executes the dynamic evaluation instruction, match the emotional keywords based on the emotional dictionary and generate an emotional keyword list; Based on the task timing timestamps of the dynamic evaluation instructions, the list of emotional keywords is divided into multiple subsets according to the task stages; The co-occurrence frequency and position offset of emotional keywords between adjacent task stage subsets are counted, and a cross-stage emotional association chain is established when the co-occurrence frequency exceeds a preset threshold.

[0036] The extraction of text semantic data is achieved through the text content entered by the user when executing dynamic assessment instructions. The text content includes but is not limited to the user's free description text in the emotion induction task, the answer text of the multiple-choice question in the cognitive assessment task, and the feedback text between tasks.

[0037] Word segmentation is implemented using a word segmentation tool based on dictionary matching, which segments text content into independent lexical units at a granular level. For example, the user input "I feel anxious but try to calm down" is segmented into "I / feel / anxious / but / try / to / calm down." The sentiment lexicon is constructed based on publicly available sentiment lexicons, such as a list of positive sentiment words (such as "happy" and "relaxed"), negative sentiment words (such as "anxious" and "depressed"), and degree adverbs (such as "very" and "slightly"). Emotional keywords are matched by traversing the segmentation results and performing a full match against the sentiment lexicon entries. For example, when the segmentation result is "anxious," the negative sentiment keyword is matched from the sentiment lexicon and added to the list.

[0038] The acquisition of task sequence timestamps is based on the system timing after the dynamic assessment instruction is sent to the user terminal. The start timestamp and end timestamp of each task stage are recorded uniformly by the server. For example, the timestamp of the emotion induction task stage is from 10 seconds to 50 seconds after the start, and the timestamp of the cognitive assessment task stage is from 55 seconds to 100 seconds after the start.

[0039] The emotional keyword list is divided based on the timestamp range of the task phase, extracting the corresponding text content from the segmentation results. For example, emotional keywords between seconds 10 and 50 are classified as the emotion induction task subset, while keywords between seconds 55 and 100 are classified as the cognitive assessment task subset. The subset division rules also include the logical relevance of the task types. For example, the word "difficulty" entered by the user in the cognitive assessment task and the word "pressure" in the subsequent emotion induction task are classified as adjacent subsets.

[0040] The co-occurrence frequency of adjacent task stage subsets was calculated by traversing the emotional keywords of the two subsets and counting the number of times the same keyword appeared in adjacent stages. For example, if "anxiety" appeared 3 times in the emotion induction task subset and 2 times in the cognitive assessment task subset, the co-occurrence frequency would be 2. The position offset was calculated by recording the difference in the first appearance position of the same keyword in adjacent subsets. For example, if "anxiety" first appeared at the 5th word position in the emotion induction task subset and the 3rd word position in the cognitive assessment task subset, the position offset would be 2.

[0041] The preset threshold is dynamically adjusted based on the total number of keywords in each task phase. For example, when the total number of keywords in the adjacent subset exceeds 20, the co-occurrence frequency threshold is set to 3; when the total number of keywords is less than 20, the threshold is set to 2. Cross-stage emotional association chains are established by connecting keywords and their offsets that meet the threshold conditions. For example, when the co-occurrence frequency of "anxiety" is 2 times and the offset is 2, the association chain "anxiety-emotional induction→cognitive evaluation" is generated.

[0042] S5. Combine the pitch-shifted speech spectrum waveform reconstructed from the fundamental frequency waveform with the cross-stage emotional association chain, and adjust the fusion weight of video and speech based on the privacy interference level to generate stage-by-stage evaluation parameters and time series labels, including: Dynamic compensation is applied to the video feature vector after regional pixel segmentation based on the privacy interference level value. The compensation strength is inversely correlated with the privacy interference level value. When the interference level is high, the contribution of the features in the periocular area is enhanced to offset the blurring loss. Semantic consistency parameters are generated based on the co-occurrence frequency and position offset in the cross-stage emotional association chain, and the intonation feature vector of the pitch-shifted speech spectrum waveform is context-calibrated. The calibration direction of the context calibration is consistent with the emotional polarity of the keywords in the cross-stage emotional association chain. The compensated video feature vector and the calibrated intonation feature vector are aligned with each other according to the task timing timestamp. Stage-by-stage evaluation parameters are generated based on the continuity of the emotional feature evolution in the aligned events, and task timing association labels are marked according to the distribution density of position offsets in the cross-stage emotional association chain.

[0043] The application of the privacy interference level value is achieved by dynamically compensating the contribution of the periocular area features in the video feature vector after regional pixel segmentation. The intensity of dynamic compensation is inversely correlated with the privacy interference level value. Specifically, the higher the privacy interference level value, the greater the intensity of feature compensation for the periocular area.

[0044] The compensation strength is calculated by determining the compensation coefficient based on the percentile of the privacy interference level. For example, when the privacy interference level is 70, the feature contribution of the periocular area is increased to 1.5 times the original value to offset the loss of detail caused by Gaussian blurring. Compensation is achieved by adjusting the weight distribution of the video feature vector. For example, in the video feature vector, the grayscale change frequency feature of the periocular area is weighted more, while the weights of other areas are decreased in proportion to the interference level.

[0045] The co-occurrence frequency and position offset in the cross-stage emotional association chain are applied by generating a semantic consistency parameter. The value of the semantic consistency parameter is positively correlated with the co-occurrence frequency and negatively correlated with the position offset. The semantic consistency parameter is calculated by dividing the co-occurrence frequency by the position offset plus 1. For example, if the co-occurrence frequency is 5 and the position offset is 2, the semantic consistency parameter is 5 / (2+1)=1.67.

[0046] The semantic consistency parameter is used to contextually calibrate the intonation feature vector of the pitch-shifted speech spectrum waveform. The calibration direction is determined by the emotional polarity of the keywords in the cross-stage emotional association chain. For example, when the proportion of negative emotional keywords in the emotional chain exceeds 60%, the fundamental frequency fluctuation amplitude of the intonation feature vector is calibrated to increase by 20% to enhance the consistency of emotional expression.

[0047] The inter-modal event alignment of the compensated video feature vector and the calibrated intonation feature vector is achieved by precisely matching the task timing timestamps. The precision of the task timing timestamps is set to millisecond level to ensure the accuracy of event alignment.

[0048] The logic behind event alignment is to identify the temporal relationship between micro-expression events (such as eye muscle contractions) in the video feature vector and intonation events (such as a sudden increase in fundamental frequency) in the intonation feature vector within the same or adjacent time windows. For example, if the time difference between an eye muscle contraction and a sudden increase in fundamental frequency is less than 500 milliseconds, it is considered an emotional synchronization event. Continuity analysis of emotional feature evolution is achieved by statistically analyzing the distribution density of synchronization events along the task time axis. For example, during the emotion induction task, if more than three synchronization events are detected every 10 seconds, it is marked as a period of high emotional continuity.

[0049] The generation of stage evaluation parameters is achieved by calculating the continuity index of the evolution of emotional features. The continuity index is the product of the distribution density of synchronous events and the semantic consistency parameter. For example, when the distribution density is 0.3 times / second and the semantic consistency parameter is 1.5, the continuity index is 0.45.

[0050] The marking rule of task temporal association labels is based on the distribution density of position offsets in the cross-stage emotional association chain. The distribution density is calculated by counting the frequency of occurrence of position offsets within a preset time window. For example, in the cognitive assessment task stage, if the standard deviation of the position offset is less than 1, it is marked as "high association", and if the standard deviation is greater than or equal to 1, it is marked as "low association".

[0051] S6. Construct a psychological state evolution map based on the stage-by-stage evaluation parameters and time series labels and generate a psychological assessment report, including: Integrate the phased evaluation parameters into the timeline according to the task sequence timestamps, and generate the node weights of the mental state evolution graph based on the continuity of the emotional feature evolution; The inter-node connection strength is generated based on the distribution density and standard deviation of the position offset in the cross-stage emotional association chain. The inter-node connection strength is dynamically adjusted by the ratio of distribution density to standard deviation. When the distribution density increases, the inter-node connection strength is strengthened, and when the standard deviation increases, the inter-node connection strength is weakened. Generate a psychological assessment report based on node weights and inter-node connection strengths; the psychological assessment report includes a depression tendency fluctuation curve and an anxiety state heat map.

[0052] The integration of phase-specific assessment parameters is achieved by sequentially arranging the task timestamps. The task timestamps are set to millisecond precision to ensure accurate data alignment. During the integration process, the assessment parameters corresponding to different task phases are arranged on the timeline in timestamp order. For example, the emotion induction task phase has timestamps from 10 to 50 seconds, while the cognitive assessment task phase has timestamps from 55 to 100 seconds. The assessment parameters are then populated sequentially within this time range.

[0053] The node weights of the mental state evolution graph are generated based on the continuity index of the emotional feature evolution in the stage-by-stage assessment parameters. The higher the continuity index, the greater the node weight. For example, when the continuity index of the emotion induction task stage is 0.8, the corresponding node weight is set to 0.8, and when the continuity index of the cognitive assessment task stage is 0.5, the node weight is set to 0.5. The specific values ​​of the node weights are converted to the range of 0 to 1 through a linear mapping function. For example, a continuity index of 0.45 is mapped to a node weight of 0.45.

[0054] The distribution density and standard deviation of position offsets in the cross-stage emotional association chain were calculated by counting the position differences between the occurrences of the same emotional keyword in different task stages. The distribution density was calculated by counting the number of occurrences of position offsets within a preset time window. For example, in the emotion induction task, if the word "anxiety" appeared three times in the adjacent subset with position offsets of 1, 2, and 3, the distribution density was 3 times / 10 seconds. The standard deviation was calculated based on the dispersion of the position offsets. For example, the standard deviation of offsets 1, 2, and 3 was approximately 0.82. The strength of the connection between nodes was generated as the ratio of the distribution density to the standard deviation. When the distribution density increases or the standard deviation decreases, the connection strength increases. For example, if the distribution density is 3 times / 10 seconds and the standard deviation is 0.82, the connection strength is 3 / 0.82≈3.66; if the distribution density is 2 times / 10 seconds and the standard deviation is 1.5, the connection strength is 2 / 1.5≈1.33.

[0055] Psychological assessment reports are generated by combining node weights and connection strengths into a visual chart. The depressive tendency fluctuation curve is plotted based on the decay trend of node weights over time. This decay trend is determined by calculating the difference between adjacent node weights. For example, if a node weight drops from 0.8 to 0.5 within a 10-second time interval, the decay rate is (0.8-0.5) / 10 = 0.03 / second.

[0056] The rendering of the anxiety state heatmap is based on the spatial distribution density of node weights along the task sequence timeline. Distribution density is calculated by counting the cumulative value of node weights per unit time. For example, if the cumulative node weight is 2.4 within 10 seconds of the emotion induction task phase, the distribution density is 0.24 / second. The heatmap's color gradient dynamically adjusts based on the distribution density value, with higher density resulting in darker colors. For example, a density of 0.24 / second corresponds to light red, and 0.5 / second corresponds to dark red.

[0057] Example 2: Figure 2 The present invention provides a structural diagram of a multimodal artificial intelligence psychological assessment device, which includes: Behavior instruction generation module: generates dynamic evaluation instructions based on user behavior history and sends them to users for execution; Facial Micro-Privacy Grading Module: This module collects user facial micro-expression videos and performs pixel segmentation, generating a privacy interference level based on the pixel activity intensity of the segmented regions. Voiceprint spectrum modulation module: collects the user's voice spectrum waveform and performs voiceprint pitch shifting to generate a pitch-shifted voice spectrum waveform; Cross-stage emotion chain construction module: extracts emotional keywords from users' text semantic data and constructs cross-stage emotion association chains based on task sequence; Multi-modal fusion weight adjustment module: This module combines the pitch-shifted speech spectrum waveform reconstructed from the fundamental frequency waveform with the cross-stage emotional association chain. It adjusts the fusion weight of video and speech based on the privacy interference level to generate stage-by-stage evaluation parameters and time series labels. Mental map evolution report module: Constructs a mental state evolution map based on stage evaluation parameters and time series labels and generates a psychological assessment report.

[0058] The above formulas are all dimensionless and numerical calculations. The formula is a formula that is closest to the actual situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.

[0059] It should be noted that the present invention can be deployed on the device itself to implement embedded applications, and can also be run on a PC or other terminal with a user interface, thereby meeting various hardware environments and usage requirements.

[0060] The above embodiments can be implemented in whole or in part via software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in the embodiments of this application are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0061] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and modules described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0062] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.

[0063] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, and may be located in one place or distributed across multiple network modules. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.

[0064] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.

[0065] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0066] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0067] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A multimodal artificial intelligence psychological assessment method, characterized in that: The steps include: S1. Generate dynamic evaluation instructions based on user behavior history and send them to users for execution; S2. Collect the user's facial micro-expression video and perform pixel segmentation on the region, and generate a privacy interference level based on the pixel activity intensity of the segmented region; S3, collecting the user's voice spectrum waveform and performing voiceprint pitch shifting processing to generate a pitch-shifted voice spectrum waveform; S4, extract emotional keywords from the user's text semantic data and build a cross-stage emotional association chain based on the task sequence; S5. Combining the pitch-shifted speech spectrum waveform reconstructed from the fundamental frequency waveform with the cross-stage emotional association chain, and adjusting the fusion weight of the video and speech according to the privacy interference level to generate stage-by-stage evaluation parameters and time series labels. S6. Construct a psychological state evolution map based on the stage-by-stage evaluation parameters and time series labels and generate a psychological assessment report.

2. The multimodal artificial intelligence psychological assessment method according to claim 1, characterized in that: Generate dynamic evaluation instructions based on user behavior history and send them to users for execution, including: Extract interaction frequency, task completion time, and emotional labels from user historical interaction records as user behavior history; Match the difficulty level of cognitive assessment tasks based on interaction frequency, divide the trigger threshold of emotion-inducing tasks based on task completion time, and generate dynamic assessment instructions containing difficulty level and trigger threshold based on emotion tags; The dynamic evaluation instructions are sent to the user terminal through the visual operation interface, and the task execution countdown module is loaded on the user terminal to synchronously start the evaluation process.

3. The multimodal artificial intelligence psychological assessment method according to claim 2, characterized in that: The user's facial micro-expression video is collected and segmented into pixels. The privacy interference level is generated based on the pixel activity intensity of the segmented area, including: The camera collects the user's facial micro-expression video when executing the dynamic assessment command, and captures the video clips including the eye area, mouth corner area and forehead area; The eye area and mouth corner area are processed by edge detection algorithm to segment the micro-expression muscle motor unit area, retaining the original pixel value of the micro-expression muscle motor unit area and performing Gaussian blur encryption on other facial areas; Calculate the pixel grayscale change frequency of the segmented area within the preset time window as the pixel activity intensity. When the pixel grayscale change frequency exceeds the preset grayscale threshold, it is marked as a high-activity area. Based on the proportion of high-activity areas and the average grayscale change amplitude, a privacy interference level value is generated through a linear mapping function.

4. The multimodal artificial intelligence psychological assessment method according to claim 3, characterized in that: Collect the user's voice spectrum waveform and perform voiceprint pitch shifting processing to generate a pitch-shifted voice spectrum waveform, including: The microphone array is used to collect the user's voice signal when executing the dynamic assessment command, and the original voice spectrum waveform is generated through Fourier transform; Perform voiceprint pitch shifting on the original speech spectrum waveform, adjust the fundamental frequency to the preset gender-neutral frequency band, and randomly perturb the formant distribution to eliminate voiceprint characteristics; The fundamental frequency waveform of the pitch-shifted speech spectrum waveform is reconstructed, the pitch-shifted speech spectrum waveform is retained, and the pitch-shifted speech spectrum waveform is generated and stored in the buffer queue.

5. The multimodal artificial intelligence psychological assessment method according to claim 4, characterized in that: Extract emotional keywords from the user's text semantic data and build a cross-stage emotional association chain based on the task sequence, including: Perform word segmentation on the text semantic data when the user executes the dynamic evaluation instruction, match the emotional keywords based on the emotional dictionary and generate an emotional keyword list; Based on the task timing timestamps of the dynamic evaluation instructions, the list of emotional keywords is divided into multiple subsets according to the task stages; The co-occurrence frequency and position offset of emotional keywords between adjacent task stage subsets are counted, and a cross-stage emotional association chain is established when the co-occurrence frequency exceeds a preset threshold.

6. The multimodal artificial intelligence psychological assessment method according to claim 5, characterized in that: The pitch-shifted speech spectrum waveform reconstructed from the fundamental frequency waveform is combined with the cross-stage emotional association chain. The fusion weight of video and speech is adjusted according to the privacy interference level to generate stage-by-stage evaluation parameters and time series labels, including: Apply dynamic compensation to the video feature vector after regional pixel segmentation based on the privacy interference level value; Semantic consistency parameters are generated based on the co-occurrence frequency and position offset in the cross-stage emotional association chain, and the intonation feature vector of the pitch-shifted speech spectrum waveform is contextually calibrated. The compensated video feature vector and the calibrated intonation feature vector are aligned with each other according to the task timing timestamp. Stage-by-stage evaluation parameters are generated based on the continuity of the emotional feature evolution in the aligned events, and task timing association labels are marked according to the distribution density of position offsets in the cross-stage emotional association chain.

7. The multimodal artificial intelligence psychological assessment method according to claim 6, characterized in that: The intensity of dynamic compensation is inversely correlated with the privacy interference level. When the interference level is high, the contribution of features in the periocular area is enhanced to offset the blurring loss. The calibration direction of contextual calibration is consistent with the keyword emotion polarity in the cross-stage emotion association chain.

8. The multimodal artificial intelligence psychological assessment method according to claim 7, characterized in that: Construct a psychological state evolution map based on phased assessment parameters and time series labels and generate a psychological assessment report, including: Integrate the phased evaluation parameters into the timeline according to the task sequence timestamps, and generate the node weights of the mental state evolution graph based on the continuity of the emotional feature evolution; Generate the connection strength between nodes based on the distribution density and standard deviation of the position offset in the cross-stage emotional association chain; Generate a psychological assessment report based on node weights and inter-node connection strengths; the psychological assessment report includes a depression tendency fluctuation curve and an anxiety state heat map.

9. The multimodal artificial intelligence psychological assessment method according to claim 8, characterized in that: The connection strength between nodes is dynamically adjusted by the ratio of distribution density to standard deviation. When the distribution density increases, the connection strength between nodes is enhanced, and when the standard deviation increases, the connection strength between nodes is weakened.

10. A multimodal artificial intelligence psychological assessment device, used to implement the multimodal artificial intelligence psychological assessment method according to any one of claims 1 to 9, characterized in that: include: Behavior instruction generation module: generates dynamic evaluation instructions based on user behavior history and sends them to users for execution; Facial Micro-Privacy Grading Module: This module collects user facial micro-expression videos and performs pixel segmentation, generating a privacy interference level based on the pixel activity intensity of the segmented regions. Voiceprint spectrum modulation module: collects the user's voice spectrum waveform and performs voiceprint pitch shifting to generate a pitch-shifted voice spectrum waveform; Cross-stage emotion chain construction module: extracts emotional keywords from users' text semantic data and constructs cross-stage emotion association chains based on task sequence; Multi-modal fusion weight adjustment module: This module combines the pitch-shifted speech spectrum waveform reconstructed from the fundamental frequency waveform with the cross-stage emotional association chain. It adjusts the fusion weight of video and speech based on the privacy interference level to generate stage-by-stage evaluation parameters and time series labels. Mental map evolution report module: Constructs a mental state evolution map based on stage evaluation parameters and time series labels and generates a psychological assessment report.

Citation Information

Patent Citations

  • Speech data processing method, apparatus, electronic device and readable storage medium

    CN108269579A

  • Prisoner emotion recognition method for multi-modal feature fusion based on self-weight differential encoder

    CN110751208A

  • Multi-modal sentiment analysis method and device based on multi-agent collaboration

    CN119908724A

  • User identity authenticity verification method and system based on multi-modal AI

    CN120068037A

  • Cognitive impairment assessment system and method based on general artificial intelligence

    CN120189067A

Cited By

  • Biological characteristic time sequence monitoring system based on micro expression-motion collaborative analysis

    CN121890941A