Psychological state analysis method and apparatus, computer device, and storage medium
Patent Information
- Application Number
- PCT/CN2025/135601
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-18
- Filing Date
- 2025-11-18
- Publication Date
- 2026-08-27
Smart Images

Figure CN2025135601_27082026_PF_FP_ABST
Abstract
Description
Methods, devices, computer equipment and storage media for analyzing psychological states
[0001] This application claims priority to Chinese Patent Application No. 2025101783897, filed on February 18, 2025, entitled “Method, Apparatus, Computer Equipment and Storage Medium for Psychological State Analysis”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the fields of artificial intelligence technology, fintech, and healthcare, specifically to a method, apparatus, computer device, and non-volatile computer-readable storage medium for analyzing mental states. Background Technology
[0003] Currently, with the ever-accelerating pace of modern society, people are experiencing increasing pressure, making mental health issues a growing focus of attention. Especially in the workplace, education, and family life, more and more people are facing anxiety, depression, and other emotional distress, leading to a rising demand for mental health screening and intervention. Current traditional mental health screening methods often rely on manual screening or standardized scales, which have limitations such as high subjectivity, time-consuming processes, and constraints related to personnel and resources.
[0004] In the fintech field, the application scenarios for psychological state analysis technology are gradually increasing. For example, financial institutions need to assess customers' credit risk and investment preferences, and customers' psychological states (such as emotional stability and risk tolerance) have a significant impact on their decision-making behavior. However, traditional technologies still have many shortcomings in achieving efficient, comprehensive, and accurate psychological state analysis. Traditional psychological assessment methods mainly rely on questionnaires or manual interviews, which are not only time-consuming and labor-intensive but also easily affected by subjective factors, making it difficult to accurately reflect customers' real-time psychological states. In addition, the customer data in the fintech field is massive and complex, and traditional technologies struggle to effectively integrate multimodal data (such as voice and video data) to comprehensively assess users' psychological states, resulting in limited accuracy and reliability of the analysis results.
[0005] In the healthcare field, psychological state analysis technology also has important applications. Mental health issues not only affect patients' quality of life but can also negatively impact physical health. For example, the mental state of patients with chronic diseases is crucial to their recovery process, and healthcare professionals need to understand patients' psychological state in a timely manner to provide personalized interventions. However, existing psychological state assessment methods also face many challenges in medical settings. Traditional psychological scale assessments rely on patient self-reporting, which may be subject to subjective bias and concealment, making it difficult to accurately reflect the true psychological state. Furthermore, the limited availability of medical resources makes it difficult for healthcare professionals to conduct frequent psychological assessments for every patient, resulting in some patients' potential psychological problems going undetected and unaddressed.
[0006] In summary, the inventors recognized the following main shortcomings of traditional techniques in analyzing psychological states:
[0007] 1. High subjectivity: Traditional psychological assessment methods rely on manual screening or standardized scales, which are easily affected by the subjective factors of the assessor and the assessee, and are difficult to objectively reflect the true psychological state.
[0008] 2. Inefficiency: Manual screening and questionnaires are time-consuming and labor-intensive, making it difficult to meet the needs of the fintech and healthcare sectors for large-scale, real-time psychological state assessments;
[0009] 3. Insufficient data integration capabilities: Traditional technologies struggle to effectively integrate multimodal data (such as voice and video data) and cannot comprehensively analyze multiple dimensions of psychological states.
[0010] Therefore, how to provide a psychological state analysis method, device, computer equipment, and non-volatile computer-readable storage medium that can efficiently, comprehensively, and accurately analyze users' psychological states is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0011] In view of the shortcomings of the prior art, the purpose of this application is to provide a psychological state analysis method, apparatus, computer device and non-volatile computer-readable storage medium, aiming to solve the problem of how to achieve efficient, comprehensive and accurate analysis of users' psychological states.
[0012] To achieve the above objectives, this application adopts the following technical solution:
[0013] Firstly, this application provides a method for analyzing mental states, comprising:
[0014] Acquire audio and facial video of target users answering target questions;
[0015] The question-and-answer audio and the facial video are respectively represented by features to obtain corresponding semantic feature vectors and facial feature vectors;
[0016] The semantic feature vector and the facial feature vector are fused to obtain a fused feature vector;
[0017] Using the fused feature vector as input, the psychological state assessment result of the target user is generated through the psychological state assessment model.
[0018] Secondly, this application provides a mental state analysis device, comprising:
[0019] The acquisition module is used to acquire the audio and facial video of the target user answering the target question;
[0020] The feature representation module is used to perform feature representation on the question-and-answer audio and the facial video respectively to obtain the corresponding semantic feature vector and facial feature vector;
[0021] The feature fusion module is used to fuse the semantic feature vector with the facial feature vector to obtain a fused feature vector.
[0022] The result generation module is used to take the fused feature vector as input and generate the psychological state assessment result of the target user through the psychological state assessment model.
[0023] Thirdly, this application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the psychological state analysis method as described above.
[0024] Fourthly, this application provides a non-volatile computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the mental state analysis method described above.
[0025] Compared to existing technologies, this application provides a method, apparatus, computer device, and non-volatile computer-readable storage medium for analyzing psychological states. The method involves acquiring audio and facial video of a target user answering a target question; representing the audio and facial video separately to obtain corresponding semantic feature vectors and facial feature vectors; fusing the semantic feature vectors and facial feature vectors to obtain a fused feature vector; and using the fused feature vector as input to generate a psychological state assessment result for the target user through a psychological state assessment model. Therefore, this application can achieve efficient, comprehensive, and accurate analysis of the user's psychological state. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 is a schematic diagram of the application environment of a psychological state analysis method provided in an embodiment of this application.
[0028] Figure 2 is a flowchart illustrating a psychological state analysis method provided in an embodiment of this application.
[0029] Figure 3 is a schematic diagram of the program modules of a psychological state analysis device provided in an embodiment of this application.
[0030] Figure 4 is a schematic diagram of the structure of a computer device provided in an embodiment of this application.
[0031] Figure 5 is another structural schematic diagram of a computer device provided in an embodiment of this application. Detailed Implementation
[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0033] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0034] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0035] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [the described condition or event] is detected," or "in response to detection of [the described condition or event]."
[0036] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0037] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0038] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0039] To illustrate the technical solution of this application, specific embodiments are described below.
[0040] This application provides a psychological state analysis method according to one embodiment, applicable to the application environment shown in Figure 1, wherein the client and server communicate via a network. The client includes, but is not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, and personal digital assistants (PDAs). The server can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0041] Please refer to Figure 2. One embodiment of this application provides a method for analyzing psychological states, wherein the method includes the following steps:
[0042] S100: Obtain audio and facial video of the target user answering the target question;
[0043] S200. Perform feature representation on the question-and-answer audio and the facial video respectively to obtain the corresponding semantic feature vector and facial feature vector;
[0044] S300. The semantic feature vector and the facial feature vector are fused to obtain a fused feature vector;
[0045] S400. Using the fused feature vector as input, the psychological state assessment result of the target user is generated through the psychological state assessment model.
[0046] In practical implementation, the psychological state analysis method of this embodiment achieves efficient, comprehensive, and accurate analysis of users' psychological states through multimodal data acquisition and fusion technology. Specifically, the method of this embodiment first acquires the audio and facial video of the target user answering the target question (S100). These two types of data capture the psychological state information of the target user from two dimensions: language expression and facial expression, respectively. Subsequently, by extracting the semantic feature vector of the audio and the facial feature vector of the video (S200), key information can be extracted from the language content and facial expressions. This multimodal feature extraction method ensures the comprehensiveness of the analysis. In S300, the semantic feature vector and the facial feature vector are fused to generate a fused feature vector. This fusion process integrates complementary information from language and facial expressions, thereby more accurately reflecting the user's true psychological state. Finally, the fused feature vector is analyzed using a psychological state assessment model to generate a psychological state assessment result (S400). This method based on multimodal data fusion not only improves the accuracy of psychological state analysis but also enhances analysis efficiency through automated processing, overcoming the limitations of traditional single-modal analysis methods and providing an efficient, comprehensive, and accurate technical solution for psychological state analysis.
[0047] Understandably, the psychological state analysis method provided in this application can be applied to psychological state analysis scenarios in the medical and health field. The following is a specific example:
[0048] In the healthcare field, the mental state of patients with chronic diseases is crucial to their recovery process. Healthcare professionals need to understand patients' mental state in order to provide personalized interventions. Traditional methods rely primarily on patient self-reporting, but these methods are susceptible to subjective bias and concealment.
[0049] The psychological state analysis method applied in this application:
[0050] Scenario Description: During regular patient follow-up appointments, audio recordings and facial videos of the patient's answers to questions are obtained through video calls or in-person interviews. These questions may cover the patient's feelings about their condition, life stress, and expectations regarding treatment.
[0051] Technical Implementation: Semantic analysis is performed on the question-and-answer audio to extract the patient's language expression features (such as descriptions of their condition and anxieties about the future), generating semantic feature vectors. Micro-expression analysis is performed on facial videos to extract the patient's facial expression features (such as anxiety or depression), generating facial feature vectors. The semantic feature vectors and facial feature vectors are fused to generate a fused feature vector. A psychological state assessment model is used to analyze the patient's psychological state (such as anxiety level and depression level), generating a psychological state assessment result.
[0052] Technological Benefits: Automated analysis rapidly generates psychological state assessment results, saving medical staff time and effort. By combining verbal and facial expression information, it comprehensively assesses the patient's psychological state, avoiding the limitations of single data dimensions. It accurately identifies the patient's psychological state, helping medical staff to promptly identify potential psychological problems and provide more effective interventions.
[0053] Understandably, the psychological state analysis method provided in this application embodiment can also be applied to psychological state analysis scenarios related to the financial technology field. The following is a specific example:
[0054] In the fintech sector, financial institutions need to accurately assess customers' credit risk and investment preferences to optimize credit decisions and investment recommendations. Traditional methods rely primarily on customers' financial data and credit history, but this information cannot fully reflect customers' psychological state and may lead to assessment biases.
[0055] The psychological state analysis method applied in this application:
[0056] Scenario Description: When a customer applies for a loan or credit card, financial institutions obtain audio and facial video of the customer's answers to questions via video call or in-person interview (S100). These questions may involve the customer's financial situation, repayment plan, and expectations for the future.
[0057] Technical Implementation: Semantic analysis is performed on the question-and-answer audio to extract the customer's language expression features (such as speech rate, tone, and word choice), generating a semantic feature vector. Micro-expression analysis is performed on the facial video to extract the customer's facial expression features (such as tension or anxiety), generating a facial feature vector. The semantic feature vector and the facial feature vector are fused to generate a fused feature vector. A psychological state assessment model is used to analyze the customer's psychological state (such as anxiety level and confidence level), generating a psychological state assessment result.
[0058] Technical Benefits: Automated analysis rapidly generates psychological state assessment results, saving time and costs associated with manual assessments. By combining verbal and facial expression information, it comprehensively assesses a customer's psychological state, avoiding the limitations of single data dimensions. It accurately identifies a customer's psychological state, helping financial institutions better assess credit risk and reduce default rates.
[0059] Furthermore, in one embodiment, the psychological state analysis method, wherein step S100, acquiring the audio and facial video of the target user answering the target question, specifically includes the following steps:
[0060] Determine the open-ended questions to be asked to the target user;
[0061] The audio and facial video of the target user answering the open-ended question are captured using audio and video capture devices, respectively.
[0062] The audio of the question and answer and the facial video are synchronized in time.
[0063] In practical implementation, this embodiment achieves the following technical effects by identifying open-ended questions and collecting audio and facial video of the target user's responses, while simultaneously performing time-synchronized processing: First, the design of open-ended questions guides users to express themselves freely, thereby obtaining richer and more authentic psychological state information and avoiding the information limitations that may arise from traditional closed-ended questions; second, the simultaneous acquisition of audio and video not only captures the user's verbal content but also records their facial expressions and nonverbal information, providing a comprehensive data foundation for subsequent multimodal analysis; finally, time-synchronized processing ensures the consistency of audio and video data in the temporal dimension, making subsequent feature extraction and fusion more accurate, thereby improving the precision and reliability of psychological state analysis. This process provides solid technical support for achieving efficient, comprehensive, and accurate psychological state analysis, and is particularly suitable for in-depth assessment of psychological states in the fields of fintech and healthcare.
[0064] The specific implementation process of the steps in this embodiment is roughly as follows:
[0065] (1) Identify open-ended questions
[0066] Question Design: Based on the goals and application scenarios of the psychological state analysis, design a series of open-ended questions. These questions should guide users to freely express their thoughts, feelings, or experiences, avoiding limiting the scope of their answers.
[0067] Examples: "Please describe a recent instance where you felt stressed." "What are your expectations and concerns about the future?"
[0068] Question Validation: Pre-test the designed open-ended questions to ensure they effectively guide user responses and avoid misunderstandings or ambiguities. Invite a small group of target users to test the answers and adjust the question wording based on feedback.
[0069] (2) Prepare the data acquisition equipment
[0070] Choose audio capture equipment: Select high-quality audio capture equipment, such as professional microphones or recording devices, to ensure clear capture of the user's voice. The equipment should have a high sampling rate (e.g., 44.1kHz or higher) and low noise characteristics.
[0071] Choose a video capture device: Select a high-definition video capture device, such as a webcam or smartphone camera, to ensure it can clearly capture the user's facial expressions. The device should have high resolution (e.g., 1080p or higher) and a stable frame rate (e.g., 30fps or higher).
[0072] Equipment calibration: Audio and video equipment is calibrated before acquisition to ensure time synchronization. This can be achieved using synchronization signals or timestamp technology to ensure temporal consistency between audio and video data.
[0073] (3) Collect audio of question and answer and facial video.
[0074] Guide users to answer: Clearly explain the content and answer requirements of open-ended questions to the target users, ensuring they understand the questions and can express themselves freely. During the data collection process, avoid guiding or interrupting users' answers to obtain the most authentic information.
[0075] Audio capture: Record the user's voice while answering questions using audio capture equipment. Ensure a quiet recording environment to minimize background noise interference.
[0076] Video capture: Use video capture equipment to record the user's facial expressions as they answer questions. Ensure the camera is pointed at the user's face and that there is sufficient lighting to clearly capture micro-expressions.
[0077] (4) Time synchronization processing
[0078] Data alignment: Time alignment is performed on the acquired audio and video data. This can be achieved using timestamps or synchronization signals to ensure temporal consistency between audio and video. For example, specialized software can be used to align audio and video data, ensuring a perfect timeline match.
[0079] Data verification: Check the time synchronization of audio and video data to ensure there is no time discrepancy. If a time discrepancy is found, readjust the data alignment until satisfactory synchronization is achieved.
[0080] (5) Data storage and preliminary processing
[0081] Data storage: The acquired audio and video data are saved to secure storage media, and each file is clearly labeled for subsequent processing and analysis.
[0082] Preliminary processing: Perform preliminary processing on the audio and video data, such as removing background noise, adjusting volume, and cropping unnecessary parts, to ensure that the data quality meets the requirements of subsequent analysis.
[0083] This embodiment, through the above process, efficiently acquires audio and facial video of target users answering open-ended questions, ensuring temporal consistency between the two. This process provides a high-quality multimodal data foundation for subsequent psychological state analysis, ensuring the comprehensiveness and accuracy of the analysis.
[0084] Furthermore, in one embodiment, the psychological state analysis method, wherein step S200, performing feature representation on the question-and-answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors, specifically includes the following steps:
[0085] The question-and-answer audio is converted into question-and-answer text using automatic speech recognition technology, and the question-and-answer text is input into a pre-trained text encoding model to generate the semantic feature vector;
[0086] The facial video is input into a pre-trained video coding model to generate the facial feature vector.
[0087] In practical implementation, this embodiment efficiently extracts key information reflecting a user's psychological state by performing feature representation on both the question-and-answer audio and facial video. Specifically, automatic speech recognition technology is used to convert the question-and-answer audio into question-and-answer text, and a pre-trained text encoding model is used to generate semantic feature vectors. This process accurately captures the user's language content and emotional inclination, transforming complex speech information into quantifiable semantic features. Simultaneously, facial video is input into a pre-trained video encoding model to generate facial feature vectors, enabling the extraction of subtle changes in the user's facial expressions, such as micro-expressions and emotional states. This multimodal feature extraction method not only fully utilizes the complementarity of linguistic and visual information but also ensures the accuracy and efficiency of feature extraction through the efficient processing of the pre-trained model. Ultimately, the generated semantic and facial feature vectors provide a comprehensive and accurate data foundation for subsequent psychological state assessment, significantly improving the reliability and depth of psychological state analysis, enabling its widespread application in complex scenarios in fields such as fintech and healthcare.
[0088] The specific implementation process of the steps in this embodiment is roughly as follows:
[0089] (1) Audio processing and semantic feature extraction
[0090] Speech Recognition: Automatic Speech Recognition (ASR) technology is used to convert the acquired question-and-answer audio into question-and-answer text. ASR technology can convert the language content in the speech signal into a processable text format. A high-precision ASR model (such as a deep learning-based model) is selected to ensure the accuracy of transcription.
[0091] Text encoding: The converted question-and-answer text is input into a pre-trained text encoding model to extract semantic features. The text encoding model can transform text content into high-dimensional feature vectors that reflect the semantic information and sentiment of the language.
[0092] Generating semantic feature vectors: The feature vectors output by the text encoding model are semantic feature vectors. These vectors can quantitatively represent the linguistic content and sentiment in user responses, providing linguistic-dimensional information for subsequent psychological state analysis.
[0093] (2) Video processing and facial feature extraction
[0094] Video Input and Preprocessing: The acquired facial videos are input into the pre-trained video coding model. Before input, the videos undergo preprocessing, including cropping, denoising, and frame rate adjustment, to ensure that the video quality meets the model's input requirements.
[0095] Facial feature extraction: Each frame of the video is analyzed using a pre-trained video coding model (such as a deep learning-based CNN or Transformer model) to extract facial expression features. These models are able to recognize micro-expressions, emotional states, and other facial movements.
[0096] Generate facial feature vectors: The feature vectors output by the video coding model are facial feature vectors. These vectors can quantitatively represent the user's facial expressions and emotional state when answering questions, providing visual information for subsequent psychological state analysis.
[0097] This embodiment, through the above process, can efficiently extract semantic feature vectors and facial feature vectors from question-and-answer audio and facial video. Through normalization and integration, these feature vectors provide a high-quality multimodal data foundation for subsequent psychological state analysis, significantly improving the comprehensiveness and accuracy of the analysis.
[0098] Furthermore, in one embodiment, the psychological state analysis method, wherein step S300, fusing the semantic feature vector with the facial feature vector to obtain a fused feature vector, specifically includes the following steps:
[0099] The semantic feature vector and the facial feature vector are normalized.
[0100] According to the preset fusion strategy, the normalized semantic feature vector and the facial feature vector are fused to obtain the fused preliminary feature vector;
[0101] The preliminary feature vector is verified, and when the verification result meets the preset requirements, the preliminary feature vector is used as the fused feature vector.
[0102] Furthermore, in the aforementioned psychological state analysis method, the step of verifying the preliminary feature vector, and using the preliminary feature vector as the fused feature vector when the verification result meets preset requirements, specifically includes the following steps:
[0103] Pre-set verification metrics for the initial feature vector;
[0104] The preliminary feature vector is evaluated using the verification metrics to generate an evaluation result;
[0105] When the evaluation result meets the preset requirements, the preliminary feature vector is used as the fused feature vector.
[0106] In practical implementation, this embodiment achieves effective integration of semantic feature vectors and facial feature vectors through normalization, feature fusion, and verification processes, thereby generating a high-quality fused feature vector. Specifically, normalization ensures the consistency of the numerical range and distribution of the two modal feature vectors, eliminating deviations caused by differences in units and providing a unified foundation for subsequent fusion. Subsequently, the normalized feature vectors are fused according to a preset fusion strategy (such as weighted averaging, concatenation, or deep learning fusion methods) to generate a preliminary feature vector. This process fully utilizes the complementarity of linguistic and visual information, enhancing the expressive power of the features. Finally, a verification step checks the quality of the preliminary feature vector; only when it meets preset requirements (such as feature completeness, validity, or matching degree with known psychological states) is it used as the final fused feature vector. This verification mechanism further guarantees the reliability and accuracy of the fused feature vector, providing high-quality data support for the input of subsequent psychological state assessment models, thereby significantly improving the accuracy and stability of psychological state analysis.
[0107] The specific implementation process of the steps in this embodiment is roughly as follows:
[0108] (1) Normalization of eigenvectors
[0109] Normalization method selection: Choose an appropriate normalization method (such as Min-Max normalization or Z-score standardization) to process the semantic feature vector and facial feature vector. The goal of normalization is to adjust the numerical range of the feature vector to a uniform interval (such as [0, 1]) or to standardize the distribution of the feature vector (mean of 0, variance of 1).
[0110] Semantic feature vector normalization: Normalize each dimension of the semantic feature vector.
[0111] Facial feature vector normalization: Each dimension of the facial feature vector is also normalized to ensure its numerical range is consistent with the semantic feature vector. The normalized feature vector eliminates dimensional differences between different modalities, providing a unified data foundation for subsequent fusion.
[0112] (2) Feature fusion
[0113] Choose a fusion strategy: Select an appropriate fusion strategy based on the application scenario and requirements. Common fusion strategies include weighted averaging, feature concatenation, or deep learning-based fusion methods (such as learning the optimal fusion method through neural networks).
[0114] Weighted average fusion: If weighted average fusion is selected, weights (such as α and 1-α) are assigned according to the importance of semantic features and facial features, and the preliminary feature vector after fusion is calculated. Here, α is the weight of semantic features, which can be adjusted according to experimental results or domain knowledge.
[0115] Feature concatenation and fusion: If feature concatenation and fusion is selected, the normalized semantic feature vector and facial feature vector will be directly concatenated into a longer vector.
[0116] Deep learning-based fusion: If a deep learning-based fusion method is chosen, the normalized feature vectors can be input into a pre-trained neural network model (such as a fully connected layer or a Transformer), allowing the model to automatically learn the optimal fusion method.
[0117] (3) Verification of preliminary eigenvectors
[0118] Validation metric setting: Pre-set validation metrics for the initial feature vectors, such as feature completeness, validity, or matching degree with known psychological states. These metrics can be adjusted according to application scenarios and needs.
[0119] Validation process: The preliminary feature vector is evaluated using pre-defined validation metrics. For example, it may be checked whether the preliminary feature vector contains all the necessary dimensions, or its effectiveness may be verified by comparing it with known psychological state data.
[0120] Result judgment and adjustment: If the evaluation result meets the preset requirements, the preliminary feature vector will be used as the final fused feature vector; if the preset requirements are not met, the fusion strategy will be adjusted according to the evaluation result, the preliminary feature vector will be regenerated and verified again until the final fused feature vector is obtained.
[0121] (4) Output of fused feature vectors
[0122] Store the fused feature vector: Store the validated fused feature vector into the specified data structure, ensuring that its format meets the input requirements of the subsequent psychological state assessment model.
[0123] Labeling and Recording: Label the fused feature vectors and record their sources (such as the weighting of semantic features and facial features) and verification results for subsequent analysis and traceability.
[0124] This embodiment efficiently fuses semantic feature vectors and facial feature vectors through the above process. Normalization ensures the consistency of data from different modalities, the selection and implementation of the fusion strategy achieves the integration of multimodal information, and the verification step guarantees the quality and reliability of the fused feature vector. The final generated fused feature vector provides high-quality data support for subsequent psychological state assessment, significantly improving the comprehensiveness and accuracy of the analysis.
[0125] Furthermore, in one embodiment, the psychological state analysis method, wherein step S400, taking the fused feature vector as input and generating the psychological state assessment result of the target user through a psychological state assessment model, specifically includes the following steps:
[0126] The fused feature vectors are input into the pre-trained emotion state assessment model and the psychological state assessment model, respectively, to generate corresponding emotion state assessment results and psychological state assessment results.
[0127] By integrating the emotional state assessment results and the psychological state assessment results using a weighted average method, a multidimensional state assessment result for the target user is obtained.
[0128] Furthermore, the aforementioned psychological state analysis method, wherein the step of integrating the emotional state assessment results and the psychological state assessment results using a weighted average method to obtain the multidimensional state assessment results of the target user specifically includes the following steps:
[0129] Pre-set the weighting parameters for integrating the emotional state assessment results and the psychological state assessment results;
[0130] Based on the weight parameters, the emotional state assessment results and the psychological state assessment results are integrated using a weighted average method to generate a multidimensional state assessment result for the target user.
[0131] Furthermore, the psychological state analysis method, after integrating the emotional state assessment results and the psychological state assessment results using a weighted average method to obtain the multidimensional state assessment results of the target user, specifically includes the following steps:
[0132] Based on the multidimensional state assessment results, risk factors of the target user in terms of emotional and psychological states are identified;
[0133] Develop a psychological state adjustment plan for the target users based on the aforementioned risk factors.
[0134] Furthermore, the aforementioned psychological state analysis method, wherein the step of formulating a psychological state adjustment plan for the target user based on the risk factors specifically includes the following steps:
[0135] The risk factors are classified and ranked to obtain the classification and ranking results;
[0136] Based on the classification and sorting results, a psychological state adjustment plan is formulated, and the psychological state adjustment plan is fed back to the target user terminal.
[0137] In practical implementation, this embodiment achieves comprehensive, accurate, and personalized analysis and adjustment of the target user's psychological state through hierarchical model evaluation, multi-dimensional result integration, and targeted risk identification and intervention plan formulation. Specifically, firstly, pre-trained emotion state assessment models and psychological state assessment models are used to analyze the fused feature vectors, generating assessment results for emotion and psychological states. This dual-model assessment method can capture the user's psychological characteristics from different perspectives, ensuring the comprehensiveness of the analysis. Subsequently, the two assessment results are integrated using a weighted average method to generate multi-dimensional state assessment results, further improving the accuracy of the analysis. Based on this, the method of this embodiment further identifies the user's risk factors in terms of emotion and psychological state, and formulates personalized psychological state adjustment plans based on the classification and ranking of risk factors. This process not only ensures the targeting of intervention measures, but also achieves closed-loop management from assessment to intervention by feeding the plan back to the user. Overall, the method of this embodiment can efficiently identify psychological state problems and provide scientific and personalized adjustment suggestions, applicable to complex application scenarios in fields such as fintech and healthcare, significantly improving the practicality and effectiveness of psychological state analysis.
[0138] The specific implementation process of the steps in this embodiment is roughly as follows:
[0139] (1) Multi-model evaluation
[0140] Input fused feature vectors: Input the normalized and fused feature vectors into the pre-trained emotion state assessment model and psychological state assessment model, respectively. The emotion state assessment model focuses on analyzing the user's emotional characteristics (such as anxiety, depression, mood swings), while the psychological state assessment model focuses more on assessing the user's overall psychological state (such as stress level, attention concentration, psychological resilience, etc.).
[0141] Generate assessment results: The emotion state assessment model and the psychological state assessment model output emotion state assessment results and psychological state assessment results, respectively. These results can be numerical scores, classification labels, or probability distributions, depending on the model design.
[0142] (2) Integration of multidimensional state assessment results
[0143] Weighted average integration: Based on preset weight parameters, the emotional state assessment results and psychological state assessment results are integrated using a weighted average method. The weight parameters can be adjusted according to the model's reliability and the importance of the application scenario. For example, if emotional state has a greater impact on the target user, then the emotional state assessment result is given a higher weight.
[0144] Generating multidimensional state assessment results: The integrated result is the multidimensional state assessment result of the target user, which comprehensively reflects the user's emotional and psychological state. This result provides a foundation for subsequent risk factor identification and adjustment plan development.
[0145] (3) Risk factor identification
[0146] Analyzing Multidimensional State Assessment Results: A thorough analysis of the multidimensional state assessment results identifies potential risk factors for users' emotional and psychological states. For example, high anxiety scores in the emotional state assessment or low attention span in the psychological state assessment may be considered risk factors.
[0147] Risk factors are categorized and prioritized based on their nature (e.g., emotional problems, psychological state problems) and severity (e.g., mild, moderate, severe). The purpose of prioritizing these risk factors is to determine which ones require priority intervention.
[0148] (4) Development of a psychological state adjustment plan
[0149] Develop personalized adjustment plans: Based on the classification and ranking results, develop personalized psychological state adjustment plans for users. These plans may include specific interventions such as relaxation training, emotion regulation techniques, and lifestyle adjustments.
[0150] Solution Feedback: The psychological adjustment plan is then fed back to the target user's terminal device. Feedback can be achieved in various ways, such as generating a detailed adjustment suggestion report, providing online guidance, or communicating with professionals.
[0151] This embodiment, through the above process, can efficiently generate multidimensional state assessment results for target users, identify risk factors based on these results, and formulate and provide feedback on personalized psychological state adjustment plans. This process not only ensures the comprehensiveness and accuracy of the assessment but also enhances the effectiveness of psychological state adjustment through personalized intervention measures, making it suitable for complex application scenarios in fields such as fintech and healthcare.
[0152] As can be seen from the above method embodiments, the psychological state analysis method provided in this application includes: acquiring audio and facial video of a target user answering a target question; performing feature representation on the audio and facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors; fusing the semantic feature vectors and facial feature vectors to obtain a fused feature vector; and using the fused feature vector as input to generate a psychological state assessment result for the target user through a psychological state assessment model. Thus, the method of this application can achieve efficient, comprehensive, and accurate analysis of the user's psychological state.
[0153] It should be understood that although this application provides the method operation steps as described in the embodiments or flowcharts, conventional or non-inventive labor may include more or fewer operation steps, and these operation steps are not necessarily executed sequentially according to the order of the embodiments or flowcharts. The order of steps listed in the embodiments or flowcharts is merely one way of executing many steps and does not represent the only execution order. It should be noted that there is no necessary sequential order between the above steps. Those skilled in the art can understand from the description of the embodiments of this application that the above steps may have different execution orders in different embodiments, that is, they may be executed in parallel or in exchange, etc. Moreover, at least some steps in the embodiments or flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but may be executed in turn, alternately, or synchronously with other steps or at least a part of the sub-steps or stages of other steps.
[0154] Based on the above method embodiments, please refer to Figure 3. Another embodiment of this application also provides a psychological state analysis device, wherein the device includes:
[0155] Module 11 is used to acquire the audio and facial video of the target user answering the target question;
[0156] Feature representation module 12 is used to perform feature representation on the question-and-answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors;
[0157] Feature fusion module 13 is used to fuse the semantic feature vector with the facial feature vector to obtain a fused feature vector;
[0158] The result generation module 14 is used to take the fused feature vector as input and generate the psychological state assessment result of the target user through the psychological state assessment model.
[0159] Furthermore, in one embodiment, the psychological state analysis device, wherein acquiring the audio and facial video of the target user answering the target question specifically includes:
[0160] Determine the open-ended questions to be asked to the target user;
[0161] The audio and facial video of the target user answering the open-ended question are captured using audio and video capture devices, respectively.
[0162] The audio of the question and answer and the facial video are synchronized in time.
[0163] Furthermore, in one embodiment, the psychological state analysis device, wherein performing feature representation on the question-and-answer audio and the facial video to obtain corresponding semantic feature vectors and facial feature vectors, specifically includes:
[0164] The question-and-answer audio is converted into question-and-answer text using automatic speech recognition technology, and the question-and-answer text is input into a pre-trained text encoding model to generate the semantic feature vector;
[0165] The facial video is input into a pre-trained video coding model to generate the facial feature vector.
[0166] Furthermore, in one embodiment, the psychological state analysis device, wherein fusing the semantic feature vector with the facial feature vector to obtain a fused feature vector specifically includes:
[0167] The semantic feature vector and the facial feature vector are normalized.
[0168] According to the preset fusion strategy, the normalized semantic feature vector and the facial feature vector are fused to obtain the fused preliminary feature vector;
[0169] The preliminary feature vector is verified, and when the verification result meets the preset requirements, the preliminary feature vector is used as the fused feature vector.
[0170] Furthermore, in the aforementioned psychological state analysis device, the step of verifying the preliminary feature vector, and using the preliminary feature vector as the fused feature vector when the verification result meets preset requirements, specifically includes:
[0171] Pre-set verification metrics for the initial feature vector;
[0172] The preliminary feature vector is evaluated using the verification metrics to generate an evaluation result;
[0173] When the evaluation result meets the preset requirements, the preliminary feature vector is used as the fused feature vector.
[0174] Furthermore, in one embodiment, the psychological state analysis device, wherein the step of using the fused feature vector as input to generate the psychological state assessment result of the target user through a psychological state assessment model specifically includes:
[0175] The fused feature vectors are input into the pre-trained emotion state assessment model and the psychological state assessment model, respectively, to generate corresponding emotion state assessment results and psychological state assessment results.
[0176] By integrating the emotional state assessment results and the psychological state assessment results using a weighted average method, a multidimensional state assessment result for the target user is obtained.
[0177] Furthermore, in the aforementioned psychological state analysis device, the step of integrating the emotional state assessment results and the psychological state assessment results using a weighted average method to obtain the multidimensional state assessment results of the target user specifically includes:
[0178] Pre-set the weighting parameters for integrating the emotional state assessment results and the psychological state assessment results;
[0179] Based on the weight parameters, the emotional state assessment results and the psychological state assessment results are integrated using a weighted average method to generate a multidimensional state assessment result for the target user.
[0180] Furthermore, the psychological state analysis device, after integrating the emotional state assessment results and the psychological state assessment results using a weighted average method to obtain the multidimensional state assessment results of the target user, specifically further includes:
[0181] Based on the multidimensional state assessment results, risk factors of the target user in terms of emotional and psychological states are identified;
[0182] Develop a psychological state adjustment plan for the target users based on the aforementioned risk factors.
[0183] Furthermore, in the aforementioned psychological state analysis device, the step of formulating a psychological state adjustment plan for the target user based on the risk factors specifically includes:
[0184] The risk factors are classified and ranked to obtain the classification and ranking results;
[0185] Based on the classification and sorting results, a psychological state adjustment plan is formulated, and the psychological state adjustment plan is fed back to the target user terminal.
[0186] It should be noted that, in the device embodiments of this application, the information interaction and execution process between the above modules are based on the same concept as the method embodiments of this application. For their specific functions and technical effects, please refer to the aforementioned method embodiments section, which will not be repeated here.
[0187] Based on the above method embodiments, another embodiment of this application provides a computer device, which can be a server, and its internal structure diagram is shown in Figure 4. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the psychological state analysis method server-side as described in any of the above method embodiments.
[0188] Based on the above method embodiments, another embodiment of this application provides a computer device, which can be a client, and its internal structure diagram is shown in Figure 5. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of the psychological state analysis method client-side as described in any of the above method embodiments.
[0189] Those skilled in the art will understand that the structural schematic diagrams shown in Figures 4 and 5 are merely schematic diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more components than shown in the figures, or combine certain components, or have different component arrangements.
[0190] The processor referred to here can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0191] The memory includes readable storage media, internal memory, etc., where internal memory can be the RAM of a computer device. Internal memory provides an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of the computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory can include both internal storage units and external storage devices of the computer device. The memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. The memory can also be used to temporarily store data that has been output or will be output.
[0192] Based on the above method embodiments, another embodiment of this application provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the mental state analysis method as described in any of the above method embodiments. The computer-readable storage medium may be non-volatile or volatile.
[0193] It should be noted that the functions or steps that can be achieved by the computer-readable storage medium or computer device, and the technical effects brought about by the functions / steps, can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0194] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc. The disclosed memory components or memories of the operating environment described herein are intended to include one or more of these and / or any other suitable types of memory.
[0195] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the device embodiments of this application only illustrate the division of the above-mentioned functional units and modules. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0196] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0197] In the embodiments provided in this application, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0198] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0199] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for analyzing psychological states, wherein, include: Acquire audio and facial video of target users answering target questions; The question-and-answer audio and the facial video are respectively represented by features to obtain corresponding semantic feature vectors and facial feature vectors; The semantic feature vector and the facial feature vector are fused to obtain a fused feature vector; Using the fused feature vector as input, the psychological state assessment result of the target user is generated through the psychological state assessment model.
2. The psychological state analysis method according to claim 1, wherein, The acquisition of the target user's audio and facial video of answering the target question includes: Determine the open-ended questions to be asked to the target user; The audio and facial video of the target user answering the open-ended question are captured using audio and video capture devices, respectively. The question-and-answer audio and the facial video are time-synchronized.
3. The psychological state analysis method according to claim 1, wherein, The step of performing feature representation on the question-and-answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors includes: The question-and-answer audio is converted into question-and-answer text using automatic speech recognition technology, and the question-and-answer text is input into a pre-trained text encoding model to generate the semantic feature vector; The facial video is input into a pre-trained video coding model to generate the facial feature vector.
4. The psychological state analysis method according to claim 1, wherein, The step of fusing the semantic feature vector with the facial feature vector to obtain a fused feature vector includes: The semantic feature vector and the facial feature vector are normalized. According to the preset fusion strategy, the normalized semantic feature vector and the facial feature vector are fused to obtain the fused preliminary feature vector; The preliminary feature vector is verified, and when the verification result meets the preset requirements, the preliminary feature vector is used as the fused feature vector.
5. The psychological state analysis method according to claim 4, wherein, The verification of the preliminary feature vector, wherein when the verification result meets preset requirements, the preliminary feature vector is used as the fused feature vector, includes: Pre-set verification metrics for the initial feature vector; The preliminary feature vector is evaluated using the verification metrics to generate an evaluation result; When the evaluation result meets the preset requirements, the preliminary feature vector is used as the fused feature vector.
6. The psychological state analysis method according to any one of claims 1-5, wherein, The step of using the fused feature vector as input to generate the psychological state assessment result of the target user through the psychological state assessment model includes: The fused feature vectors are input into the pre-trained emotion state assessment model and the psychological state assessment model, respectively, to generate corresponding emotion state assessment results and psychological state assessment results. By integrating the emotional state assessment results and the psychological state assessment results using a weighted average method, a multidimensional state assessment result for the target user is obtained.
7. The psychological state analysis method according to claim 6, wherein, The process of integrating the emotional state assessment results and the psychological state assessment results using a weighted average method to obtain the multidimensional state assessment results of the target user includes: Pre-set the weighting parameters for integrating the emotional state assessment results and the psychological state assessment results; Based on the weight parameters, the emotional state assessment results and the psychological state assessment results are integrated using a weighted average method to generate a multidimensional state assessment result for the target user.
8. The psychological state analysis method according to claim 6, wherein, After integrating the emotional state assessment results and the psychological state assessment results using a weighted average method to obtain the multidimensional state assessment results of the target user, the process further includes: Based on the multidimensional state assessment results, risk factors for the target user in terms of emotional and psychological states are identified. Develop a psychological state adjustment plan for the target users based on the aforementioned risk factors.
9. The psychological state analysis method according to claim 8, wherein, The step of developing a psychological state adjustment plan for the target user based on the risk factors includes: The risk factors are classified and ranked to obtain the classification and ranking results; Based on the classification and sorting results, a psychological state adjustment plan is formulated, and the psychological state adjustment plan is fed back to the target user terminal.
10. A psychological state analysis device, wherein, include: The acquisition module is used to acquire the audio and facial video of the target user answering the target question; The feature representation module is used to perform feature representation on the question-and-answer audio and the facial video respectively to obtain the corresponding semantic feature vector and facial feature vector; The feature fusion module is used to fuse the semantic feature vector with the facial feature vector to obtain a fused feature vector. The result generation module is used to take the fused feature vector as input and generate the psychological state assessment result of the target user through the psychological state assessment model.
11. The psychological state analysis device according to claim 10, wherein, The acquisition of the target user's audio and facial video of answering the target question includes: Determine the open-ended questions to be asked to the target user; The audio and facial video of the target user answering the open-ended question are captured using audio and video capture devices, respectively. The question-and-answer audio and the facial video are time-synchronized.
12. The psychological state analysis device according to claim 10, wherein, The step of performing feature representation on the question-and-answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors includes: The question-and-answer audio is converted into question-and-answer text using automatic speech recognition technology, and the question-and-answer text is input into a pre-trained text encoding model to generate the semantic feature vector; The facial video is input into a pre-trained video coding model to generate the facial feature vector.
13. The psychological state analysis device according to claim 10, wherein, The step of fusing the semantic feature vector with the facial feature vector to obtain a fused feature vector includes: The semantic feature vector and the facial feature vector are normalized. According to the preset fusion strategy, the normalized semantic feature vector and the facial feature vector are fused to obtain the fused preliminary feature vector; The preliminary feature vector is verified, and when the verification result meets the preset requirements, the preliminary feature vector is used as the fused feature vector.
14. The psychological state analysis device according to claim 13, wherein, The verification of the preliminary feature vector, wherein when the verification result meets preset requirements, the preliminary feature vector is used as the fused feature vector, includes: Pre-set verification metrics for the initial feature vector; The preliminary feature vector is evaluated using the verification metrics to generate an evaluation result; When the evaluation result meets the preset requirements, the preliminary feature vector is used as the fused feature vector.
15. The psychological state analysis device according to any one of claims 10-14, wherein, The step of using the fused feature vector as input to generate the psychological state assessment result of the target user through the psychological state assessment model includes: The fused feature vectors are input into the pre-trained emotion state assessment model and the psychological state assessment model, respectively, to generate corresponding emotion state assessment results and psychological state assessment results. By integrating the emotional state assessment results and the psychological state assessment results using a weighted average method, a multidimensional state assessment result for the target user is obtained.
16. The psychological state analysis device according to claim 15, wherein, The process of integrating the emotional state assessment results and the psychological state assessment results using a weighted average method to obtain the multidimensional state assessment results of the target user includes: Pre-set the weighting parameters for integrating the emotional state assessment results and the psychological state assessment results; Based on the weight parameters, the emotional state assessment results and the psychological state assessment results are integrated using a weighted average method to generate a multidimensional state assessment result for the target user.
17. The psychological state analysis device according to claim 15, wherein, After integrating the emotional state assessment results and the psychological state assessment results using a weighted average method to obtain the multidimensional state assessment results of the target user, the process further includes: Based on the multidimensional state assessment results, risk factors for the target user in terms of emotional and psychological states are identified. Develop a psychological state adjustment plan for the target users based on the aforementioned risk factors.
18. The psychological state analysis device according to claim 17, wherein, The step of developing a psychological state adjustment plan for the target user based on the risk factors includes: The risk factors are classified and ranked to obtain the classification and ranking results; Based on the classification and sorting results, a psychological state adjustment plan is formulated, and the psychological state adjustment plan is fed back to the target user terminal.
19. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the psychological state analysis method as described in any one of claims 1-9.
20. A non-volatile computer-readable storage medium storing a computer program, wherein, When the computer program is executed by the processor, it implements the psychological state analysis method as described in any one of claims 1-9.