Psychological state analysis method and device, computer equipment and storage medium
By obtaining and integrating user's Q&A audio and facial video features and inputting a psychological state evaluation model, the subjectivity and inefficiency of psychological state analysis in the existing technology are solved, and efficient, comprehensive and accurate analysis results are achieved.
Patent Information
- Application Number
- CN202510178389.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-06
AI Technical Summary
The existing technology has problems such as strong subjectivity, low efficiency and insufficient data integration capabilities in psychological state analysis, making it difficult to achieve efficient, comprehensive and accurate analysis.
By obtaining the Q&A audio and facial video when the target user answers the question, feature representation is performed separately, semantic feature vectors and facial feature vectors are obtained, feature fusion, fusion feature vectors are generated, and psychological state evaluation model is input to generate evaluation results.
It realizes efficient, comprehensive and precise analysis of the user's psychological state, overcomes the problems of subjectivity and inefficiency of traditional methods, and improves the accuracy and reliability of the analysis through the integration of multimodal data.
Smart Images

Figure CN120105153A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology, financial technology and medical health, and in particular to a psychological state analysis method, device, computer equipment and computer-readable storage medium. Background Art
[0002] At present, with the accelerating pace of modern society, people are under increasing pressure, and mental health issues are becoming the focus of people's attention. Especially in the workplace, education and family life, more and more people are facing emotional distress such as anxiety and depression, and the demand for mental health screening and intervention is increasing. At present, traditional psychological screening methods often rely on manual screening or standardized scales, which are highly subjective, time-consuming, and limited by personnel and resources.
[0003] In the field of financial technology, the application scenarios of psychological state analysis technology are gradually increasing. For example, financial institutions need to assess customers' credit risks and investment preferences, and customers' psychological states (such as emotional stability and risk tolerance) have an important impact on their decision-making behavior. However, existing technologies still have many shortcomings in achieving efficient, comprehensive and accurate psychological state analysis. Traditional psychological assessment methods mainly rely on questionnaires or manual interviews. These methods are not only time-consuming and labor-intensive, but also easily interfered by subjective factors, and it is difficult to accurately reflect customers' real-time psychological states. In addition, the amount of customer data in the field of financial technology is huge and complex, and existing technologies are difficult to effectively integrate multimodal data (such as voice, video and other data) to comprehensively assess users' psychological states, resulting in limited accuracy and reliability of analysis results.
[0004] In the field of medical health, psychological state analysis technology also has important application scenarios. Mental health problems not only affect the patient's quality of life, but may also have a negative impact on physical health. For example, the mental state of patients with chronic diseases is crucial to their recovery process, and medical staff need to understand the patient's mental condition in a timely manner in order to provide personalized interventions. However, existing mental state assessment methods also face many challenges in medical scenarios. Traditional psychological scale assessments rely on patient self-reports, which may have subjective biases and concealment, and it is difficult to accurately reflect the true mental state. In addition, the limited medical resources make it difficult for medical staff to conduct frequent psychological assessments on each patient, resulting in the failure to timely discover and intervene in potential psychological problems of some patients.
[0005] In summary, the existing technology has the following main deficiencies in psychological state analysis:
[0006] 1. Highly subjective: Traditional psychological assessment methods rely on manual screening or standardized scales, which are easily affected by subjective factors of the assessor and the assessed, and are difficult to objectively reflect the true psychological state;
[0007] 2. Inefficiency: Manual screening and questionnaire surveys are time-consuming and labor-intensive, and are difficult to meet the needs of the financial technology and healthcare fields for large-scale, real-time psychological state assessment;
[0008] 3. Insufficient data integration capabilities: Existing technologies are unable to effectively integrate multimodal data (such as voice, video, and other data) and are unable to comprehensively analyze multiple dimensions of psychological states.
[0009] Based on this, how to provide a psychological state analysis method, device, computer equipment and computer-readable storage medium that can efficiently, comprehensively and accurately analyze the user's psychological state is a problem that urgently needs to be solved by technical personnel in this field. Summary of the invention
[0010] In view of the above-mentioned deficiencies in the prior art, the object of the present invention is to provide a psychological state analysis method, apparatus, computer device and computer-readable storage medium, aiming to solve the problem of how to efficiently, comprehensively and accurately analyze the user's psychological state.
[0011] In order to achieve the above object, the present invention adopts the following technical solutions:
[0012] In a first aspect, the present invention provides a method for analyzing a psychological state, comprising:
[0013] Obtain the target user’s audio and facial video when answering the target question;
[0014] Performing feature representation on the question-answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors;
[0015] Performing feature fusion on the semantic feature vector and the facial feature vector to obtain a fused feature vector;
[0016] The fused feature vector is used as input to generate a psychological state assessment result of the target user through a psychological state assessment model.
[0017] In a second aspect, the present invention provides a psychological state analysis device, comprising:
[0018] An acquisition module is used to acquire the question-answering audio and facial video of the target user when answering the target question;
[0019] A feature representation module, used to perform feature representation on the question-answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors;
[0020] A feature fusion module, used for fusing the semantic feature vector with the facial feature vector to obtain a fused feature vector;
[0021] The result generating module is used to take the fused feature vector as input and generate the psychological state evaluation result of the target user through the psychological state evaluation model.
[0022] In a third aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the psychological state analysis method as described above when executing the computer program.
[0023] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program implements the psychological state analysis method as described above when executed by a processor.
[0024] Compared with the prior art, the present invention provides a psychological state analysis method, apparatus, computer device and computer-readable storage medium, wherein the question and answer audio and facial video of the target user answering the target question are obtained; the question and answer audio and the facial video are respectively represented by features to obtain corresponding semantic feature vectors and facial feature vectors; the semantic feature vector and the facial feature vector are feature fused to obtain a fused feature vector; the fused feature vector is used as input to generate the psychological state evaluation result of the target user through a psychological state evaluation model; thereby, the present invention can realize efficient, comprehensive and accurate analysis of the user's psychological state. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0026] Figure 1 A schematic diagram of an application environment of a psychological state analysis method provided by an embodiment of the present invention.
[0027] Figure 2 A flowchart of a psychological state analysis method provided by an embodiment of the present invention.
[0028] Figure 3 A schematic diagram of a program module of a psychological state analysis device provided by an embodiment of the present invention.
[0029] Figure 4 A schematic diagram of the structure of a computer device provided by an embodiment of the present invention.
[0030] Figure 5 Another structural schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0033] It should also be understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0034] As used in the present specification and the appended claims, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0035] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0036] References to "one embodiment" or "some embodiments" etc. described in the present specification mean that one or more embodiments of the present invention include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0037] It should be understood that the order of execution of the steps in the following embodiments does not imply a precedence of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0038] In order to illustrate the technical solution of the present invention, specific embodiments are provided below for illustration.
[0039] A psychological state analysis method provided by an embodiment of the present invention can be applied in Figure 1 In the application environment shown, the client and the server communicate through the network. The client includes but is not limited to PDAs, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, personal digital assistants (PDAs) and other computer devices. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0040] See also Figure 2 An embodiment of the present invention provides a method for analyzing a psychological state, wherein the method comprises the following steps:
[0041] S100, obtaining the question-answering audio and facial video of the target user when answering the target question;
[0042] S200, performing feature representation on the question-answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors;
[0043] S300, performing feature fusion on the semantic feature vector and the facial feature vector to obtain a fused feature vector;
[0044] S400: Taking the fused feature vector as input, generating a psychological state assessment result of the target user through a psychological state assessment model.
[0045] In specific implementation, the psychological state analysis method of this embodiment achieves the technical effect of efficiently, comprehensively and accurately analyzing the psychological state of the user through multimodal data collection and fusion technology. Specifically, the method of this embodiment first obtains the question-answering audio and facial video of the target user when answering the target question (S100). These two types of data capture the psychological state information of the target user from the two dimensions of language expression and facial expression. Subsequently, by extracting the semantic feature vector of the audio and the facial feature vector of the video respectively (S200), key information can be extracted from the language content and facial expression. This multimodal feature extraction method ensures the comprehensiveness of the analysis. In S300, the semantic feature vector is fused with the facial feature vector to generate a fused feature vector. This fusion process can integrate the complementary information in the language and expression, thereby more accurately reflecting the user's true psychological state. Finally, the fused feature vector is analyzed using the psychological state evaluation model to generate a psychological state evaluation result (S400). This method based on multimodal data fusion not only improves the accuracy of psychological state analysis, but also improves the analysis efficiency through automated processing procedures, overcomes the limitations of traditional single modality analysis methods, and provides an efficient, comprehensive and accurate technical solution for psychological state analysis.
[0046] It can be understood that the psychological state analysis method provided in the embodiment of the present invention can be applied to psychological state analysis scenarios related to the medical and health field. The following is a specific example:
[0047] Example: Monitoring the mental status of patients with chronic diseases
[0048] In the healthcare field, the psychological state of patients with chronic diseases is crucial to their recovery process. Medical staff need to understand the patient's psychological state in a timely manner in order to provide personalized interventions. Traditional methods mainly rely on patients' self-reports, but these methods are subject to subjective bias and concealment.
[0049] The psychological state analysis method of the present invention is applied:
[0050] Scenario description:
[0051] During the patient's regular follow-up visit, the patient's audio and facial video of the questions answered by the patient are obtained through video calls or offline interviews (S100). These questions may involve the patient's condition, life pressure, and expectations for treatment.
[0052] Technical implementation:
[0053] The question-and-answer audio is semantically analyzed to extract the patient's language expression features (such as description of the condition, concerns about the future, etc.) and generate a semantic feature vector (S200).
[0054] Perform micro-expression analysis on the facial video, extract the patient's facial expression features (such as whether he is anxious, depressed, etc.), and generate a facial feature vector (S200).
[0055] The semantic feature vector is fused with the facial feature vector to generate a fused feature vector (S300).
[0056] The psychological state assessment model is used to analyze the patient's psychological state (such as anxiety level, depression level, etc.) and generate a psychological state assessment result (S400).
[0057] Technical effects:
[0058] Efficiency: Through automated analysis, psychological status assessment results can be quickly generated, saving medical staff’s time and energy.
[0059] Comprehensiveness: Combine language and facial expression information to comprehensively assess the patient's mental state and avoid the limitations of a single data dimension.
[0060] Accuracy: Accurately identify the patient's mental state, help medical staff detect potential psychological problems in a timely manner, and provide more effective intervention measures.
[0061] It can be understood that the psychological state analysis method provided in the embodiment of the present invention can also be applied to psychological state analysis scenarios related to the financial technology field. The following is a specific example:
[0062] Example: Customer credit assessment and risk management
[0063] In the field of financial technology, financial institutions need to accurately assess customers' credit risk and investment preferences in order to optimize credit decisions and investment recommendations. Traditional methods mainly rely on customers' financial data and credit records, but this information cannot fully reflect the customer's psychological state and may lead to evaluation bias.
[0064] The psychological state analysis method of the present invention is applied:
[0065] Scenario description:
[0066] When a customer applies for a loan or credit card, the financial institution obtains the customer's audio and facial video of the questions answered by the customer through a video call or offline interview (S100). These questions may involve the customer's financial situation, repayment plan, and future expectations.
[0067] Technical implementation:
[0068] The question-and-answer audio is semantically analyzed to extract the language expression features of the customer (such as speaking speed, intonation, word choice, etc.) and generate a semantic feature vector (S200).
[0069] Perform micro-expression analysis on the facial video, extract the customer's facial expression features (such as whether they are nervous, anxious, etc.), and generate a facial feature vector (S200).
[0070] The semantic feature vector is fused with the facial feature vector to generate a fused feature vector (S300).
[0071] The psychological state assessment model is used to analyze the psychological state of the customer (such as anxiety level, confidence level, etc.) and generate a psychological state assessment result (S400).
[0072] Technical effects:
[0073] Efficiency: Through automated analysis, psychological status assessment results can be quickly generated, saving the time and cost of manual assessment.
[0074] Comprehensiveness: Combine language and facial expression information to comprehensively assess the customer's psychological state and avoid the limitations of a single data dimension.
[0075] Precision: Accurately identify the customer's psychological state to help financial institutions better assess credit risk and reduce default rates.
[0076] Furthermore, in one embodiment, the psychological state analysis method, wherein the step S100, obtaining the question and answer audio and facial video of the target user when answering the target question, specifically comprises the steps of:
[0077] determining open-ended questions to be asked of the target user;
[0078] The question-and-answer audio and the facial video of the target user when answering the open-ended question are respectively collected by an audio collection device and a video collection device;
[0079] The question-and-answer audio and the facial video are time-synchronized.
[0080] In specific implementation, this embodiment can achieve the following technical effects by determining open-ended questions and collecting the audio and facial video of the target user answering the questions, and performing time synchronization processing at the same time:
[0081] First, the design of open-ended questions can guide users to express themselves freely, thereby obtaining richer and more realistic information about their mental states, thus avoiding the information limitations that may be brought about by traditional closed-ended questions;
[0082] Secondly, the synchronous collection of audio and video not only captures the user's language content, but also records their facial expressions and non-verbal information, providing a comprehensive data basis for subsequent multimodal analysis;
[0083] Finally, time synchronization ensures the consistency of audio and video data in the time dimension, making subsequent feature extraction and fusion more accurate, thereby improving the accuracy and reliability of psychological state analysis. This process provides solid technical support for efficient, comprehensive and accurate psychological state analysis, especially for in-depth assessment of psychological state in the fields of financial technology and medical health.
[0084] The specific implementation process of the steps in this embodiment is roughly as follows:
[0085] (1) Identify open-ended questions
[0086] Question design:
[0087] Design a series of open-ended questions based on the goals and application scenarios of psychological state analysis. These questions should guide users to freely express their thoughts, feelings, or experiences, and avoid limiting the scope of users' answers.
[0088] Example:
[0089] “Please describe a recent stressful situation in which you experienced stress.”
[0090] “What are your hopes and concerns for the future?”
[0091] Problem verification:
[0092] Pre-test the designed open-ended questions to ensure that they can effectively guide user expression and will not cause misunderstanding or ambiguity. You can invite a small number of target users to test answer the questions and adjust the question wording based on the feedback.
[0093] (2) Prepare the acquisition equipment
[0094] Select the audio capture device:
[0095] Choose a high-quality audio capture device, such as a professional microphone or recording device, to ensure that the user's voice can be captured clearly. The device should have a high sampling rate (such as 44.1kHz or higher) and low noise characteristics.
[0096] Select the video capture device:
[0097] Choose a high-definition video capture device, such as a webcam or smartphone camera, to ensure that the user's facial expressions can be captured clearly. The device should have a high resolution (such as 1080p or higher) and a stable frame rate (such as 30fps or higher).
[0098] Equipment Calibration:
[0099] Calibrate audio and video equipment before acquisition to ensure time synchronization between the two. You can use synchronization signals or timestamp technology to ensure the time consistency of audio and video data.
[0100] (3) Collecting Q&A audio and facial video
[0101] Guide users to answer:
[0102] Clearly explain the content and answer requirements of open-ended questions to target users to ensure that users understand the questions and can express themselves freely. During the collection process, avoid guiding or interrupting users' answers to obtain the most authentic information.
[0103] Audio Collection:
[0104] Use an audio capture device to record the user's voice when answering questions. Make sure the recording environment is quiet and reduce background noise interference.
[0105] Video Collection:
[0106] Use a video capture device to record the user's facial expressions as they answer questions. Make sure the camera is pointed at the user's face and that there is sufficient lighting to clearly capture micro-expressions.
[0107] (4) Time synchronization processing
[0108] Data alignment:
[0109] Time-align the captured audio and video data. You can use timestamps or synchronization signals to ensure the temporal consistency of audio and video. For example, use professional software to align audio and video data to ensure that the time axes of the two are completely matched.
[0110] Data verification:
[0111] Check the time synchronization of audio and video data to ensure that there is no deviation in time between the two. If a time deviation is found, realign the data until satisfactory synchronization is achieved.
[0112] (5) Data preservation and preliminary processing
[0113] Data Retention:
[0114] The collected audio and video data are saved separately to secure storage media, and clear identification is added to each file for subsequent processing and analysis.
[0115] Initial processing:
[0116] Perform preliminary processing on audio and video data, such as removing background noise, adjusting volume, cropping unnecessary parts, etc., to ensure that the data quality meets the requirements of subsequent analysis.
[0117] Through the above process, this embodiment can efficiently obtain the target user's answer audio and facial video when answering open-ended questions, and ensure the temporal consistency of the two. This process provides a high-quality multimodal data foundation for subsequent psychological state analysis, ensuring the comprehensiveness and accuracy of the analysis.
[0118] Example description:
[0119] In the field of medical health, the psychological state analysis method of this embodiment can be applied to the psychological state monitoring of patients with chronic diseases to help medical staff better understand the patient's inner feelings and emotional state, thereby providing more accurate intervention and treatment.
[0120] Implementation process
[0121] 1) Identify open-ended questions
[0122] Question design: Based on the mental health needs of patients with chronic diseases, a series of open-ended questions are designed to guide patients to express their feelings about their illness, life pressures, and expectations for the future.
[0123] For example:
[0124] "How have you felt about your illness recently?"
[0125] “What was the biggest psychological challenge you faced during treatment?”
[0126] “What are your expectations and concerns about your future life?”
[0127] 2) Collect Q&A audio and facial video
[0128] Equipment preparation: Provide a quiet and comfortable environment for patients in a hospital or community medical center. Use a high-definition camera and a professional microphone to collect facial video and audio of patients answering questions.
[0129] Data collection: Medical staff guide patients to answer designed open-ended questions and record their answers through audio and video equipment. Ensure that external interference is avoided during the collection process to obtain high-quality data.
[0130] 3) Time synchronization processing
[0131] Data alignment: Use professional software to synchronize the captured audio and video data to ensure that they are completely aligned on the timeline. For example, use timestamp technology or synchronization signals to calibrate the time deviation of audio and video data.
[0132] Data verification: Check the synchronized data to ensure the consistency of audio and video in time. If deviation is found, readjust the data alignment until a satisfactory synchronization effect is achieved.
[0133] Technical Effects
[0134] Comprehensiveness: Through multimodal data collection of audio and video, not only the patient's language content is captured, but also their facial expressions and non-verbal information are recorded, providing a more comprehensive data basis for psychological state analysis.
[0135] Objectivity: Open-ended questions guide patients to express themselves freely, avoiding the subjectivity of traditional scales. At the same time, the objective recording of audio and video data reduces the bias of subjective judgment of medical staff.
[0136] Timeliness: Automated time synchronization processing and subsequent analysis processes can quickly generate psychological status assessment results, helping medical staff to promptly identify patients' psychological problems and provide personalized intervention measures.
[0137] Personalized intervention: Psychological state analysis based on multimodal data can more accurately identify patients' psychological needs. Medical staff can develop personalized psychological intervention plans based on the assessment results, such as psychological counseling, emotion management training, or referral to professional psychologists.
[0138] Application Scenario
[0139] Hospital outpatient clinics: During regular follow-up visits for patients with chronic diseases, psychological state analysis helps medical staff better understand the patient's psychological state and adjust the treatment plan in a timely manner.
[0140] Community Medical Center: Conduct mental health screening at the community level to detect potential psychological problems at an early stage and provide timely psychological intervention.
[0141] Telemedicine: Collect data through video calls to provide remote mental status monitoring services for patients with limited mobility.
[0142] The psychological state analysis method of this embodiment guides patients to express their inner feelings through open-ended questions, and combined with multimodal data collection and time synchronization processing, it can comprehensively and objectively evaluate the psychological state of patients with chronic diseases. This method not only improves the efficiency and accuracy of psychological state monitoring, but also provides medical staff with personalized intervention basis, which helps to improve the treatment effect and quality of life of patients.
[0143] Furthermore, in one embodiment, the psychological state analysis method, wherein the step S200, respectively performing feature representation on the question-answer audio and the facial video to obtain corresponding semantic feature vectors and facial feature vectors, specifically comprises the steps of:
[0144] Converting the question-and-answer audio into question-and-answer text using automatic speech recognition technology, inputting the question-and-answer text into a pre-trained text encoding model, and generating the semantic feature vector;
[0145] The facial video is input into a pre-trained video coding model to generate the facial feature vector.
[0146] In specific implementation, this embodiment can efficiently extract key information reflecting the user's psychological state by representing the features of the question-and-answer audio and facial video respectively. Specifically, the question-and-answer audio is converted into question-and-answer text using automatic speech recognition technology, and a semantic feature vector is generated through a pre-trained text encoding model. This process can accurately capture the user's language content and emotional tendencies, and convert complex voice information into quantifiable semantic features. At the same time, the facial video is input into the pre-trained video encoding model to generate a facial feature vector, which can extract subtle changes in the user's facial expression, such as micro-expressions and emotional states. This multimodal feature extraction method not only makes full use of the complementarity of language and visual information, but also ensures the accuracy and efficiency of feature extraction through the efficient processing of the pre-trained model. Ultimately, the generated semantic feature vectors and facial feature vectors provide a comprehensive and accurate data basis for subsequent psychological state assessment, significantly improving the reliability and depth of psychological state analysis, so that it can be widely used in complex scenarios in fields such as financial technology and medical health.
[0147] The specific implementation process of the steps in this embodiment is roughly as follows:
[0148] (1) Audio processing and semantic feature extraction
[0149] Voice Recognition:
[0150] Use automatic speech recognition (ASR) technology to convert the collected Q&A audio into Q&A text. ASR technology can convert the language content in the speech signal into a processable text format. Select a high-precision ASR model (such as a deep learning-based model) to ensure the accuracy of the transcription.
[0151] Technical details: You can use open source ASR tools or customized ASR models to ensure adaptability to different accents and speaking speeds.
[0152] Text encoding:
[0153] The converted question-answer text is input into the pre-trained text encoding model to extract the semantic features of the text. The text encoding model can convert the text content into a high-dimensional feature vector that reflects the semantic information and emotional tendency of the language.
[0154] Technical details: Select a pre-trained model suitable for mental state analysis and fine-tune it as needed to better adapt to the language expression of specific fields.
[0155] Generate semantic feature vector:
[0156] The feature vectors output by the text encoding model are semantic feature vectors. These vectors can quantitatively represent the language content and emotional tendency in the user's answer, and provide language dimension information for subsequent psychological state analysis.
[0157] (2) Video processing and facial feature extraction
[0158] Video input and preprocessing:
[0159] The collected facial video is input into the pre-trained video coding model. Before input, the video is pre-processed, including cropping, denoising, and frame rate adjustment, to ensure that the video quality meets the model input requirements.
[0160] Technical details: You can use video processing tools to pre-process the video to ensure that the resolution and frame rate of the video are consistent.
[0161] Facial feature extraction:
[0162] Each frame in the video is analyzed using a pre-trained video coding model (such as a deep learning-based CNN or Transformer model) to extract facial expression features. These models are able to recognize micro-expressions, emotional states, and other facial movements.
[0163] Technical details: Choose a pre-trained model suitable for facial expression analysis to ensure that it can accurately capture subtle facial changes.
[0164] Generate facial feature vectors:
[0165] The feature vectors output by the video coding model are facial feature vectors. These vectors can quantitatively represent the facial expressions and emotional states of users when answering, and provide visual dimension information for subsequent psychological state analysis.
[0166] Through the above process, this embodiment can efficiently extract semantic feature vectors and facial feature vectors from question-answer audio and facial video. Through normalization and integration, these feature vectors provide a high-quality multimodal data foundation for subsequent psychological state analysis, significantly improving the comprehensiveness and accuracy of the analysis.
[0167] Furthermore, in one embodiment, the psychological state analysis method, wherein the step S300, fusing the semantic feature vector with the facial feature vector to obtain a fused feature vector, specifically comprises the steps of:
[0168] Normalizing the semantic feature vector and the facial feature vector;
[0169] According to a preset fusion strategy, the normalized semantic feature vector and the facial feature vector are subjected to feature fusion to obtain a fused preliminary feature vector;
[0170] The preliminary feature vector is verified, and when the verification result meets the preset requirement, the preliminary feature vector is used as the fused feature vector.
[0171] In specific implementation, this embodiment realizes the effective integration of semantic feature vectors and facial feature vectors through normalization, feature fusion and verification process, thereby generating a high-quality fused feature vector. Specifically, the normalization process ensures the consistency of the numerical range and distribution of the two modal feature vectors, eliminates the deviation caused by the dimension difference, and provides a unified basis for subsequent fusion. Subsequently, the normalized feature vectors are fused according to the preset fusion strategy (such as weighted average, splicing or deep learning fusion method) to generate a preliminary feature vector. This process makes full use of the complementarity of language and visual information and enhances the expressive power of features. Finally, the preliminary feature vector is quality checked through the verification step, and only when it meets the preset requirements (such as feature integrity, validity or matching with a known psychological state) is it used as the final fused feature vector. This verification mechanism further ensures the reliability and accuracy of the fused feature vector, provides high-quality data support for the input of the subsequent psychological state assessment model, and thus significantly improves the accuracy and stability of psychological state analysis.
[0172] The specific implementation process of the steps in this embodiment is roughly as follows:
[0173] (1) Normalization of feature vectors
[0174] Normalization method selection:
[0175] Select an appropriate normalization method (such as Min-Max normalization or Z-score normalization) to process the semantic feature vector and facial feature vector. The goal of normalization is to adjust the numerical range of the feature vector to a uniform interval (such as [0,1]) or to standardize the distribution of the feature vector (mean is 0, variance is 1).
[0176] Semantic feature vector normalization:
[0177] Each dimension in the semantic feature vector is normalized.
[0178] Normalization of facial feature vector:
[0179] Similarly, each dimension in the facial feature vector is normalized to ensure that its value range is consistent with the semantic feature vector. The normalized feature vector can eliminate the dimensional differences between different modal data and provide a unified data basis for subsequent fusion.
[0180] (2) Feature Fusion
[0181] Select the fusion strategy:
[0182] Choose the appropriate fusion strategy based on the application scenario and requirements. Common fusion strategies include weighted averaging, feature concatenation, or deep learning-based fusion methods (such as learning the optimal fusion method through neural networks).
[0183] Weighted average fusion:
[0184] If weighted average fusion is selected, weights (such as α and 1-α) are assigned according to the importance of semantic features and facial features, and the fused preliminary feature vector is calculated, where α is the weight of the semantic feature and can be adjusted according to experimental results or domain knowledge.
[0185] Feature splicing and fusion:
[0186] If feature concatenation and fusion is selected, the normalized semantic feature vector and facial feature vector are directly concatenated into a longer vector.
[0187] Fusion based on deep learning:
[0188] If you choose a fusion method based on deep learning, you can input the normalized feature vector into a pre-trained neural network model (such as a fully connected layer or Transformer) to let the model automatically learn the optimal fusion method.
[0189] (3) Verification of preliminary feature vector
[0190] Verification indicator settings:
[0191] Set indicators to verify the preliminary feature vector, such as feature completeness, validity, or matching with known psychological states. These indicators can be adjusted according to application scenarios and requirements.
[0192] Verification process:
[0193] The preliminary feature vector is evaluated using pre-set validation indicators, for example, checking whether the preliminary feature vector contains all necessary dimensions, or verifying its validity by comparing it with known psychological state data.
[0194] Result judgment and adjustment:
[0195] If the verification result meets the preset requirements, the preliminary feature vector is used as the final fused feature vector; if it does not meet the preset requirements, the fusion strategy is adjusted according to the verification result, the preliminary feature vector is regenerated and verified again until the final fused feature vector is obtained.
[0196] (4) Output of fused feature vector
[0197] Store the fused feature vector:
[0198] The verified fused feature vector is stored in the specified data structure to ensure that its format meets the input requirements of the subsequent psychological state assessment model.
[0199] Marking and recording:
[0200] The fused feature vector is marked, and its source (such as the weight distribution of semantic features and facial features) and verification results are recorded for subsequent analysis and tracing.
[0201] Through the above process, this embodiment can efficiently fuse the semantic feature vector and the facial feature vector. The normalization process ensures the consistency of data from different modalities, the selection and implementation of the fusion strategy realizes the integration of multimodal information, and the verification step ensures the quality and reliability of the fused feature vector. The fused feature vector finally generated provides high-quality data support for subsequent psychological state assessment, significantly improving the comprehensiveness and accuracy of the analysis.
[0202] Furthermore, in one embodiment, the psychological state analysis method, wherein the step S400, taking the fused feature vector as input and generating the psychological state evaluation result of the target user through the psychological state evaluation model, specifically comprises the steps of:
[0203] Inputting the fused feature vector into the pre-trained emotional state assessment model and the psychological state assessment model respectively to generate corresponding emotional state assessment results and the psychological state assessment results;
[0204] The emotional state assessment result and the psychological state assessment result are integrated by weighted average method to obtain a multi-dimensional state assessment result of the target user.
[0205] Furthermore, the psychological state analysis method, wherein, after the emotional state evaluation result and the psychological state evaluation result are integrated by weighted average method to obtain the multi-dimensional state evaluation result of the target user, specifically further comprises the steps of:
[0206] Based on the multi-dimensional state assessment result, identifying risk factors of the target user in terms of emotional state and psychological state;
[0207] A psychological state adjustment plan for the target user is formulated according to the risk factors.
[0208] Furthermore, the psychological state analysis method, wherein the formulating of a psychological state adjustment plan for the target user according to the risk factors specifically comprises the steps of:
[0209] Classifying and ranking the risk factors to obtain classification and ranking results;
[0210] The psychological state adjustment plan is formulated according to the classification and sorting results, and the psychological state adjustment plan is fed back to the target user terminal.
[0211] In specific implementation, this embodiment achieves comprehensive, accurate and personalized analysis and adjustment of the psychological state of the target user through hierarchical model evaluation, multi-dimensional result integration, and targeted risk identification and intervention plan formulation. Specifically, the fusion feature vector is analyzed respectively using the pre-trained emotional state evaluation model and the psychological state evaluation model to generate the evaluation results of the emotional state and the psychological state. This dual-model evaluation method can capture the psychological characteristics of the user from different angles to ensure the comprehensiveness of the analysis. Subsequently, the two evaluation results are integrated by the weighted average method to generate a multi-dimensional state evaluation result, which further improves the accuracy of the analysis. On this basis, the method of this embodiment further identifies the risk factors of the user in terms of emotion and psychological state, and formulates a personalized psychological state adjustment plan based on the classification and sorting of risk factors. This process not only ensures the pertinence of the intervention measures, but also realizes closed-loop management from evaluation to intervention by feeding back the plan to the user. Overall, the method of this embodiment can efficiently identify psychological state problems and provide scientific and personalized adjustment suggestions. It is suitable for complex application scenarios in the fields of financial technology and medical health, and significantly improves the practicality and effect of psychological state analysis.
[0212] The specific implementation process of the steps in this embodiment is roughly as follows:
[0213] (1) Multi-model evaluation
[0214] Input fused feature vector:
[0215] The normalized and fused feature vectors are input into the pre-trained emotional state assessment model and psychological state assessment model respectively. The emotional state assessment model focuses on analyzing the user's emotional characteristics (such as anxiety, depression, mood swings), while the psychological state assessment model focuses more on assessing the user's overall psychological condition (such as stress level, concentration, psychological resilience, etc.).
[0216] Generate evaluation results:
[0217] The emotional state assessment model and the psychological state assessment model output emotional state assessment results and psychological state assessment results respectively. These results can be numerical scores, classification labels or probability distributions, depending on the design of the model.
[0218] (2) Integration of multi-dimensional status assessment results
[0219] Weighted average integration:
[0220] According to the preset weights, the emotional state assessment results and the psychological state assessment results are integrated through the weighted average method. The weights can be adjusted according to the reliability of the model and the importance of the application scenario. For example, if the emotional state has a greater impact on the target user, a higher weight is given to the emotional state assessment result.
[0221] Generate multi-dimensional condition assessment results:
[0222] The integrated result is the multi-dimensional status assessment result of the target user, which comprehensively reflects the user's emotional and psychological state. This result provides a basis for the subsequent risk factor identification and adjustment plan formulation.
[0223] (3) Identification of risk factors
[0224] Analyze multi-dimensional condition assessment results:
[0225] Conduct in-depth analysis of the multidimensional state assessment results to identify possible risk factors for the user in terms of emotional state and psychological state. For example, a high anxiety score in the emotional state assessment results or a low concentration score in the psychological state assessment results may be considered a risk factor.
[0226] Risk factor classification and ranking:
[0227] Risk factors are classified and ranked according to their nature (e.g., emotional problems, psychological problems) and severity (e.g., mild, moderate, severe). The purpose of ranking is to determine which risk factors require priority intervention.
[0228] (4) Development of a plan to adjust the mental state
[0229] Develop a personalized adjustment plan:
[0230] According to the classification and sorting results, a personalized mental state adjustment plan is formulated for the user. The plan may include specific intervention measures, such as relaxation training, emotion regulation techniques, lifestyle adjustments, etc.
[0231] Program feedback:
[0232] Feedback of the mental state adjustment plan to the terminal device of the target user. Feedback can be achieved in a variety of ways, such as generating a detailed adjustment suggestion report, providing online guidance, or communicating with professionals.
[0233] (5) Plan implementation and tracking (optional)
[0234] Implementation of adjustment plan:
[0235] Assist users to start implementing a plan to adjust their mental state and provide necessary support and resources. For example, provide a tutorial on relaxation training, recommend psychological counseling services, or suggest adjusting their work and rest schedule.
[0236] Tracking and evaluation:
[0237] Regularly track changes in the user's mental state and evaluate the effectiveness of the mental state adjustment plan. If the user's mental state has not improved significantly, the content of the plan can be adjusted based on feedback.
[0238] Through the above process, this embodiment can efficiently generate the multi-dimensional status assessment results of the target user, identify risk factors based on the multi-dimensional status assessment results, and formulate and feedback personalized psychological state adjustment plans. This process not only ensures the comprehensiveness and accuracy of the assessment, but also improves the effect of psychological state adjustment through personalized intervention measures, and is suitable for complex application scenarios in the fields of financial technology and medical health.
[0239] It can be seen from the above method embodiments that the psychological state analysis method provided by the present invention includes: obtaining the question and answer audio and facial video of the target user answering the target question; performing feature representation on the question and answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors; performing feature fusion on the semantic feature vector and the facial feature vector to obtain a fused feature vector; using the fused feature vector as input, and generating the psychological state evaluation result of the target user through the psychological state evaluation model. In this way, the method of the present invention can realize efficient, comprehensive and accurate analysis of the user's psychological state.
[0240] It should be understood that, although the present application provides method operation steps as described in the embodiments or flowcharts, more or less operation steps may be included based on conventional or non-creative labor, and these operation steps are not necessarily performed in sequence according to the embodiment or flowchart. The order of steps listed in the embodiment or flowchart is only one way of executing the order of many steps, and does not represent the only execution order. It should be noted that there is not necessarily a certain order between the above steps. A person of ordinary skill in the art can understand from the description of the embodiment of the present invention that in different embodiments, the above steps may have different execution orders, that is, they may be executed in parallel, or they may be executed in exchange, etc. Moreover, at least a part of the steps in the embodiment or flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but may be executed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but may be executed in turn, alternately or synchronously with other steps or at least a part of the sub-steps or stages of other steps.
[0241] Based on the above method embodiment, please refer to Figure 3 Another embodiment of the present invention further provides a psychological state analysis device, wherein the device comprises:
[0242] An acquisition module 11 is used to acquire the question-answering audio and facial video of the target user when answering the target question;
[0243] A feature representation module 12 is used to perform feature representation on the question-answer audio and the facial video to obtain corresponding semantic feature vectors and facial feature vectors;
[0244] A feature fusion module 13 is used to perform feature fusion on the semantic feature vector and the facial feature vector to obtain a fused feature vector;
[0245] The result generating module 14 is used to take the fused feature vector as input and generate the psychological state evaluation result of the target user through the psychological state evaluation model.
[0246] Furthermore, in one embodiment, in the mental state analysis device, the acquisition module 11 is specifically used for:
[0247] determining open-ended questions to be asked of the target user;
[0248] The question-and-answer audio and the facial video of the target user when answering the open-ended question are respectively collected by an audio collection device and a video collection device;
[0249] The question-and-answer audio and the facial video are time-synchronized.
[0250] Furthermore, in one embodiment, in the mental state analysis device, the feature representation module 12 is specifically used for:
[0251] Using automatic speech recognition technology to convert the question-answer audio into question-answer text, inputting the question-answer text into a pre-trained text encoding model to generate the semantic feature vector;
[0252] The facial video is input into a pre-trained video coding model to generate the facial feature vector.
[0253] Furthermore, in one embodiment, in the mental state analysis device, the feature fusion module 13 is specifically used for:
[0254] Normalizing the semantic feature vector and the facial feature vector;
[0255] According to a preset fusion strategy, the normalized semantic feature vector and the facial feature vector are subjected to feature fusion to obtain a fused preliminary feature vector;
[0256] The preliminary feature vector is verified, and when the verification result meets the preset requirement, the preliminary feature vector is used as the fused feature vector.
[0257] Furthermore, in one embodiment, in the mental state analysis device, the result generation module 14 is specifically used for:
[0258] Inputting the fused feature vector into the pre-trained emotional state assessment model and the psychological state assessment model respectively to generate corresponding emotional state assessment results and the psychological state assessment results;
[0259] The emotional state assessment result and the psychological state assessment result are integrated by weighted average method to obtain a multi-dimensional state assessment result of the target user.
[0260] Furthermore, the mental state analysis device, wherein after the emotional state evaluation result and the mental state evaluation result are integrated by weighted average method to obtain the multi-dimensional state evaluation result of the target user, specifically includes:
[0261] Based on the multi-dimensional state assessment result, identifying risk factors of the target user in terms of emotional state and psychological state;
[0262] A psychological state adjustment plan for the target user is formulated according to the risk factors.
[0263] Furthermore, the psychological state analysis device, wherein the formulating of a psychological state adjustment plan for the target user according to the risk factors specifically includes:
[0264] Classifying and ranking the risk factors to obtain classification and ranking results;
[0265] The psychological state adjustment plan is formulated according to the classification and sorting results, and the psychological state adjustment plan is fed back to the target user terminal.
[0266] It should be noted that in the embodiment of the device of the present invention, the information interaction, execution process and other contents between the above-mentioned modules are based on the same concept as the embodiment of the method of the present invention. Their specific functions and technical effects can be found in the aforementioned method embodiment part and will not be repeated here.
[0267] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the functions or steps of the server side of the psychological state analysis method in any of the above method embodiments are implemented.
[0268] Based on the above method embodiment, another embodiment of the present invention further provides a computer device, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the functions or steps of the client side of the psychological state analysis method in any of the above method embodiments are implemented.
[0269] Those skilled in the art will understand that Figure 4 and Figure 5 The structural schematic diagram shown in the figure is only a schematic diagram of a partial structure related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0270] The processor may be a CPU, or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, etc.
[0271] Among them, the memory includes a readable storage medium, an internal memory, etc., wherein the internal memory can be the memory of a computer device, and the internal memory provides an environment for the operation of an operating system and computer-readable instructions in the readable storage medium. The readable storage medium can be a hard disk of a computer device, and in other embodiments, it can also be an external storage device of a computer device, for example, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on a computer device. Further, the memory can also include both an internal storage unit of a computer device and an external storage device. The memory is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of a computer program, etc. The memory can also be used to temporarily store data that has been output or is to be output.
[0272] Based on the above method embodiments, another embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the psychological state analysis method in any of the above method embodiments is implemented. The computer-readable storage medium may be non-volatile or volatile.
[0273] It should be noted that the above-mentioned functions or steps that can be implemented by the computer-readable storage medium or computer device, and the technical effects brought about by the functions / steps, can be found in the relevant descriptions in the aforementioned method embodiments. To avoid repetition, they will not be described one by one here.
[0274] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM). The disclosed memory components or memories of the operating environments described herein are intended to comprise one or more of these and / or any other suitable types of memory.
[0275] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, in the embodiment of the device of the present invention, only the division of the above-mentioned functional units and modules is used as an example. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the above-mentioned method embodiment, which will not be repeated here. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0276] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0277] In the embodiments provided by the present invention, it should be understood that the disclosed devices / computer equipment and methods can be implemented in other ways. For example, the device / computer equipment embodiments described above are only schematic, for example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0278] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0279] It should be noted that if software tools or components other than those of the Company appear in the embodiments of the present application, they are only used for illustration and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the above embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention.
Claims
1. A method for analyzing a psychological state, characterized in that: include: Obtain the target user’s audio and facial video when answering the target question; Performing feature representation on the question-answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors; Performing feature fusion on the semantic feature vector and the facial feature vector to obtain a fused feature vector; The fused feature vector is used as input to generate a psychological state assessment result of the target user through a psychological state assessment model.
2. The psychological state analysis method according to claim 1, characterized in that: The step of obtaining the question-answering audio and facial video of the target user when answering the target question includes: determining open-ended questions to be asked of the target user; The question-and-answer audio and the facial video of the target user when answering the open-ended question are respectively collected by an audio collection device and a video collection device; The question-and-answer audio and the facial video are time-synchronized.
3. The psychological state analysis method according to claim 1, characterized in that: The step of performing feature representation on the question-answer audio and the facial video to obtain corresponding semantic feature vectors and facial feature vectors includes: Using automatic speech recognition technology to convert the question-answer audio into question-answer text, inputting the question-answer text into a pre-trained text encoding model to generate the semantic feature vector; The facial video is input into a pre-trained video coding model to generate the facial feature vector.
4. The psychological state analysis method according to claim 1, characterized in that: The step of fusing the semantic feature vector with the facial feature vector to obtain a fused feature vector comprises: Normalizing the semantic feature vector and the facial feature vector; According to a preset fusion strategy, the normalized semantic feature vector and the facial feature vector are subjected to feature fusion to obtain a fused preliminary feature vector; The preliminary feature vector is verified, and when the verification result meets the preset requirement, the preliminary feature vector is used as the fused feature vector.
5. The psychological state analysis method according to any one of claims 1 to 4, characterized in that: The step of taking the fused feature vector as input and generating a psychological state evaluation result of the target user through a psychological state evaluation model includes: Inputting the fused feature vector into the pre-trained emotional state assessment model and the psychological state assessment model respectively to generate corresponding emotional state assessment results and the psychological state assessment results; The emotional state assessment result and the psychological state assessment result are integrated by weighted average method to obtain a multi-dimensional state assessment result of the target user.
6. The method for analyzing psychological state according to claim 5, characterized in that: After the emotional state evaluation result and the psychological state evaluation result are integrated by weighted average method to obtain the multi-dimensional state evaluation result of the target user, the method further includes: Based on the multi-dimensional state assessment result, identifying risk factors of the target user in terms of emotional state and psychological state; A psychological state adjustment plan for the target user is formulated according to the risk factors.
7. The psychological state analysis method according to claim 6, characterized in that: The formulating a psychological state adjustment plan for the target user according to the risk factors includes: Classifying and ranking the risk factors to obtain classification and ranking results; The psychological state adjustment plan is formulated according to the classification and sorting results, and the psychological state adjustment plan is fed back to the target user terminal.
8. A psychological state analysis device, characterized in that: include: An acquisition module is used to acquire the question-answering audio and facial video of the target user when answering the target question; A feature representation module, used to perform feature representation on the question-answer audio and the facial video respectively to obtain corresponding semantic feature vectors and facial feature vectors; A feature fusion module, used for fusing the semantic feature vector with the facial feature vector to obtain a fused feature vector; The result generating module is used to take the fused feature vector as input and generate the psychological state evaluation result of the target user through the psychological state evaluation model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the psychological state analysis method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the psychological state analysis method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Psychological state analysis method and apparatus, computer device, and storage medium
WO2026174863A1