Psychological risk assessment method and device

Through the combination of multimodal data analysis and pre-trained large language model, psychological risk assessment results are generated, and the problem of relying on a single data source and simple analysis model in the existing technology is solved, achieving a more accurate and personalized psychological risk assessment.

CN119943399APending Publication Date: 2025-05-06CHINA TELECOM CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510052074.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing sensorless psychological analysis techniques rely on a single data source and simple analytical models, resulting in a lack of accuracy and in-depth analysis results, and the inability to fully capture the individual's psychological state in a specific environment.

Method used

By obtaining the multimodal target state data (visual mode and sound mode) of the target object, combined with a pre-trained large language model, the target emotional characteristics and archival information are analyzed to generate psychological risk assessment results, including psychological risk types, analysis basis and coping strategies.

Benefits of technology

It improves the comprehensiveness and accuracy of emotional recognition and psychological state assessment, provides highly personalized evaluation results, can identify psychological risk types in specific contexts, and provides analysis basis and coping strategies that are more in line with personal needs, improving the pertinence and effectiveness of psychological interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943399A_ABST
    Figure CN119943399A_ABST
Patent Text Reader

Abstract

The invention discloses a psychological risk assessment method and device. The method comprises the steps that multi-mode target state data of a target object are acquired, and the modes of the target state data at least comprise a visual mode and a sound mode; determining a target emotional feature of the target object according to the multi-modal target state data; obtaining target archive information associated with the target object; the target emotional features and the target archive information are analyzed through a pre-trained target large language model, a psychological risk assessment result of the target object is obtained, and the psychological risk assessment result at least comprises a psychological risk type, an analysis basis and a coping strategy. The technical problems that an existing non-inductive psychological analysis technology depends on a single data source and a simple analysis model, and the analysis result is lack of accuracy and deepness are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a psychological risk assessment method and device. Background Art

[0002] In scenarios such as community corrections, security interrogations and inquiries, psychological assessment and monitoring are important links to ensure safety, improve efficiency and promote individual transformation. However, existing non-intuitive psychological assessment technologies face a series of challenges that limit their accuracy and reliability in practical applications. First, these technologies often rely on a single data source, such as analyzing only facial expressions or voice information, which cannot fully capture the psychological state of an individual in a specific environment, especially in a complex and changeable judicial and law enforcement environment. Secondly, the analytical models used, such as decision trees and discrete mathematics, are simple and easy to use, but lack deep learning capabilities and cannot deeply explore the intrinsic connections between multimodal data and the deep-level characteristics of an individual's psychological state, resulting in superficial and inaccurate assessment results. Furthermore, the lack of explicit inquiry process and basis has questioned the interpretability and scientific nature of psychological state assessment. Finally, existing technologies are also insufficient in providing customized suggestions, and it is difficult to provide managers with specific and targeted work suggestions or strategies based on psychological assessment results, which limits their application value in judicial decision-making, educational transformation and other aspects.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] The embodiments of the present application provide a psychological risk assessment method and device to at least solve the technical problems that the existing non-sensitive psychological analysis technology relies on a single data source and a simple analysis model, and the analysis results lack accuracy and depth.

[0005] According to one aspect of an embodiment of the present application, a psychological risk assessment method is provided, comprising: obtaining multimodal target state data of a target object, wherein the modalities of the target state data include at least a visual modality and a sound modality; determining a target emotional feature of the target object based on the multimodal target state data; obtaining target profile information associated with the target object; and analyzing the target emotional feature and the target profile information using a pre-trained target large language model to obtain a psychological risk assessment result of the target object, wherein the psychological risk assessment result includes at least a psychological risk type, an analysis basis, and a coping strategy.

[0006] Optionally, obtaining multimodal target state data of the target object includes: obtaining image data including the face and / or body of the target object; obtaining voice data of the communication process with the target object; and analyzing the voice data using voice recognition technology to obtain text data corresponding to the voice data.

[0007] Optionally, the target emotional features of the target object are determined based on the multimodal target state data, including: respectively determining the image features, voice features and text features corresponding to the image data, voice data and text data; performing weighted calculation on the image features, voice features and text features using a preset weight coefficient to obtain a multimodal fusion feature; analyzing the multimodal fusion feature using a valence arousal model, and using the emotional values ​​of the two dimensions output by the valence arousal model as the target emotional features of the target object.

[0008] Optionally, the image features, speech features and text features corresponding to the image data, speech data and text data are determined respectively, including: using a pre-trained image recognition model to analyze the image data to obtain image features, wherein the image features include at least one of the following: expression features, eye movement features, posture features; using a pre-trained speech recognition model to analyze the speech data to obtain speech features, wherein the speech features include at least one of the following: tone features, intonation features; using a pre-trained text recognition model to analyze the text data to obtain text features, wherein the text features include at least one of the following: context features, emotional tendency features.

[0009] Optionally, the training process of the target large language model includes: obtaining multiple groups of multimodal historical status data and historical behavior data corresponding to multiple objects, and obtaining historical archive information of each object; for each object, determining the emotional characteristics of the object based on the multimodal historical status data corresponding to the object, using the emotional characteristics and the historical archive information of the object as a training sample, and determining the psychological risk type of the object based on the historical behavior data of the object, and using the psychological risk type as a sample label of the training sample; determining an initial large language model based on verification rules, wherein the verification rules are used to construct a mapping relationship between the object's emotional characteristics, historical archive information, psychological risk type, analysis basis and coping strategies, and the analysis basis and coping strategies are determined based on a preset sociological and psychological knowledge graph; and iteratively training the initial large language model using multiple training samples and sample labels to obtain a target large language model.

[0010] Optionally, the verification rules are also used to guide the target large language model to generate a psychological risk assessment report in JSON format, wherein the psychological risk assessment report includes at least: report title, psychological risk type, risk description, analysis basis and response strategy, and the response strategy includes at least one of the following: psychological intervention strategy, environmental intervention strategy.

[0011] Optionally, the method also includes: obtaining multimodal status data and behavior data of the target object after executing the coping strategy, and determining the degree of improvement of the target object's psychological risk based on the multimodal status data and behavior data; when the target object's psychological risk improvement degree does not reach a preset standard, updating the coping strategy corresponding to the target emotional characteristics and target profile information in the verification rules of the target large language model based on the degree of improvement of the psychological risk.

[0012] According to another aspect of an embodiment of the present application, a psychological risk assessment device is also provided, including: a first acquisition module, for acquiring multimodal target state data of a target object, wherein the modalities of the target state data include at least: a visual modality and a sound modality; a feature extraction module, for determining the target emotional features of the target object based on the multimodal target state data; a second acquisition module, for acquiring target profile information associated with the target object; a psychological assessment module, for analyzing the target emotional features and target profile information using a pre-trained target large language model to obtain a psychological risk assessment result of the target object, wherein the psychological risk assessment result includes at least: a psychological risk type, an analysis basis, and a coping strategy.

[0013] According to another aspect of an embodiment of the present application, a computer program product is also provided, the computer program product comprising: a computer program, wherein the computer program implements the above-mentioned psychological risk assessment method when executed by a processor.

[0014] According to another aspect of an embodiment of the present application, an electronic device is further provided, comprising: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the above-mentioned psychological risk assessment method through the computer program.

[0015] In an embodiment of the present application, obtaining multimodal target state data of the target object (including but not limited to visual modality and sound modality) can capture the emotional state of the individual from multiple dimensions, avoiding the deviations and limitations that may be caused by single modality data. This multimodal fusion method improves the comprehensiveness and accuracy of emotion recognition and psychological state assessment, ensuring that the system can understand the complex emotions and behavior patterns of the target object in more detail; combining the target emotional characteristics and target archival information of the target object for psychological risk assessment, allowing the system to provide highly personalized assessment results based on personal history, personality characteristics and current emotional state. This personalized method can identify the type of psychological risk in a specific context, and provide an analysis basis and coping strategies that are more in line with personal needs, thereby improving the pertinence and effectiveness of psychological intervention measures; using the pre-trained target large language model for analysis, it can not only process and understand the emotional characteristics in the multimodal data, but also deeply explore the potential associations in the archival information and identify possible causes of psychological risks. The deep learning capability of the large model enables it to learn complex sentiment analysis rules and psychological risk assessment logic from massive data, and provide assessment results based on deep understanding; the psychological risk assessment results include psychological risk types, analysis basis and response strategies, which not only gives a qualitative description of the risk, but also provides an objective basis for risk assessment and specific intervention measures. Recommendations, thereby solving the technical problem that existing non-sensory psychological analysis technology relies on a single data source and a simple analysis model, and the analysis results lack accuracy and depth. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0017] Figure 1 is a flow chart of an optional psychological risk assessment method according to an embodiment of the present application;

[0018] Figure 2 is a schematic structural diagram of an optional psychological risk assessment device according to an embodiment of the present application;

[0019] Figure 3 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0021] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0022] Example 1

[0023] According to an embodiment of the present application, a psychological risk assessment method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0024] Figure 1 is a flow chart of a psychological risk assessment method provided according to an embodiment of the present application, such as Figure 1 As shown, the method comprises the following steps:

[0025] Step S102, acquiring multimodal target state data of the target object, wherein the modalities of the target state data at least include: a visual modality and a sound modality;

[0026] Step S104, determining the target emotion characteristics of the target object according to the multimodal target state data;

[0027] Step S106, obtaining target profile information associated with the target object;

[0028] Step S108, using the pre-trained target large language model to analyze the target emotional characteristics and target profile information to obtain the psychological risk assessment results of the target object, wherein the psychological risk assessment results at least include: psychological risk type, analysis basis and coping strategy.

[0029] The following describes each step of the psychological risk assessment method in conjunction with the specific implementation process.

[0030] First, multimodal target state data of the target object is obtained, wherein the modalities of the target state data include at least: visual modality and sound modality. The process can be carried out in the following ways: obtaining image data including the face and / or body of the target object; obtaining voice data of the communication process with the target object; and analyzing the voice data using voice recognition technology to obtain text data corresponding to the voice data.

[0031] For image data, for example, the facial expressions or body postures of the target objects can be captured by cameras or other image acquisition devices installed in the target environment. These image data can be static pictures or dynamic video streams, so as to capture the emotional changes of the target objects at different time points. The acquisition of visual modality data pays special attention to changes in facial micro-expressions and overall posture. These subtle visual clues can often reveal the individual's inner emotional state, even in an insensitive or hidden evaluation environment.

[0032] For voice data, for example, the target object's voice data during the communication process can be collected through a microphone or other audio collection equipment. This data may come from scenes such as daily conversations and other dialogues, providing audio information about the target object's emotional state. The analysis of voice data should not only take into account basic features such as volume and speaking speed, but also deeply analyze more complex voice features such as tone, intonation, pauses, etc. These features can reflect the individual's emotional fluctuations and psychological state.

[0033] For text data, for example, the collected speech data can be input into the Automatic Speech Recognition (ASR) system. ASR technology can convert speech signals into text form. This text data contains the verbal content of the target object and provides a basis for subsequent sentiment analysis. The text data generated by ASR not only includes literal meaning, but can also be further analyzed to extract higher-level features such as contextual information and emotional tendencies of language, which is crucial for comprehensive assessment of the psychological state of the target object.

[0034] After the target state data is obtained, the target emotion characteristics of the target object are determined according to the multimodal target state data. The process may include the following steps:

[0035] S1, respectively determining image features, voice features and text features corresponding to the image data, voice data and text data;

[0036] As an optional implementation, the following steps may be used:

[0037] Analyze the image data using a pre-trained image recognition model to obtain image features, wherein the image features include at least one of the following: expression features, eye movement features, and posture features;

[0038] For example, image data is analyzed using pre-trained image recognition models, such as Vision Transformer (ViT) or VideoMAE models. Image features include, but are not limited to, facial features (such as smile, frown, eye opening, upturned mouth corners, etc.), eye movement features (eye gaze direction, blinking frequency, etc.) and posture features (body posture, gesture, head rotation, etc.), which can reflect an individual's emotional state and psychological stress level.

[0039] Analyzing the speech data using a pre-trained speech recognition model to obtain speech features, wherein the speech features include at least one of the following: tone features and intonation features;

[0040] For example, the pre-trained speech recognition model HuBERT is used to analyze speech data and extract speech features. Speech features mainly include tone features (such as firmness, hesitation, joy, sadness) and intonation features (pitch changes, volume, speaking speed, etc.). Speech feature analysis can capture the emotional tendencies and changes in the psychological state of an individual during communication.

[0041] The text data is analyzed using a pre-trained text recognition model to obtain text features, wherein the text features include at least one of the following: context features and sentiment tendency features.

[0042] For example, the Transformer-based text analysis model extracts text features, which include contextual features (conversation context, discussion topics, etc.) and emotional tendency features (positive, negative, and neutral emotional tendencies). Text feature analysis helps to understand an individual's emotional expression and potential psychological problems in communication.

[0043] S2, using preset weight coefficients to perform weighted calculation on image features, speech features and text features to obtain multimodal fusion features;

[0044] Among them, the setting of weight coefficients can be based on the relative importance and credibility of different modal features in emotion recognition. For example, in certain situations, facial expressions may reflect emotional states more directly than voice tone. Taking these features into consideration comprehensively can generate a more comprehensive and accurate emotion assessment result.

[0045] S3, uses the valence arousal model to analyze the multimodal fusion features, and uses the two-dimensional emotion values ​​output by the valence arousal model as the target emotion features of the target object.

[0046] The valence-arousal model is a model used to analyze emotional states, usually based on Russell's valence-arousal theory, which defines the emotional space on a two-dimensional coordinate, where the horizontal axis represents valence and the vertical axis represents arousal level. In order for the model to be able to process multimodal fusion features, the model needs to be trained, and the training data should contain emotional labels (valence and arousal level) and corresponding multimodal features. Supervised learning methods such as support vector machines (SVM) and neural networks (such as deep learning models) can be used to pair multimodal features with valence and arousal values, and train the model to learn the mapping relationship from features to emotional values. Once the model training is completed, the multimodal fusion feature vector can be input into the model, and the model will output a two-dimensional emotional value of (valence, arousal). The valence value indicates the positive or negative tendency of the emotion, while the arousal value indicates the intensity or excitement of the emotion. For example, a value with high valence (positive emotion) and high arousal (strong emotion) may indicate excitement, while low valence and low arousal may indicate calmness. The output valence arousal model emotion value can be further interpreted as a specific emotional feature of the target object. For example, for workplace health monitoring scenarios, if the model outputs an emotional value with low valence and medium arousal level, this may indicate that the target object is in a melancholic or depressed emotional state. Based on this emotional feature, personalized health intervention recommendations can be generated, such as recommending relaxation exercises, providing career counseling, or adjusting the work environment to improve the mental health of employees.

[0047] In addition to determining the target emotional characteristics of the target object, it is also necessary to obtain the target profile information associated with the target object.

[0048] After obtaining the target emotional characteristics and target profile information of the target object, the target emotional characteristics and target profile information are analyzed using the pre-trained target large language model to obtain the psychological risk assessment results of the target object, wherein the psychological risk assessment results at least include: psychological risk type, analysis basis and coping strategy.

[0049] As an optional implementation, the training process of the target large language model may include the following steps:

[0050] S1, obtaining multiple groups of multi-modal historical state data and historical behavior data corresponding to multiple objects, and obtaining historical archive information of each object;

[0051] S2, for each subject, determine the subject's emotional characteristics based on the subject's corresponding multimodal historical state data, use the emotional characteristics and the subject's historical archive information as a training sample, and determine the subject's psychological risk type based on the subject's historical behavior data, and use the psychological risk type as a sample label of the training sample;

[0052] S3, determining the initial large language model based on the verification rules, wherein the verification rules are used to construct the mapping relationship between the emotional characteristics, historical archive information, psychological risk type, analysis basis and coping strategy of the object, and the analysis basis and coping strategy are determined based on the preset sociological and psychological knowledge map;

[0053] S4, iteratively train the initial large language model using multiple training samples and sample labels to obtain a target large language model.

[0054] After the target large language model is trained using the above method, the target emotional characteristics and target profile information of the target object are analyzed using the model to obtain the psychological risk assessment results of the target object.

[0055] As an optional implementation method, the verification rules are also used to guide the target large language model to generate a psychological risk assessment report in JSON format, wherein the psychological risk assessment report at least includes: report title, psychological risk type, risk description, analysis basis and coping strategies, and the coping strategies include at least one of the following: psychological intervention strategies (such as emotion regulation techniques, cognitive behavioral therapy recommendations), environmental intervention strategies (such as work environment adjustments, recommendations for reducing work stress).

[0056] As an optional implementation method, after executing the corresponding coping strategy, the multimodal state data and behavior data of the target object after executing the coping strategy can also be obtained, and the degree of improvement of the psychological risk of the target object can be determined based on the multimodal state data and behavior data. The quantification of the degree of improvement of psychological risk may be based on the change of emotion value in the valence-arousal model, or it can be completed by comparing statistical indicators of behavioral data (such as the frequency of positive behavior), or evaluated based on expert experience. Among them, when the degree of improvement of the psychological risk of the target object does not meet the preset standard, the coping strategy corresponding to the target emotional characteristics and target archival information in the verification rules of the target large language model is updated according to the degree of improvement of psychological risk. For example, according to the feedback of the psychological risk assessment results, the importance of certain features in the rules is adjusted, or new coping strategies are introduced, so that the system can propose more effective intervention measures when facing similar emotional characteristics and archival information.

[0057] For example, in the scenario of workplace health monitoring, suppose the target is a software engineer who is identified as having the emotional characteristics of "high pressure and low valence" in the multimodal non-sensory psychological assessment system, which may indicate that he is experiencing high-intensity work pressure and is dissatisfied with the current work environment or tasks. Based on his personal profile information (such as work history, personality traits, performance records) and emotional characteristics, a preliminary psychological risk assessment report was generated, and the company's human resources department was recommended to adopt the following coping strategies: Psychological intervention strategy: recommend stress management seminars and emotion regulation courses to help him learn effective ways to deal with high pressure and negative emotions; Environmental intervention strategy: adjust his workload, reduce overtime, and provide a quieter and more comfortable working environment to improve job satisfaction. One month after the human resources department implemented these recommended strategies, the system collected the engineer's multimodal state data and behavioral data again for evaluation. However, according to the new data, his emotional characteristics still showed "high pressure and medium valence", which means that his psychological pressure has been reduced, but his job satisfaction has not been significantly improved, and the degree of improvement in psychological risk has not reached the preset standard. In this case, the system will update the coping strategies corresponding to the emotional characteristics of "high pressure and low valence" and the personal profile information of the engineer in the verification rules of the target large language model based on the results of the psychological risk improvement analysis. The updates may include: Optimization of psychological intervention strategies: Given that emotion regulation courses may have limited effects on improving work efficiency and satisfaction, the system may recommend adding regular one-on-one psychological counseling to more directly address their psychological distress and provide more personalized support; Deepening of environmental intervention strategies: The system may recommend a deeper transformation of the work environment, such as providing more flexible work arrangements (for example, allowing remote work or flexible working hours), or establishing a more positive team atmosphere and enhancing communication and collaboration among team members through team building activities. The updated verification rules and coping strategies will be reintegrated into the target large language model for fine-tuning or training to reflect the new intervention measures and their expected effects. For example, the system will learn the emotional characteristics of "high pressure and low valence", combined with the personal profile information of the engineer (such as the tendency to need personal space and deep work), regular one-on-one psychological counseling plus flexible work environment adjustment may be more effective than the original strategy.

[0058] This iterative process ensures that the system can continuously optimize its coping strategies based on actual feedback and provide more precise and personalized psychological intervention suggestions, thereby more effectively helping the target subjects improve their psychological conditions and reduce psychological risks until the expected level of improvement is achieved.

[0059] In an embodiment of the present application, obtaining multimodal target state data of the target object (including but not limited to visual modality and sound modality) can capture the emotional state of the individual from multiple dimensions, avoiding the deviations and limitations that may be caused by single modality data. This multimodal fusion method improves the comprehensiveness and accuracy of emotion recognition and psychological state assessment, ensuring that the system can understand the complex emotions and behavior patterns of the target object in more detail; combining the target emotional characteristics and target archival information of the target object for psychological risk assessment, allowing the system to provide highly personalized assessment results based on personal history, personality characteristics and current emotional state. This personalized method can identify the type of psychological risk in a specific context, and provide an analysis basis and coping strategies that are more in line with personal needs, thereby improving the pertinence and effectiveness of psychological intervention measures; using the pre-trained target large language model for analysis, it can not only process and understand the emotional characteristics in the multimodal data, but also deeply explore the potential associations in the archival information and identify possible causes of psychological risks. The deep learning capability of the large model enables it to learn complex sentiment analysis rules and psychological risk assessment logic from massive data, and provide assessment results based on deep understanding; the psychological risk assessment results include psychological risk types, analysis basis and response strategies, which not only gives a qualitative description of the risk, but also provides an objective basis for risk assessment and specific intervention measures. Recommendations, thereby solving the technical problem that existing non-sensory psychological analysis technology relies on a single data source and a simple analysis model, and the analysis results lack accuracy and depth.

[0060] Example 2

[0061] According to an embodiment of the present application, a psychological risk assessment device for implementing the psychological risk assessment method in embodiment 1 is also provided. Figure 2 As shown, the psychological risk assessment device at least includes: a first acquisition module 21, a feature extraction module 22, a second acquisition module 23 and a psychological assessment module 24, wherein:

[0062] A first acquisition module 21 acquires multimodal target state data of a target object, wherein the modalities of the target state data at least include: a visual modality and a sound modality;

[0063] A feature extraction module 22, used to determine the target emotion feature of the target object based on the multimodal target state data;

[0064] The second acquisition module 23 is used to acquire target profile information associated with the target object;

[0065] The psychological assessment module 24 is used to analyze the target emotional characteristics and target profile information using the pre-trained target large language model to obtain the psychological risk assessment results of the target object, wherein the psychological risk assessment results at least include: psychological risk type, analysis basis and coping strategy.

[0066] The functions of each module of the psychological risk assessment device are explained below in conjunction with the specific implementation process.

[0067] The first acquisition module acquires multi-modal target state data of the target object, wherein the modalities of the target state data at least include: visual modality and sound modality. The process can be performed in the following manner:

[0068] Acquiring image data including the face and / or body of a target object;

[0069] Acquire voice data of the communication process with the target object;

[0070] The speech data is analyzed using speech recognition technology to obtain text data corresponding to the speech data.

[0071] After obtaining the target state data, the feature extraction module determines the target emotion feature of the target object according to the multimodal target state data. The process may include the following steps:

[0072] S1, respectively determining image features, voice features and text features corresponding to the image data, voice data and text data;

[0073] As an optional implementation, the following steps may be used:

[0074] Analyze the image data using a pre-trained image recognition model to obtain image features, wherein the image features include at least one of the following: expression features, eye movement features, and posture features;

[0075] Analyzing the speech data using a pre-trained speech recognition model to obtain speech features, wherein the speech features include at least one of the following: tone features and intonation features;

[0076] The text data is analyzed using a pre-trained text recognition model to obtain text features, wherein the text features include at least one of the following: context features and sentiment tendency features.

[0077] S2, using preset weight coefficients to perform weighted calculation on image features, speech features and text features to obtain multimodal fusion features;

[0078] S3, uses the valence arousal model to analyze the multimodal fusion features, and uses the two-dimensional emotion values ​​output by the valence arousal model as the target emotion features of the target object.

[0079] In addition to obtaining the target emotion characteristics of the determined target object, the second acquisition module acquires the target profile information associated with the target object.

[0080] After obtaining the target emotional characteristics and target profile information of the target object, the psychological assessment module uses the pre-trained target large language model to analyze the target emotional characteristics and target profile information to obtain the psychological risk assessment results of the target object, wherein the psychological risk assessment results at least include: psychological risk type, analysis basis and coping strategy.

[0081] As an optional implementation, the training process of the target large language model may include the following steps:

[0082] Obtain multiple sets of multi-modal historical status data and historical behavior data corresponding to multiple objects, and obtain historical archive information for each object;

[0083] For each subject, the emotional characteristics of the subject are determined based on the multimodal historical state data corresponding to the subject, and the emotional characteristics and the historical archive information of the subject are used as a training sample. The psychological risk type of the subject is determined based on the historical behavior data of the subject, and the psychological risk type is used as the sample label of the training sample.

[0084] Determine the initial large language model based on the verification rules, where the verification rules are used to construct the mapping relationship between the emotional characteristics, historical archive information, psychological risk type, analysis basis and coping strategy of the object, and the analysis basis and coping strategy are determined based on the preset sociological and psychological knowledge map;

[0085] The initial large language model is iteratively trained using multiple training samples and sample labels to obtain a target large language model.

[0086] Among them, the verification rules are also used to guide the target large language model to generate a psychological risk assessment report in JSON format, wherein the psychological risk assessment report at least includes: report title, psychological risk type, risk description, analysis basis and response strategy, and the response strategy includes at least one of the following: psychological intervention strategy, environmental intervention strategy.

[0087] The target object's multimodal status data and behavior data are obtained after the execution of the coping strategy, and the degree of improvement of the target object's psychological risk is determined based on the multimodal status data and behavior data; when the target object's psychological risk improvement degree does not reach the preset standard, the coping strategy corresponding to the target emotional characteristics and target profile information in the verification rules of the target large language model is updated based on the degree of improvement of the psychological risk.

[0088] It should be noted that each module in the psychological risk assessment device in the embodiment of the present application corresponds one by one to each implementation step of the psychological risk assessment method in Example 1. Since a detailed description has been given in Example 1, some details not reflected in this embodiment can be referred to Example 1 and will not be repeated here.

[0089] Example 3

[0090] According to an embodiment of the present application, a computer program product is also provided, which includes a computer program, wherein when the computer program is executed by a processor, the psychological risk assessment method in Example 1 is implemented.

[0091] According to an embodiment of the present application, a non-volatile storage medium is also provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the psychological risk assessment method in Example 1 by running the computer program.

[0092] According to an embodiment of the present application, a processor is also provided, which is used to run a computer program, wherein the psychological risk assessment method in Example 1 is executed when the computer program is running.

[0093] According to an embodiment of the present application, an electronic device is also provided, which includes: a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the psychological risk assessment method in Example 1 through the computer program.

[0094] Specifically, the computer program executes the following steps when it is running: obtaining multimodal target state data of the target object, wherein the modalities of the target state data include at least: visual modality and sound modality; determining the target emotional characteristics of the target object based on the multimodal target state data; obtaining target profile information associated with the target object; and analyzing the target emotional characteristics and target profile information using a pre-trained target large language model to obtain a psychological risk assessment result of the target object, wherein the psychological risk assessment result includes at least: psychological risk type, analysis basis, and coping strategy.

[0095] As an optional implementation, the electronic device may be in the form of a mobile terminal, a computer terminal or a similar computing device. Figure 3 FIG. 1 shows a hardware structure block diagram of an electronic device for implementing a psychological risk assessment method. Figure 3 As shown, the electronic device 30 may include one or more (302a, 302b, ..., 302n are used to illustrate) processors 302 (the processor 302 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission device 306 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 3The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 3 More or fewer components as shown, or with Figure 3 Different configurations are shown.

[0096] It should be noted that the one or more processors 302 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the electronic device 30. As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0097] The memory 304 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the psychological risk assessment method in the embodiment of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, to implement the vulnerability detection method of the above-mentioned application. The memory 304 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 304 may further include a memory remotely arranged relative to the processor 302, and these remote memories may be connected to the electronic device 30 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0098] The transmission device 306 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the electronic device 30. In one example, the transmission device 306 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 306 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0099] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the electronic device 30 .

[0100] The serial numbers of the above embodiments are only for description and do not represent the advantages or disadvantages of the embodiments.

[0101] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0103] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0104] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0105] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.

[0106] The above are only preferred implementations of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A psychological risk assessment method, characterized in that: include: Acquire multimodal target state data of the target object, wherein the modalities of the target state data at least include: a visual modality and a sound modality; Determining a target emotion feature of the target object according to the multimodal target state data; Acquire target profile information associated with the target object; The target emotional characteristics and the target profile information are analyzed using a pre-trained target large language model to obtain a psychological risk assessment result of the target object, wherein the psychological risk assessment result at least includes: psychological risk type, analysis basis and coping strategy.

2. The method according to claim 1, characterized in that Get the multi-modal target state data of the target object, including: Acquiring image data including the face and / or body of the target object; Acquiring voice data of the communication process with the target object; The speech data is analyzed using speech recognition technology to obtain text data corresponding to the speech data.

3. The method according to claim 2, characterized in that Determining the target emotion feature of the target object according to the multimodal target state data includes: Determine the image features, voice features and text features corresponding to the image data, the voice data and the text data respectively; The image feature, the speech feature and the text feature are weighted by using a preset weight coefficient to obtain a multimodal fusion feature; The multimodal fusion feature is analyzed using a valence arousal model, and the emotion values ​​of the two dimensions output by the valence arousal model are used as the target emotion features of the target object.

4. The method according to claim 3, characterized in that Respectively determining image features, voice features, and text features corresponding to the image data, the voice data, and the text data, including: Analyzing the image data using a pre-trained image recognition model to obtain the image features, wherein the image features include at least one of the following: expression features, eye movement features, and posture features; Analyzing the speech data using a pre-trained speech recognition model to obtain the speech features, wherein the speech features include at least one of the following: tone features and intonation features; The text data is analyzed using a pre-trained text recognition model to obtain the text features, wherein the text features include at least one of the following: context features and sentiment tendency features.

5. The method according to claim 1, characterized in that The training process of the target large language model includes: Obtain multiple sets of multi-modal historical status data and historical behavior data corresponding to multiple objects, and obtain historical archive information for each object; For each subject, determining the emotional characteristics of the subject according to the multimodal historical state data corresponding to the subject, using the emotional characteristics and the historical archive information of the subject as a training sample, and determining the psychological risk type of the subject according to the historical behavior data of the subject, using the psychological risk type as a sample label of the training sample; Determine an initial large language model based on verification rules, wherein the verification rules are used to construct a mapping relationship between the emotional characteristics, historical archive information, psychological risk type, analysis basis and coping strategy of the object, and the analysis basis and the coping strategy are determined based on a preset sociological and psychological knowledge map; The initial large language model is iteratively trained using a plurality of the training samples and sample labels to obtain the target large language model.

6. The method according to claim 5, characterized in that The verification rules are also used to guide the target large language model to generate a psychological risk assessment report in JSON format, wherein the psychological risk assessment report at least includes: report title, psychological risk type, risk description, analysis basis and response strategy, and the response strategy includes at least one of the following: psychological intervention strategy, environmental intervention strategy.

7. The method according to claim 5, characterized in that The method further comprises: Acquiring multimodal status data and behavior data of the target object after executing the coping strategy, and determining the degree of improvement of the psychological risk of the target object based on the multimodal status data and behavior data; When the psychological risk improvement degree of the target object does not reach the preset standard, the coping strategy corresponding to the target emotional characteristics and the target profile information in the verification rule of the target large language model is updated according to the psychological risk improvement degree.

8. A psychological risk assessment device, characterized in that: include: A first acquisition module is configured to acquire multimodal target state data of a target object, wherein the modalities of the target state data at least include: a visual modality and a sound modality; A feature extraction module, used to determine the target emotion feature of the target object according to the multimodal target state data; A second acquisition module is used to acquire target profile information associated with the target object; The psychological assessment module is used to analyze the target emotional characteristics and the target profile information using a pre-trained target large language model to obtain a psychological risk assessment result of the target object, wherein the psychological risk assessment result at least includes: psychological risk type, analysis basis and coping strategy.

9. A computer program product, characterized in that include: A computer program, wherein when the computer program is executed by a processor, the psychological risk assessment method according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the psychological risk assessment method according to any one of claims 1 to 7 through the computer program.

Citation Information

Cited By

  • Method and device for psychological assessment, medium and program product

    CN120203586A