Data processing method and device, electronic equipment and storage medium

By extracting mental state features from video, audio, and text data from interview transcripts, and utilizing multimodal information fusion processing and classification models, the problem of slow classification speed of interview transcript data was solved, achieving fast and accurate data classification.

CN115565672BActive Publication Date: 2026-04-21ANHUI IFLYHEALTH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI IFLYHEALTH CO LTD
Filing Date
2022-10-20
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In existing technologies, the classification of interview record data is slow and requires a high level of expertise, making it difficult to classify quickly and accurately.

Method used

By extracting mental state characteristics from video, audio, and text data from interview transcripts, and utilizing multimodal information fusion processing and classification models, rapid and accurate data classification can be achieved.

Benefits of technology

It enables rapid and accurate classification of interview record data, improving classification efficiency and accuracy while reducing reliance on professionals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115565672B_ABST
    Figure CN115565672B_ABST
Patent Text Reader

Abstract

The application provides a data processing method and device, electronic equipment and storage medium. The mental state features of a test subject are extracted from interview record data of the test subject, the interview record data including at least two of video data, audio data and text data recorded during an interview process. Then, mental state features corresponding to the same mental state category are fused to obtain fused mental state features. Based on the fused mental state features, data classification processing based on a set classification label is performed on the interview record data to determine the data type to which the interview record data belongs. The set classification label includes a plurality of set interview record data types, and different interview record data types correspond to different mental states. The application classifies the interview record data by means of multi-modal information such as audio, video and text of the test subject, and can quickly and accurately classify the interview record data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a data processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] With rapid societal development, the pace of life is accelerating, and people face increasing pressure, leading to a rise in mental health issues such as depression and anxiety. Analyzing interview data from psychological testing can help determine the underlying causes of mental health problems, allowing for the development of different coping strategies to alleviate depression, anxiety, and other mental health issues.

[0003] Analyzing interview transcripts related to the same mental health issues is crucial for obtaining more accurate results. Therefore, before analyzing the interview transcripts, it is necessary to categorize them according to different mental health issues. How to quickly and accurately categorize interview transcripts is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] Based on the above requirements, this application proposes a data processing method, apparatus, electronic device, and storage medium, which can quickly and accurately classify interview record data.

[0005] The technical solution proposed in this application is as follows:

[0006] On the one hand, this application provides a data processing method, including:

[0007] The mental state characteristics of the test subjects are extracted from the interview record data of the test subjects. The interview record data includes at least two of the video data, audio data and text data recorded during the interview. The mental state characteristics include the mental state characteristics corresponding to each preset mental state category.

[0008] The mental state features corresponding to the same mental state category are fused to obtain the fused mental state features;

[0009] Based on the fused mental state characteristics, the interview record data is classified according to a set classification label to determine the data type to which the interview record data belongs; wherein, the set classification label includes a variety of set interview record data types, and different interview record data types correspond to different mental states.

[0010] Furthermore, in the method described above, the mental state features corresponding to the same mental state category are fused to obtain fused mental state features, including:

[0011] The interview record data was divided into multiple interview record data segments according to different test questions;

[0012] The mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused mental state features for each interview record data segment.

[0013] Furthermore, in the method described above, the mental state feature carries the confidence level of the mental state feature; the mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused mental state feature corresponding to each interview record data segment, including:

[0014] The confidence scores of mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused confidence scores of mental state features for each mental state category corresponding to each interview record data segment.

[0015] Furthermore, in the method described above, based on the fused mental state characteristics, the interview record data is subjected to data classification processing based on set classification labels to determine the data type to which the interview record data belongs, including:

[0016] The confidence level of the mental state features for each mental state category corresponding to each interview record data segment is determined based on the fused confidence level of the mental state features for each mental state category corresponding to each interview record data segment.

[0017] Based on the confidence level of the mental state characteristics of each mental state category corresponding to the interview record data, the interview record data is classified according to the set classification labels to determine the data type to which the interview record data belongs.

[0018] Furthermore, in the method described above, based on the confidence level of the mental state characteristics corresponding to each mental state category of the interview record data, data classification processing based on set classification labels is performed on the interview record data to determine the data type to which the interview record data belongs, including:

[0019] Calculate the confidence scores of the mental state characteristics for each mental state category corresponding to the interview record data.

[0020] Determine the confidence level and the confidence level interval in which it falls; wherein, different confidence levels correspond to different set classification labels;

[0021] The data type of the interview record corresponding to the set classification label of the confidence level and the confidence level interval is determined as the data type to which the interview record data belongs.

[0022] Furthermore, in the method described above, the interview record data includes video data and audio data; extracting the mental state characteristics of the test subject from the interview record data includes:

[0023] Text data is obtained by performing speech recognition on the audio data;

[0024] The mental state characteristics of the test subject are extracted from the text data, video data, and audio data.

[0025] Furthermore, in the method described above, the mental state characteristics include mental state characteristics extracted from the audio data, including at least one of depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics.

[0026] The mental state characteristics extracted from the audio data include at least one of the following: depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics;

[0027] The mental state features extracted from the text data include at least one of the following: depressive mood features, guilt features, difficulty falling asleep features, shallow sleep features, early awakening features, work and interest features, mental anxiety features, physical anxiety features, gastrointestinal symptoms features, systemic symptoms features, hypochondriasis features, weight loss features, and insight features.

[0028] Furthermore, in the method described above, the mental state characteristics of the test subjects are extracted from the interview record data of the test subjects; mental state characteristics corresponding to the same mental state category are fused to obtain fused mental state characteristics; and based on the fused mental state characteristics, the interview record data is classified according to a set classification label to determine the data type to which the interview record data belongs, including:

[0029] The interview record data of the test subjects is input into a pre-trained interview record data classification model, so that the interview record data classification model can extract the mental state characteristics of the test subjects from the interview record data of the test subjects. The mental state characteristics corresponding to the same mental state category are fused to obtain fused mental state characteristics. Based on the fused mental state characteristics, the interview record data is classified according to the set classification labels to determine the data type to which the interview record data belongs.

[0030] On the other hand, this application provides a data processing apparatus, including:

[0031] The extraction module is used to extract the mental state characteristics of the test subject from the interview record data of the test subject. The interview record data includes at least two of the following: video data, audio data and text data recorded during the interview. The mental state characteristics include mental state characteristics corresponding to each preset mental state category.

[0032] The fusion module is used to fuse mental state features that correspond to the same mental state category to obtain fused mental state features.

[0033] The classification module is used to perform data classification processing on the interview record data based on the fused mental state characteristics, and to determine the data type to which the interview record data belongs; wherein, the set classification labels include a variety of set interview record data types, and different interview record data types correspond to different mental states.

[0034] On the other hand, this application provides an electronic device, including:

[0035] Memory and processor;

[0036] The memory is used to store programs;

[0037] The processor is configured to implement any of the above-described data processing methods by running a program in the memory.

[0038] On the other hand, this application provides a storage medium, including: a computer program stored on the storage medium, wherein when the computer program is executed by a processor, it implements the data processing method described in any one of the above.

[0039] The data processing method proposed in this application extracts the mental state characteristics of the test subjects from their interview record data. The interview record data includes at least two of the following: video data, audio data, and text data recorded during the interview. The mental state characteristics include those corresponding to various preset mental state categories. Then, the mental state characteristics corresponding to the same category are fused to obtain fused mental state characteristics. Based on these fused characteristics, the interview record data is classified according to predefined classification labels to determine the data type to which the interview record data belongs. The predefined classification labels include multiple types of interview record data, with different data types corresponding to different mental states. This application utilizes multimodal information such as audio, video, and text from the test subjects to classify interview record data, achieving the goal of rapid and accurate classification of interview record data. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0041] Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0042] Figure 2 This is a flowchart illustrating a method for determining the type of interview record data, as provided in an embodiment of this application.

[0043] Figure 3 This is a schematic diagram of a process for extracting mental state features according to an embodiment of this application;

[0044] Figure 4 This is an architecture diagram of an interview record data classification model provided in an embodiment of this application;

[0045] Figure 5 This is an architecture diagram of another interview record data classification model provided in an embodiment of this application;

[0046] Figure 6 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application;

[0047] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0048] Application Overview

[0049] The technical solution of this application is applicable to the application scenario of classifying interview record data. By adopting the technical solution of this application, the purpose of quickly and accurately classifying interview record data can be achieved.

[0050] With societal development, people's pace of life is accelerating, and they face increasing pressure. The accumulation of negative emotions, difficult to release, leads to mental health issues such as depression and anxiety. Currently, the detection process for mental health problems generally involves: a doctor interviewing the subject; the doctor then uses the interview record data to determine if the subject has a mental health issue; if so, the type of mental health issue is further determined.

[0051] To determine the pathogenesis and coping strategies of mental health issues such as depression and anxiety, the current approach involves recruiting a large number of volunteers, collecting their interview records, and analyzing this data to identify the underlying causes and coping strategies for these issues.

[0052] Classifying interview transcripts based on different mental states facilitates analysis of transcripts within the same mental state, leading to more accurate results. Therefore, before analyzing interview transcripts, it is necessary to classify them according to different mental states. Currently, this classification is done by professional medical personnel, which not only demands high levels of expertise from the classifiers but is also slow.

[0053] Based on this, this application proposes a data processing method, apparatus, electronic device, and storage medium. This technical solution can quickly and accurately classify interview record data based on the multimodal data recorded during the interview process.

[0054] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0055] Exemplary methods

[0056] This application provides a data processing method that can be executed by an electronic device. The electronic device can be any device with data and instruction processing capabilities, such as a computer, a smart terminal, or a server. See also... Figure 1 As shown, the method includes:

[0057] S101. Extract the mental state characteristics of the test subjects from the interview records of the test subjects.

[0058] The test subjects mentioned above refer to those who undergo psychological interviews. They can be people or intelligent electronic devices, such as intelligent robots, smartphones, and intelligent cars. This embodiment does not impose any limitations.

[0059] The aforementioned interview record data includes at least two of the following: video data, audio data, and text data recorded during the interview.

[0060] Video data includes video images of the test subject's face and limbs during the psychological interview. The video images are processed into grayscale data and then subjected to frame extraction, for example, retaining only 25 frames per second, to obtain the processed video data. The video data is acquired by a camera device, which can be a webcam or an electronic device equipped with a camera, etc., and this embodiment is not limited to this. Audio data includes the test subject's speech data during the psychological interview, generally the speech data of the test subject answering test questions during the psychological interview. Through framing, windowing, discrete Fourier transform, and filtering, the speech data is converted into FBank40 features commonly used in the field of automatic speech recognition, and the FBank40 features are used as audio data. The audio data is acquired by an audio acquisition device, which can be a recorder or an electronic device equipped with recording capabilities, etc., and this embodiment is not limited to this.

[0061] It should be noted that some existing electronic devices can simultaneously collect audio and video data. Such electronic devices can be used to collect audio and video data during interviews, and then audio and video data can be extracted from the audio and video data.

[0062] The text data includes the text content of the test subjects' answers to test questions during the psychological interview. A text recorder can be assigned to record the text content of the test subjects' answers to test questions during the psychological interview. The text recorder can be the questioner, other personnel, or the test subjects themselves; this embodiment does not impose any limitations. Furthermore, the test subjects' audio data during the psychological interview can be converted into text data.

[0063] The aforementioned mental state characteristics include those corresponding to pre-defined mental state categories. Different types of interview record data can be pre-defined with different mental state categories, and these categories are generally related to the mental state corresponding to that type of interview record data. For example, if the interview record data is used to detect depression using the Hamilton Depression Scale (HAMD), the mental state categories can include the 17, 21, or 24 categories set in the HAMD. As another example, if the interview record data is used to detect depression using the PHQ-9 scale, the mental state categories can include the 9 categories set in the PHQ-9 scale. Similarly, if the interview record data is used to detect anxiety disorders using the Hamilton Anxiety Scale (HAMA), the mental state categories can include the 14 categories set in the HAMA.

[0064] The specific contents of the 17, 21, and 24 mental state categories set in HAMD, the 9 mental state categories set in the PHQ-9 scale, and the 14 mental state categories set in HAMA can be obtained by those skilled in the art without expending creativity, and will not be elaborated here.

[0065] In this embodiment, the mental state characteristics of the test subjects are extracted from their interview transcripts. Specifically, mental state characteristics can be extracted from at least two of the video, audio, and text data.

[0066] Specifically, the video data contains the test subjects' facial expressions and body movements. Based on these expressions and movements, the test subjects' emotional state and reaction speed can be reflected. Therefore, mental state features related to emotional state and reaction speed can be extracted from the video data. For example, if the mental state categories include the 17 mental state categories set in HAMD, at least one of the following can be extracted from the video data: depressive mood features, mental retardation features, mental agitation features, and mental anxiety features.

[0067] Audio data contains the vocal characteristics of the test subjects, which can reflect their emotional state and reaction speed. Therefore, mental state characteristics related to emotional state and reaction speed can be extracted from audio data. For example, if the mental state categories include the 17 mental state categories set in HAMD, at least one of the following can be extracted from the audio data: depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics.

[0068] Text data includes the test subjects' responses, but these responses alone cannot reflect their reaction speed. Therefore, it is possible to extract mental state features that are not related to reaction speed from the text data. For example, if the mental state categories include the 17 categories set in HAMD, the mental retardation and agitation features extracted from the text features may not be accurate. Instead, at least one of the following features can be extracted from the text data: depressive mood features, guilt features, difficulty falling asleep, shallow sleep features, early awakening features, work and interest features, mental anxiety features, somatic anxiety features, gastrointestinal symptoms features, general symptoms features, hypochondriasis features, weight loss features, insight features, etc.

[0069] Neural network models can be used to extract mental state features. Specifically, video feature extraction models can be used to extract mental state features from video data, audio feature extraction models can be used to extract mental state features from audio data, and text feature extraction models can be used to extract mental state features from text data.

[0070] The mental state characteristics of test subjects extracted from their interview records can also carry confidence levels that characterize the severity of those characteristics. Higher confidence levels indicate a more severe mental state characteristic. If the mental state categories include the 17 categories defined in the HAMD (Hardware and Advanced Mindset) framework, then according to HAMD specifications, the confidence levels are generally 0-4, 0-2, or 0-3. A confidence level of 0 for a mental state characteristic indicates that the test subject does not possess that characteristic, while a higher confidence level indicates a more severe manifestation of that characteristic in the test subject.

[0071] The aforementioned video feature extraction model can select a wav2vec2.0 pre-trained model trained on large-scale video data as the basic training model. Only the feature encoding related layer of the basic training model is retained. The feature embedding vector extraction capability of this basic training model is trained by inputting a large number of training samples to obtain the video feature extraction model. In this embodiment, the training samples of the video feature extraction model are sample video data. The processing procedure for the sample video data is the same as the video data processing procedure in the above embodiment, which is to first convert the video image into a grayscale image and then perform frame extraction. The training labels are the mental state features and confidence levels corresponding to different sample data labeled by professionals. The training samples are input into the video feature extraction model to train the feature embedding vector extraction capability. When the loss value of the video feature extraction model reaches within a set value, the training is complete.

[0072] The aforementioned audio feature extraction model can select the Conformer automatic speech recognition pre-trained model trained with large-scale speech recognition data as the basic training model. Only the feature encoding related layer of the basic training model is retained. The feature embedding vector extraction capability of this basic training model is trained by inputting a large number of training samples, thus obtaining the audio feature extraction model. In this embodiment, the training samples for the audio feature extraction model are sample audio data. The processing procedure for the sample audio data is the same as the audio data processing procedure in the above embodiments, which involves extracting FBank40 features. The training labels are the mental state features and confidence levels corresponding to different sample data labeled by professionals. The training samples are input into the audio feature extraction model to train its feature embedding vector extraction capability. Training is complete when the loss value of the audio feature extraction model reaches within a set value.

[0073] The aforementioned text feature extraction model can select a BERT model trained on large-scale reading text data as the base training model. Only the feature encoding layers of the base training model are retained. The feature embedding vector extraction capability of this base training model is trained by inputting a large number of training samples, thus obtaining the text feature extraction model. In this embodiment, the training samples for the text feature extraction model are sample text data, and the training labels are the mental state features and confidence levels corresponding to different sample data labeled by professionals. The training samples are input into the text feature extraction model to train its feature embedding vector extraction capability. Training is complete when the loss value of the text feature extraction model reaches within a set value.

[0074] The training labels for the video feature extraction model, audio feature extraction model, and text feature extraction model can include mental state features corresponding to different sample data labeled by professionals. These mental state features also carry corresponding confidence levels. With this setup, the outputs of the video, audio, and text feature extraction models are the mental state features and their confidence levels. Optionally, the mental state features and confidence levels corresponding to the different sample data labeled by professionals can be one-hot encoded to obtain the training labels for each model.

[0075] It should also be noted that if the number of training samples is too small to meet the training requirements, the smaller proportion of data can be expanded through methods such as data resampling to increase the amount of training samples and balance the data types.

[0076] S102. Merge the mental state features corresponding to the same mental state category to obtain the merged mental state features.

[0077] After extracting the mental state characteristics of the test subject from at least two of the video data, audio data, and text data recorded during the interview through the steps of the above embodiments, the mental state characteristics of the same mental state type are fused together.

[0078] Specifically, through the above embodiments, the confidence scores of mental state features extracted from at least two of the video data, audio data, and text data can be fused together to obtain the fused mental state feature confidence scores as the fused mental state features.

[0079] A pre-trained feature confidence fusion model can be used to fuse the confidence levels of mental state features. The training samples for the feature confidence fusion model are the confidence levels corresponding to the mental state features of samples from various mental state categories, with the labels being the fusion results. The training samples are input into the feature confidence fusion model for training. Training is complete when the loss value of the feature confidence fusion model is less than a set value.

[0080] S103. Based on the integrated mental state characteristics, perform data classification processing on the interview record data based on the set classification labels to determine the data type to which the interview record data belongs.

[0081] The aforementioned classification labels include pre-defined data types of interview records, with different data types corresponding to different mental states. Mental states can include several common mental health issues, such as normal, depression, and anxiety; they can also include severity levels of the same mental health issue, such as normal, mild depression, moderate depression, and severe depression. Those skilled in the art can set these according to data classification needs, and this embodiment does not impose limitations. Based on different mental states, the interview record data types can include data types corresponding to several common mental health issues, such as normal data types, depression data types, and anxiety data types; they can also include data types corresponding to the severity levels of the same mental health issue, such as normal level data types, mild depression level data types, moderate depression level data types, and severe depression level data types.

[0082] In this embodiment, different data type classification labels are pre-set. For example, if the interview record data is classified according to the level of depression, the classification labels include normal level data type label, mild depression level data type label, moderate depression level data type label and severe depression level data type label.

[0083] Each defined category label corresponds to a defined confidence level interval. If the confidence level of the interview record data is determined to fall within a certain target defined confidence level interval based on the fused mental state characteristics, then the data type corresponding to that target defined confidence level interval is determined as the data type corresponding to the interview record data.

[0084] For example, the confidence level of interview record data can be obtained according to the following example: As described in the above embodiments, the confidence levels of mental state features corresponding to the same mental state category can be fused to obtain a fused confidence level of mental state features. Then, the confidence level of interview record data can be determined based on the fused confidence level of mental state features. For example, the sum of the confidence levels of the fused mental state features can be calculated, and the sum of the confidence levels can be used as the confidence level of the interview record data.

[0085] In the above embodiments, the mental state characteristics of the test subjects are extracted from the interview record data. The interview record data includes at least two of the following: video data, audio data, and text data recorded during the interview. The mental state characteristics include mental state characteristics corresponding to each preset mental state category. Then, the mental state characteristics corresponding to the same mental state category are fused to obtain fused mental state characteristics. Based on the fused mental state characteristics, the interview record data is classified according to a set classification label to determine the data type to which the interview record data belongs. The set classification label includes multiple set interview record data types, and different interview record data types correspond to different mental states. This application uses the multimodal information of the test subjects, such as audio, video, and text, to classify interview record data, which can achieve the purpose of quickly and accurately classifying interview record data.

[0086] As an optional implementation, another embodiment of this application discloses that the steps of the above embodiments fuse mental state features corresponding to the same mental state category to obtain fused mental state features, which may specifically include the following steps:

[0087] Based on different test questions, the interview record data was divided into multiple interview record data segments; the mental state features corresponding to the same mental state category in each interview record data segment were fused to obtain the fused mental state features corresponding to each interview record data segment.

[0088] In the embodiments of this application, the interview record data is segmented according to different test questions.

[0089] Specifically, if the interview transcript data includes text, audio, and video data, the audio and video data of the interview transcript data can be obtained. Voice endpoint detection technology is used to segment the audio and video data to obtain audio and video data segments corresponding to different test questions. Then, video images are extracted from each audio and video data segment. After performing grayscale processing and frame extraction as described in the above embodiments, the video images are obtained to obtain the video data segment corresponding to each test question. Voice data is extracted from each audio and video data segment, and FBank40 features are extracted from the voice data as described in the above embodiments to obtain the audio data segment corresponding to each test question. Text conversion processing is performed on the audio data segment corresponding to each test question to obtain the text data segment corresponding to each test question. The audio data segment, text segment, and video data segment corresponding to the same test question are combined to form the interview transcript data segment corresponding to that test question.

[0090] Furthermore, if audio and video data are extracted separately, speech endpoint detection technology can be used to segment the speech data, obtaining audio data segments corresponding to different test questions. For video data acquisition, after the test subject answers a question, they can make a specific gesture or posture, or another person can make a specific gesture or posture within the camera's field of view. After acquiring the video image, it can be segmented into video images corresponding to each test question based on the specific gesture or posture. These video images can then be processed to obtain the video data segment corresponding to each test question.

[0091] The mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused mental state features for each interview record data segment. The specific fusion process for fusing the confidence levels of mental state features corresponding to the same mental state category has been described in detail in the above embodiments, and those skilled in the art can refer to the descriptions in the above embodiments; it will not be repeated here.

[0092] In the above embodiments, the mental state features corresponding to the same mental state category for each test question are fused and processed to achieve the purpose of analyzing the data type of the interview record data for each question. This can avoid the omission of effective information and improve the reliability of the classification results.

[0093] As an optional implementation, another embodiment of this application discloses that if the mental state feature carries the confidence level of the mental state feature, then the mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused mental state feature corresponding to each interview record data segment. Specifically, this may include the following steps:

[0094] The confidence scores of mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused confidence scores of mental state features for each mental state category corresponding to each interview record data segment.

[0095] Specifically, the confidence scores of mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused mental state feature confidence scores of each mental state category corresponding to each interview record data segment as the fused mental state feature.

[0096] For example, if the mental state categories include the 17 mental state categories set in HAMD, the following features are extracted from the video data segment: depressive mood features, mental retardation features, mental agitation features, and mental anxiety features, as well as the confidence scores corresponding to each mental state feature; depressive mood features, mental retardation features, mental agitation features, and mental anxiety features, as well as the confidence scores corresponding to each mental state feature, are extracted from the audio data segment; and all features except mental retardation features and mental agitation features, as well as the confidence scores corresponding to each mental state feature, are extracted from the text data segment.

[0097] The confidence scores of depressive mood features extracted from video, audio, and text data segments can be fused to obtain the fused confidence score of the depressive mood features for this interview transcript data segment. Similarly, the confidence scores of anxiety features extracted from video, audio, and text data segments can be fused to obtain the fused confidence score of the anxiety features for this interview transcript data segment. Since no mental retardation features were extracted from the text data segment, the confidence scores of mental retardation features extracted from video and audio data segments are fused to obtain the fused confidence score of the anxiety features for this interview transcript data segment. Confidence scores for the mental retardation feature; since no agitation feature was extracted from the text data segment, the confidence scores of the agitation features extracted from the video data segment and the audio data segment were fused to obtain the fused confidence scores for the agitation features corresponding to this interview record data segment; since all mental state features other than depressive mood features, mental retardation features, agitation features, and anxiety features were only extracted from the text data segment, the confidence scores of all mental state features other than depressive mood features, mental retardation features, agitation features, and anxiety features were fused with an empty set to obtain the fused confidence scores for all mental state features other than depressive mood features, mental retardation features, agitation features, and anxiety features corresponding to this interview record data segment. The confidence scores for all mental state features other than depressive mood features, mental retardation features, agitation features, and anxiety features after fusion are the same as before fusion.

[0098] A pre-trained feature confidence fusion model can be used to fuse the confidence levels of mental state features corresponding to the same mental state category in each segment of interview records. The training samples for the feature confidence fusion model are the confidence levels corresponding to the mental state features of various mental state categories, with the labels being the fusion results. The training samples are input into the feature confidence fusion model for training, and training is complete when the loss value of the feature confidence fusion model is detected to be less than a set value.

[0099] By inputting the confidence scores of the mental state features corresponding to the same mental state category in each interview record data segment into the above feature confidence fusion model, we can obtain the fused mental state feature confidence scores of each mental state category corresponding to each interview record data segment output by the feature confidence fusion model.

[0100] In the above embodiments, by determining the confidence level of the fused mental state features for each category of mental state corresponding to each interview record data segment, the purpose of analyzing the data type of the interview record data for each question is achieved. This avoids the omission of effective information and improves the reliability of the classification results. Furthermore, the above embodiments employ a method that first uses a feature extraction network model to extract the confidence level of each mental state feature in each interview record data segment, then uses a fusion network to fuse the confidence levels to obtain the fused mental state feature confidence level for each interview record data segment. Based on the fused mental state feature confidence level for each interview record data segment, the confidence level of the interview record data is determined. Different computational steps are distributed across different neural networks, thereby reducing the computational power required for individual models and improving the accuracy of the classification results.

[0101] As an optional implementation method, such as Figure 2 As shown in another embodiment of this application, the steps of the above embodiments, which perform data classification processing on the interview record data based on set classification labels according to the fused mental state characteristics, and determine the data type to which the interview record data belongs, may specifically include the following steps:

[0102] S201. Based on the confidence level of the fused mental state features of each mental state category corresponding to each interview record data segment, determine the confidence level of the mental state features of each mental state category corresponding to the interview record data.

[0103] In this embodiment, the confidence level of the mental state features corresponding to each mental state category of the interview record data is determined based on the confidence level of the mental state features after fusion of each mental state category corresponding to each interview record data segment.

[0104] The interview transcript data includes the mental state categories corresponding to all interview transcript data segments.

[0105] If all interview record data segments contain fused mental state features of the same mental state category, the average confidence level of the fused mental state features of the same mental state category can be calculated and determined as the confidence level of the mental state feature of that mental state category corresponding to the interview record data. In addition to calculating the average value, a weighted average value can also be calculated according to the set weights, which is not limited in this embodiment.

[0106] If all interview record data segments contain a fused mental state feature that is different from the mental state category of other fused mental state features, then the confidence level corresponding to the fused mental state feature of that mental state category can be determined as the confidence level of the mental state feature of that mental state category corresponding to the interview record data.

[0107] For example, if the interview record data contains three interview data segments, the first interview data segment contains fused mental state features for mental state category A with a fused confidence level of a1, fused mental state features for mental state category B with a fused confidence level of b1, and fused mental state features for mental state category C with a fused confidence level of c1. The second interview data segment contains fused mental state features for mental state category A with a fused confidence level of a2, fused mental state features for mental state category B with a fused confidence level of b2, and the third interview data segment contains fused mental state features for mental state category A with a fused confidence level of a3, and fused mental state features for mental state category D with a fused confidence level of d1.

[0108] The interview record data includes mental state characteristics for mental state category A, with confidence levels of a1, a2, and a3 (average); mental state characteristics for mental state category B, with confidence levels of b1 and b2 (average); mental state characteristics for mental state category C, with confidence level of c1; and mental state characteristics for mental state category D, with confidence level of d1.

[0109] S202. Based on the confidence level of the mental state characteristics of each mental state category corresponding to the interview record data, perform data classification processing based on the set classification labels to determine the data type to which the interview record data belongs.

[0110] After determining the confidence level of the mental state characteristics for each mental state category corresponding to the interview record data, the confidence level of the mental state characteristics for each mental state category corresponding to the interview record data can be calculated. Based on the confidence level and the sum of the sums, the set classification labels for the interview record data are determined, and the data type corresponding to the set classification labels is determined as the data type of the interview record data.

[0111] In the above embodiments, the confidence level of the mental state features of each mental state category corresponding to each interview record data segment is determined based on the fused confidence level of the mental state features of each mental state category corresponding to each interview record data segment, thereby obtaining the data type to which the interview record data belongs.

[0112] As an optional implementation, another embodiment of this application discloses that the steps of the above embodiments, based on the confidence level of the mental state characteristics of each mental state category corresponding to the interview record data, perform data classification processing on the interview record data based on a set classification label to determine the data type to which the interview record data belongs, may specifically include the following steps:

[0113] Calculate the confidence scores of the mental state characteristics for each mental state category corresponding to the interview record data; determine the confidence interval in which the confidence scores are located; and determine the data type of the interview record data corresponding to the classification label set by the confidence scores and the confidence interval in which they are located as the data type to which the interview record data belongs.

[0114] Each defined category label corresponds to a defined confidence level interval. If the confidence level of the mental state characteristics of each mental state category corresponding to the interview record data falls within a certain target defined confidence level interval, then the data type corresponding to that target defined confidence level interval is determined as the data type corresponding to the interview record data.

[0115] For example, if the mental state categories include the 17 mental state categories set in the HAMD and the interview record data is classified according to the level of depression, according to the HAMD, a confidence interval of less than or equal to 7 corresponds to the normal level data type; a confidence interval of greater than 7 and less than or equal to 17 corresponds to the mild depression level data type; a confidence interval of greater than 17 and less than or equal to 24 corresponds to the moderate depression level data type; and a confidence interval of greater than 24 corresponds to the severe depression level data type. When the sum of the above confidence levels is 6, the interview record data of the test subject can be determined to be the normal level data type; when the sum of the above confidence levels is 18, the interview record data of the test subject can be determined to be the moderate depression level data type.

[0116] In the above embodiments, the interview transcript data is classified according to the mental state features extracted from the video, audio and text modal data, which can achieve the purpose of quickly and accurately classifying the interview transcript data.

[0117] As an optional implementation method, such as Figure 3 As shown in another embodiment of this application, the steps of the above embodiments to extract the mental state characteristics of the test subjects from the interview record data of the test subjects may specifically include the following steps:

[0118] S301. Text data is obtained by performing speech recognition on audio data.

[0119] Specifically, the interview transcript data in this embodiment includes video data and audio data. Text data is obtained by performing speech recognition on the audio data. Converting audio data into text data using speech recognition technology is a conventional existing technique in the art, and those skilled in the art can refer to existing descriptions; therefore, speech recognition will not be described in detail here.

[0120] S302. Extract the mental state characteristics of the test subject from text data, video data, and audio data.

[0121] The mental state characteristics of the test subjects were extracted from text data, video data, and audio data, respectively.

[0122] For example, if the mental state categories include the 17 mental state categories set in HAMD, then the mental state characteristics of the test subject extracted from video data include at least one of the following: depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics; the mental state characteristics of the test subject extracted from audio data include at least one of the following: depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics; the mental state characteristics of the test subject extracted from text data include at least one of the following fifteen mental state characteristics other than mental retardation and mental agitation characteristics, such as at least one of the following: depressive mood characteristics, guilt characteristics, difficulty falling asleep characteristics, shallow sleep characteristics, early awakening characteristics, work and interest characteristics, mental anxiety characteristics, somatic anxiety characteristics, gastrointestinal symptoms characteristics, general symptoms characteristics, hypochondriasis characteristics, weight loss characteristics, and insight characteristics.

[0123] In the above embodiments, the audio data is converted into text features without the need for manual input by personnel, which not only saves manpower but also ensures the speed of text feature generation.

[0124] As an optional implementation, another embodiment of this application discloses that the steps of the above embodiments extract the mental state characteristics of the test subjects from the interview record data, fuse the mental state characteristics corresponding to the same mental state category to obtain the fused mental state characteristics, and perform data classification processing on the interview record data based on the set classification labels according to the fused mental state characteristics to determine the data type to which the interview record data belongs. Specifically, it may include the following steps:

[0125] The interview record data of the test subjects is input into a pre-trained interview record data classification model, so that the interview record data classification model can extract the mental state characteristics of the test subjects from the interview record data. The mental state characteristics corresponding to the same mental state category are fused to obtain the fused mental state characteristics. Based on the fused mental state characteristics, the interview record data is classified according to the set classification labels to determine the data type of the interview record data.

[0126] Specifically, in this embodiment, an interview record data classification model is pre-trained to facilitate the extraction and fusion of mental state characteristics based on the interview record data classification model, and to classify the interview record data according to the fused mental state characteristics.

[0127] For example, such as Figure 4 As shown, the interview record data classification model includes an automatic speech recognition layer, a speech feature extraction layer, a speech coding module, a speech item classification layer, a text feature extraction layer, a text coding module, a text item classification layer, a video feature extraction layer, a video coding module, a video item classification layer, a modal result fusion layer, and a decision layer.

[0128] The system comprises several modules: a speech feature extraction layer to extract mental state feature vectors from audio data of various interview transcripts; and a speech encoding module to encode these feature vectors and input them into a speech item classification layer. The encoded feature vectors are then classified according to predefined mental state categories, yielding the extracted mental state features from the audio data. The posterior probability of these features is used as the confidence level. For example, the speech encoding module can utilize a pre-trained feature encoding layer from the Conformer automatic speech recognition model, allowing for fine-tuning of the data in this layer during later training. Using a pre-trained Conformer automatic speech recognition model reduces the need for training samples, and the feature encoding layer only requires fine-tuning, eliminating the need for initial training and simplifying the training process.

[0129] The video feature extraction layer extracts mental state feature vectors from video data of various interview transcript segments. The video encoding module encodes these mental state feature vectors and inputs them into the video item classification layer. The mental state feature vectors are then classified according to preset mental state categories, yielding the mental state features extracted from the video data. The posterior probability of these mental state features is used as the confidence level of the mental state features. For example, the video encoding module can use the feature encoding-related layer of a pre-trained wav2vec2.0 model, which can be fine-tuned during later training. Using a pre-trained wav2vec2.0 model not only reduces the amount of training samples required but also simplifies the training process by eliminating the need for initial training, as the feature encoding-related layer only requires fine-tuning.

[0130] The automatic speech recognition layer is used to perform speech recognition on audio data to obtain text features. The text feature extraction layer is used to extract mental state feature vectors from the text data of each interview transcript segment. The text encoding module is used to encode the mental state feature vectors and input the encoded mental state feature vectors into the text item classification layer, classifying the mental state feature vectors according to preset mental state categories to obtain the mental state features extracted from the text data. The posterior probability of the mental state features is used as the confidence level of the mental state features. For example, the text encoding module can use the feature encoding related layer in a pre-trained BERT model, and the data of the feature encoding related layer can be fine-tuned during the later training process. Using a pre-trained BERT model not only reduces the amount of training samples required, but also simplifies the training process because the data of the feature encoding related layer only needs to be fine-tuned, without having to train from the beginning.

[0131] The modality result fusion layer is used to fuse the confidence scores of mental state features of the same mental state type in the same interview record data segment. Then, the decision layer obtains the posterior probability of the fused mental state features of each mental state category corresponding to each interview record data segment. This posterior probability is used as the confidence score of the fused mental state features of each mental state category corresponding to each interview record data segment.

[0132] Then, based on the confidence level of the fused mental state features of each mental state category corresponding to each interview record data segment, the confidence level of the mental state features of each mental state category corresponding to the interview record data is determined; based on the confidence level of the mental state features of each mental state category corresponding to the interview record data, the interview record data is classified according to the set classification labels to determine the data type to which the interview record data belongs.

[0133] For example, such as Figure 4As shown, if the mental state categories include the 17 mental state categories set in HAMD and the interview record data is classified according to the level of depression, after determining the confidence level of the mental state features of each mental state category corresponding to the interview record data, the confidence level of the mental state features of each mental state category corresponding to the interview record data is calculated, and then the data type to which the interview record data belongs is determined according to the mapping relationship of HAMD.

[0134] Specifically, if the above confidence level is less than or equal to 7, the interview record data is of normal level data type; if the above confidence level is greater than 7 and less than or equal to 17, the interview record data is of mild depression level data type; if the above confidence level is greater than 17 and less than or equal to 24, the interview record data is of moderate depression level data type; and if the above confidence level is greater than 24, the interview record data is of severe depression level data type.

[0135] When training the aforementioned interview record data classification model, the automatic speech recognition layer can be trained separately. The training samples are audio data, and the training labels are the text data corresponding to each training sample.

[0136] The speech feature extraction layer, speech coding module, speech item classification layer, text feature extraction layer, text coding module, text item classification layer, video feature extraction layer, video coding module, video item classification layer, modality result fusion layer, and decision layer can be trained as a unified network combination. Training samples consist of audio, text, and video data corresponding to the same test question. Mental state features and confidence levels determined by professional medical personnel for different test questions based on the corresponding audio, text, and video data are obtained. These mental state features and confidence levels are then mapped onto one-hot vectors to obtain training labels.

[0137] The training samples are input into the above network combination for training, and the output of the network combination is obtained. The loss value of the network combination is determined based on the output of the network combination and the training labels. Training is complete when the loss value is less than a set value.

[0138] Alternatively, the automatic speech recognition layer, speech feature extraction layer, speech coding module, speech item classification layer, text feature extraction layer, text coding module, text item classification layer, video feature extraction layer, video coding module, video item classification layer, modal result fusion layer, and decision layer can be trained simultaneously. The training samples are video data and audio data, and the labels are the same as the above network combination. The training process is also the same. Those skilled in the art can refer to the training process of the above network combination, which will not be elaborated here.

[0139] Another example, such as Figure 5As shown, the interview record data classification model includes an automatic speech recognition layer, a speech feature extraction layer, a text feature extraction layer, a video feature extraction layer, a modal feature fusion layer, a modal item classification layer, and a decision layer.

[0140] The input audio, text, and video data are aligned along the time dimension. The speech feature extraction layer extracts mental state feature vectors from the audio data of each interview transcript segment. The video feature extraction layer extracts mental state feature vectors from the video data of each interview transcript segment. The automatic speech recognition layer performs speech recognition on the audio data to obtain text features, and the text feature extraction layer extracts mental state feature vectors from the text features of each interview transcript segment.

[0141] The modal feature fusion layer is used to fuse and encode the mental state feature vectors in the same interview record data segment. Then, the modal item classification layer classifies each fused feature according to the set mental state type to obtain the mental state features of each type. The decision layer further obtains the posterior probability of the mental state features of each mental state type, which is used as the confidence level of each mental state type.

[0142] Then, based on the confidence level of the fused mental state features for each mental state category corresponding to each interview record data segment, the confidence level of the mental state features for each mental state category corresponding to the interview record data is determined. Based on the confidence level of the mental state features for each mental state category corresponding to the interview record data, data classification processing based on set classification labels is performed on the interview record data to determine the data type to which the interview record data belongs. The specific processing procedure is as follows... Figure 4 The processing procedures in the illustrated embodiments are the same, and those skilled in the art can refer to the descriptions in the above embodiments.

[0143] The training process of the interview record data classification model in this embodiment and Figure 5 The training process of the interview record data classification model in the illustrated embodiment is the same. Those skilled in the art can refer to the description in the above embodiments, and it will not be repeated here.

[0144] It should be noted that the technical methods described in the above embodiments can be used to diagnose mental illnesses. Specifically, interview record data of the subject undergoing mental examination can be obtained. By analyzing the interview record data using the technical methods described in the above embodiments, the data type corresponding to the interview record data can be determined, and the mental state corresponding to that data type can be identified as the mental health test result of the subject undergoing mental examination.

[0145] For example, if a mental health examination subject is being tested for depression, interview transcripts generated during the depression testing process can be obtained. Mental state characteristics of the test subject can be extracted from these transcripts. If the HAMD (Hearing Aptitude Test) with seventeen common test results is used for depression testing, the extracted mental state characteristics include seventeen items such as depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics. Mental state characteristics corresponding to the same mental state category are fused to obtain fused mental state characteristics. Based on these fused characteristics, the interview transcript data is classified according to a set classification label to determine the data type to which the interview transcript data belongs. The depression level corresponding to the data type to which the interview transcript data belongs is determined as the test subject's depression level, for example, the test subject's depression level is normal, mild depression, moderate depression, or severe depression, etc.

[0146] Exemplary device

[0147] Corresponding to the above data processing method, this application also discloses a data processing apparatus, see [link to relevant documentation]. Figure 6 As shown, the device includes:

[0148] The extraction module 100 is used to extract the mental state characteristics of the test subjects from the interview record data of the test subjects. The interview record data includes at least two of the video data, audio data and text data recorded during the interview. The mental state characteristics include the mental state characteristics corresponding to each preset mental state category.

[0149] The fusion module 110 is used to fuse mental state features corresponding to the same mental state category to obtain fused mental state features.

[0150] The classification module 120 is used to classify the interview record data based on the fused mental state characteristics and determine the data type of the interview record data. The classification labels include multiple types of interview record data, and different types of interview record data correspond to different mental states.

[0151] As an optional implementation, another embodiment of this application discloses the fusion module 110 of the above embodiments, which includes:

[0152] The segmentation unit is used to divide the interview record data into multiple segments according to different test questions;

[0153] The fusion unit is used to fuse the mental state features corresponding to the same mental state category in each interview record data segment to obtain the fused mental state features corresponding to each interview record data segment.

[0154] As an optional implementation, another embodiment of this application discloses that the mental state feature carries the confidence level of the mental state feature; when the fusion unit of the above embodiment performs fusion processing on the mental state features corresponding to the same mental state category in each interview record data segment to obtain the fused mental state feature corresponding to each interview record data segment, it is specifically used for:

[0155] The confidence scores of mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused confidence scores of mental state features for each mental state category corresponding to each interview record data segment.

[0156] As an optional implementation, another embodiment of this application discloses the classification module 120 of the above embodiments, which includes:

[0157] The unit is defined to determine the confidence level of the mental state features of each mental state category corresponding to each interview record data segment based on the confidence level of the fused mental state features of each mental state category.

[0158] The classification unit is used to classify the interview record data based on the confidence level of the mental state characteristics of each mental state category corresponding to the interview record data, and to determine the data type to which the interview record data belongs.

[0159] As an optional implementation, another embodiment of this application discloses that, when the classification unit of the above embodiments performs data classification processing on the interview record data based on the confidence level of the mental state characteristics of each mental state category corresponding to the interview record data, and determines the data type to which the interview record data belongs, it is specifically used for:

[0160] Calculate the confidence scores of the mental state characteristics for each mental state category corresponding to the interview record data; determine the confidence interval in which the confidence scores fall; different confidence intervals correspond to different set classification labels; determine the data type of the interview record data corresponding to the set classification label of the confidence score and its confidence interval as the data type to which the interview record data belongs.

[0161] As an optional implementation, another embodiment of this application discloses that the interview record data includes video data and audio data; the extraction module 100 of the above embodiment includes:

[0162] The recognition unit is used to obtain text data by performing speech recognition on audio data;

[0163] The extraction unit is used to extract the mental state characteristics of the test subject from text data, video data, and audio data.

[0164] As an optional implementation, another embodiment of this application discloses that the mental state characteristics include mental state characteristics extracted from audio data, including at least one of depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics.

[0165] The mental state characteristics extracted from the audio data include at least one of the following: depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics;

[0166] The mental state features extracted from the text data include at least one of the following: depressive mood features, guilt features, difficulty falling asleep features, shallow sleep features, early awakening features, work and interest features, mental anxiety features, physical anxiety features, gastrointestinal symptoms features, systemic symptoms features, hypochondriasis features, weight loss features, and insight features.

[0167] As an optional implementation, another embodiment of this application discloses an extraction module 100 that extracts the mental state characteristics of the test subjects from the interview record data; a fusion module 110 that fuses mental state characteristics corresponding to the same mental state category to obtain fused mental state characteristics; and a classification module 120 that performs data classification processing on the interview record data based on set classification labels according to the fused mental state characteristics. Specifically, when determining the data type of the interview record data, the module 120 is used for:

[0168] The interview record data of the test subjects is input into a pre-trained interview record data classification model, so that the interview record data classification model can extract the mental state characteristics of the test subjects from the interview record data. The mental state characteristics corresponding to the same mental state category are fused to obtain the fused mental state characteristics. Based on the fused mental state characteristics, the interview record data is classified according to the set classification labels to determine the data type of the interview record data.

[0169] For details on the specific working functions of each unit of the aforementioned data processing device, please refer to the above method embodiments; they will not be repeated here.

[0170] Exemplary electronic devices, computer program products, and storage media

[0171] Another embodiment of this application also provides an electronic device, see [link to relevant documentation] Figure 7As shown, the electronic device includes:

[0172] Memory 200 and processor 210;

[0173] The memory 200 is connected to the processor 210 and is used to store programs;

[0174] The processor 210 is configured to implement the data processing method disclosed in any of the above embodiments by running a program stored in the memory 200.

[0175] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 220, an input device 230, and an output device 240.

[0176] The processor 210, memory 200, communication interface 220, input device 230, and output device 240 are interconnected via a bus. Among them:

[0177] A bus can include a pathway for transmitting information between various components of a computer system.

[0178] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present application. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0179] Processor 210 may include a main processor, as well as a baseband chip, modem, etc.

[0180] The memory 200 stores a program that executes the technical solution of this application, and may also store an operating system and other critical business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 200 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.

[0181] Input device 230 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.

[0182] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.

[0183] The communication interface 220 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.

[0184] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of the data processing method provided in the above embodiments of this application.

[0185] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by processor 210, cause processor 210 to perform the various steps of the data processing method provided in the above embodiments.

[0186] Computer program products can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0187] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor 210 to perform various steps of the data processing method provided in the above embodiments.

[0188] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0189] Specifically, the specific working content of each part of the aforementioned electronic device, computer program product, and storage medium, as well as the specific processing content of the computer program product or the computer program on the aforementioned storage medium when run by the processor, can all be found in the various embodiments of the aforementioned data processing method, and will not be repeated here.

[0190] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0191] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0192] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.

[0193] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.

[0194] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0195] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.

[0196] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.

[0197] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0198] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0199] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0200] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method of processing data, characterized by, include: The mental state characteristics of the test subjects are extracted from the interview record data of the test subjects. The interview record data includes at least two of the following: video data, audio data and text data recorded during the interview. The mental state characteristics include mental state characteristics corresponding to each preset mental state category. The mental state characteristics carry the confidence level of the mental state characteristics. The confidence level is used to characterize the severity of the mental state characteristics. The higher the confidence level, the deeper the degree of the mental state characteristics corresponding to that confidence level. The mental state features corresponding to the same mental state category are fused to obtain the fused mental state features; Based on the fused mental state characteristics, the interview record data is classified according to a set classification label to determine the data type to which the interview record data belongs; wherein, the set classification label includes a variety of set interview record data types, and different interview record data types correspond to different mental states.

2. The method of claim 1, wherein, Mental state features corresponding to the same mental state category are fused to obtain fused mental state features, including: The interview record data was divided into multiple interview record data segments according to different test questions; The mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused mental state features for each interview record data segment.

3. The method of claim 2, wherein, The mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused mental state features for each interview record data segment, including: The confidence scores of mental state features corresponding to the same mental state category in each interview record data segment are fused to obtain the fused confidence scores of mental state features for each mental state category corresponding to each interview record data segment.

4. The method of claim 3, wherein, Based on the fused mental state characteristics, the interview record data is classified according to predefined classification labels to determine the data type to which the interview record data belongs, including: The confidence level of the mental state features for each mental state category corresponding to each interview record data segment is determined based on the fused confidence level of the mental state features for each mental state category corresponding to each interview record data segment. Based on the confidence level of the mental state characteristics of each mental state category corresponding to the interview record data, the interview record data is classified according to the set classification labels to determine the data type to which the interview record data belongs.

5. The method of claim 4, wherein, Based on the confidence levels of mental state characteristics corresponding to each mental state category in the interview record data, the interview record data is classified according to predefined classification labels to determine the data type to which the interview record data belongs, including: Calculate the confidence scores of the mental state characteristics for each mental state category corresponding to the interview record data. Determine the confidence level and the confidence level interval in which it falls; wherein, different confidence levels correspond to different set classification labels; The data type of the interview record corresponding to the set classification label of the confidence level and the confidence level interval is determined as the data type to which the interview record data belongs.

6. The method of claim 1, wherein, The interview transcript data includes video and audio data; the mental state characteristics of the test subjects are extracted from the interview transcript data, including: Text data is obtained by performing speech recognition on the audio data; The mental state characteristics of the test subject are extracted from the text data, video data, and audio data.

7. The method of claim 6, wherein, The mental state characteristics include mental state characteristics extracted from the audio data, including at least one of depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics; The mental state characteristics extracted from the audio data include at least one of the following: depressive mood characteristics, mental retardation characteristics, mental agitation characteristics, and mental anxiety characteristics; The mental state features extracted from the text data include at least one of the following: depressive mood features, guilt features, difficulty falling asleep features, shallow sleep features, early awakening features, work and interest features, mental anxiety features, physical anxiety features, gastrointestinal symptoms features, systemic symptoms features, hypochondriasis features, weight loss features, and insight features.

8. The method of claim 1, wherein, The mental state characteristics of the test subjects are extracted from their interview records. Mental state characteristics corresponding to the same category are then fused to obtain fused mental state characteristics. Based on these fused characteristics, the interview records are classified according to predefined labels to determine the data type to which the interview records belong, including: The interview record data of the test subjects is input into a pre-trained interview record data classification model, so that the interview record data classification model can extract the mental state characteristics of the test subjects from the interview record data of the test subjects. The mental state characteristics corresponding to the same mental state category are fused to obtain fused mental state characteristics. Based on the fused mental state characteristics, the interview record data is classified according to the set classification labels to determine the data type to which the interview record data belongs.

9. A data processing device, characterized by comprising: include: An extraction module is used to extract the mental state characteristics of the test subjects from the interview record data of the test subjects. The interview record data includes at least two of the following: video data, audio data, and text data recorded during the interview. The mental state characteristics include mental state characteristics corresponding to each preset mental state category. The mental state characteristics carry the confidence level of the mental state characteristics. The confidence level is used to characterize the severity of the mental state characteristics. The higher the confidence level, the deeper the degree of the mental state characteristics corresponding to that confidence level. The fusion module is used to fuse mental state features that correspond to the same mental state category to obtain fused mental state features. The classification module is used to perform data classification processing on the interview record data based on the fused mental state characteristics, and to determine the data type to which the interview record data belongs; wherein, the set classification labels include a variety of set interview record data types, and different interview record data types correspond to different mental states.

10. An electronic device, comprising: include: Memory and processor; The memory is used to store programs; The processor is configured to implement the data processing method as described in any one of claims 1 to 8 by running a program in the memory.

11. A storage medium, characterized by include: The storage medium stores a computer program, which, when executed by a processor, implements the data processing method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Mental health evaluation method and system based on dialogue communication

    CN112768070A