Dilemma Children Attachment State Assessment Method and Device Based on Multimodal Large Model
Through the automated evaluation method of multimodal large model, the problem of information loss and interaction impact unmodeled caused by artificial feature screening in attachment state evaluation is solved, and efficient and accurate attachment state evaluation is achieved.
Patent Information
- Application Number
- CN202410745472.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-11
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-06-11
AI Technical Summary
The existing attachment state assessment methods rely on manual screening characteristics, resulting in information loss and insufficient reliability and consistency of evaluation results, and failure to effectively model the interaction effects during the interview process.
An automated evaluation method based on multimodal large model is adopted, through dialogue round segmentation, multimodal feature extraction and fusion, pre-trained models and large language models are used to predict attachment states, and an evaluation model based on sequence computing architecture is constructed.
It improves the efficiency and accuracy of attachment assessment, reduces human resources requirements, improves the consistency and accuracy of assessment results, and can effectively model the interaction impact during the interview process.
Smart Images

Figure CN118737432B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and device for evaluating the attachment status of children in distress based on a multimodal large model. Background Art
[0002] Attachment Theory is a core theory for analyzing the social and emotional development of children and their adulthood. The theory emphasizes that the early emotional connection between children and their caregivers has a decisive influence on their long-term mental health and behavioral development. Studies have shown that a stable attachment relationship can promote the development of children's positive self-cognition, social adaptability and emotional regulation, while the lack or problems of attachment relationships may lead to mental health challenges such as anxiety, depression and behavioral problems.
[0003] Traditionally, the assessment of children's attachment status mainly relies on methods such as questionnaires, semi-structured interviews and projective tests. Although these methods provide a path to understand children's attachment, they each have their limitations. For example, although the questionnaire is easy to operate, its results are highly dependent on the subjective feelings of the participants and may not truly reflect the children's attachment status. Semi-structured interviews and projective tests can deeply explore the emotional state of children, but require long-term professional analysis, and their interpretation is subjective. In recent years, some studies have begun to explore automated detection methods based on multimodal data, such as its application in the treatment of depression and autism assessment. These methods attempt to build a more comprehensive psychological state assessment model by analyzing multiple data types such as language, voice and facial expressions. However, these technologies are still rarely used in the assessment of attachment status; related research also has widespread deficiencies: for example, many methods use predictions based on artificially designed and manually screened features. In the process of feature extraction and feature screening, not only a large amount of manpower is required, but also the implicit information in the data is lost; in addition, many methods fail to fully consider the impact of the interactive process on the subjects during the detection process, such as the interaction between counselors and children in the interview, such as the lack of effective modeling of the interview dialogue process.
[0004] Existing attachment state assessment methods have several significant drawbacks. For traditional attachment state detection methods: using questionnaires makes the results dependent on the subjective willingness of the subjects to express; traditional methods based on semi-structured interviews and projective tests not only require a large amount of manpower for assessment, but also introduce the subjective judgment of the assessors, which may affect the reliability and consistency of the assessment results. In the research on the detection of abnormal mental and psychological states based on multi-modal data, the target group is children in distress, and there is little attention paid to the detection of attachment states; in these few existing attachment detection studies, they often focus on artificially designed and manually screened features, losing some implicit information. At the same time, these detection methods for abnormal mental and psychological states based on multi-modal data do not consider the impact of the interaction process on the subjects during the data collection process, such as the impact of doctors or counselors on the target children, and lack effective modeling of multi-round conversations in interviews, which not only limits the standardization of the method and the consistency of the results, but also has a negative impact on the accuracy of the assessment.
[0005] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0006] Based on this, it is necessary to provide a method and device for assessing the attachment state of children in distress based on a multi-modal large model to address the above technical problems. The present invention reduces the human resources required for assessing the attachment state of children in distress through automated assessment means, thereby improving the assessment efficiency, and uses an attachment state assessment model based on a sequence computing architecture to model the multi-round conversation process, improving the accuracy and robustness of attachment assessment.
[0007] According to the first aspect of the present invention, there is provided a method for predicting the attachment state of children in distress based on a multi-modal large model, including:
[0008] A dialogue turn segmentation step, which segments the dialogue turns between the child in distress and the doctor through an automated recognition method;
[0009] A feature extraction step, which extracts expression, action, acoustic, and language channels from the single-round dialogue data, and performs data fusion through a large language model to obtain multi-modal fusion features based on the single-round dialogue or the attachment state of the child in distress;
[0010] An attachment state prediction step, which inputs multiple multi-modal fusion features or the attachment state of the child in distress obtained from the single-round dialogue data into the attachment state assessment model of the child in distress based on multi-round conversations to predict the attachment state of the child in distress.
[0011] In some embodiments, the extracting of expressions, actions, acoustics, and language channels from the single-round conversation data, and performing data fusion through a large language model to obtain multimodal fusion features based on the single-round conversation or the attachment status of the child in distress includes:
[0012] Extract expressions, actions, and acoustic channels from single-round conversation data as input, and use the hidden layer output of the pre-trained model as deep features. Project the deep features into the word embedding space used by the large language model to form word embedding space data.
[0013] For the language channel, the text features in the current round of dialogue are obtained through the speech recognition model;
[0014] The word embedding space data, the text features, the metadata of the current round of dialogue, and the prompt words requiring the large language model to predict the attachment state are input into the large language model, the hidden layer features of the large language model are extracted as the multimodal fusion features of the single-round dialogue, and the results are output as the attachment state of the distressed child in the single-round dialogue.
[0015] In some embodiments, the method for constructing the attachment status assessment model for children in distress based on multiple rounds of dialogue is:
[0016] The long-term dependencies in multi-round conversations are captured by an attachment state evaluation model based on a sequential computing architecture, thereby modeling the content of multi-round conversations.
[0017] In some embodiments, the pre-trained model includes: SynFace for the expression channel; C3D of OpenMM for the action channel; and HuBERT for the acoustic channel.
[0018] In some embodiments, projecting the deep features into the word embedding space used by the large language model to form word embedding space data includes: projecting the deep features into the word embedding space used by the large language model to form word embedding space data through a neural network composed of a Q-Former and a multi-layer perceptron.
[0019] In some embodiments, the training method of the neural network composed of Q-Former and multi-layer perceptron is: supervised training is performed by using (I, t) with labeled data, the content of (I, t) is about the attachment of children in distress, so as to project the input information into the word embedding area related to attachment, input I into the pre-trained model used, and input it into the large language model through the neural network composed of Q-Former and multi-layer perceptron to generate text; then generate error L through the text txt-gen Update the parameters of the neural network composed of Q-Former and multi-layer perceptron. The update formula is:
[0020] argmin Θ L txt-gen (LLM(P, Θ(PT_model(I))), t)
[0021] Wherein, I represents the video or audio input in the data, t is the text description, Θ represents the neural network composed of Q-Former and multi-layer perceptron, LLM represents the large language model used, P is the prompt used for training, and PT_model represents the pre-trained model used to extract features for the current channel.
[0022] In some embodiments, the speech recognition model is Whisper.
[0023] According to a second aspect of the present invention, there is provided a device for predicting the attachment status of children in distress based on a multi-modal large model, including:
[0024] A dialogue turn segmentation module for segmenting the dialogue turns between children in distress and doctors through an automatic recognition method;
[0025] A feature extraction module for extracting expression, action, acoustic, and language channels from single-turn dialogue data and performing data fusion through a large language model to obtain multi-modal fusion features based on single-turn dialogue or the attachment status of children in distress;
[0026] An attachment status prediction module for inputting multiple multi-modal fusion features or the attachment status of children in distress obtained from single-turn dialogue data into an evaluation model for the attachment status of children in distress based on multi-turn dialogue to predict the attachment status of children in distress.
[0027] According to a third aspect of the present invention, there is provided a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the computer program, the steps of the method in any of the above embodiments are implemented.
[0028] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method in any of the above embodiments are implemented.
[0029] Based on the above technical solutions, the present invention has at least the following beneficial effects:
[0030] 1. Introduce a multi-modal feature extraction and fusion method based on pre-trained model multi-modal feature extraction - large language model multi-modal fusion understanding in the field of automated attachment assessment, effectively making up for the problem that the subjective factors of evaluators may be introduced in the existing traditional attachment assessment methods, improving the assessment efficiency and consistency, solving the hidden information omitted when setting and manually screening features in the existing automated attachment assessment methods, and improving the stability of the assessment method.
[0031] 2. A channel feature projection method constrained by attachment state evaluation and detection is used, which can constrain features into attachment-related word embeddings during the feature projection process. This method can ensure that large language models pre-trained based on a general large-scale corpus can effectively extract multi-modal information features related to attachment and improve the accuracy of the evaluation model.
[0032] 3. An automated evaluation process framework for attachment state based on dialogue turns is introduced. It can effectively model the interaction process during the attachment evaluation process and fully consider the relationship between the variables during the inquiry by doctors or counselors and the response feedback of children in difficult situations when outputting the evaluation results, thereby improving the accuracy and robustness of the evaluation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 is a flowchart of some embodiments of the method for predicting the attachment state of children in difficult situations based on a multi-modal large model of the present invention;
[0034] Figure 2 is a schematic diagram of feature extraction of single-round dialogue data of the present invention;
[0035] Figure 3 is a framework diagram of the method for predicting attachment state based on multi-round dialogue of the present invention;
[0036] Figure 4 is a flowchart of some other embodiments of the method for predicting the attachment state of children in difficult situations based on a multi-modal large model of the present invention;
[0037] Figure 5 is a schematic structural diagram of some embodiments of the device for predicting the attachment state of children in difficult situations based on a multi-modal large model of the present invention;
[0038] Figure 6 is an internal structure diagram of a computer device for implementing some embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] The embodiments of the present invention will be described more fully hereinafter with reference to the accompanying drawings, in which embodiments of the present invention are shown. However, the present invention can be implemented in many different forms and should not be construed as limited to the embodiments set forth herein.
[0040] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. As used herein, the singular forms "a", "an", "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that when used herein, the term "comprising" specifies the presence of the claimed features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof.
[0041] Unless otherwise defined, the terms (including technical and scientific terms) used herein have the same meanings as commonly understood by those of ordinary skill in the art to which the present disclosure belongs. The terms used herein should be interpreted as having the meanings consistent with their meanings in the context of this specification and in the relevant field, and should not be interpreted in an idealized or overly formal sense, unless specifically defined as such herein.
[0042] The present invention provides a method for predicting the attachment status of children in distress based on a multimodal large model. The method uses audio and video data of interviews and questions between children in distress and doctors about attachment-related issues, and uses a pre-trained recognition method to segment the conversation rounds in the interview process; for each round of conversation, the pre-trained deep learning model is used to extract facial expressions, actions, voices, and sub-language features, and these features are projected into the word embedding space used by the large language model through a channel feature projector fine-tuned for attachment-related features, and by extracting the deep features of the large language model, a multimodal fusion deep embedding is obtained as the fusion feature of the conversation round; prediction is performed based on a deep network model with a sequence computing architecture. The method extracts features from multi-round conversation audio and video, maps them to a deep large model for data feature fusion, and predicts through a sequence model, providing a more efficient, accurate, and standardized method for attachment detection of children in distress.
[0043] Figure 1 The flowchart of some embodiments of the method for predicting the attachment status of children in distress based on a multimodal large model of the present invention is shown. Figure 1 As shown, the method includes:
[0044] In the dialogue turn segmentation step S102, the dialogue turns between the distressed child and the doctor are segmented by an automated recognition method; this step enables subsequent multimodal feature extraction and fusion understanding to perform feature extraction based on single-round data.
[0045] In the feature extraction step S104, facial expressions, actions, acoustics, and language channels are extracted from the single-round conversation data, and data fusion is performed through a large language model to obtain multimodal fusion features based on a single-round conversation or the attachment status of the child in distress. By using objective indicators extracted from audio and video data to evaluate the attachment status of the subject, the evaluation results are prevented from being affected by the subjective factors of the evaluator, and the consistency of the model results is improved.
[0046] Figure 2 FIG. 4 shows a schematic diagram of feature extraction of single-round dialogue data in the present invention. Figure 2 As shown, for expression, action and acoustic channels, the hidden layer features of the pre-trained deep recognition model are used as the representation features of the channel, and the features are projected into the word embedding feature space of the large language model by some method, and then the projected features are input into the large language model. For the audio-language channel, the text of the conversation is identified from the audio data of the current conversation by the pre-trained speech recognition model, and it is input into the large language model. At the same time, the metadata about this round of conversation is input into the large language model in the form of text, and the large language model is required to predict the attachment status of the child in distress based on the input information by the prompt method. Finally, the result predicted by the large language model is used as the prediction result of the attachment status of the child in distress based on a single round of conversation, and the features of the hidden layer of the large language model are extracted as multimodal fusion feature data based on a single round of conversation.
[0047] By using pre-trained models to extract features, and using projectors and large language models based on attachment state theory keyword constraints to perform multimodal fusion understanding on them. Using pre-trained models based on large-scale data sets, we extract easy-to-generalize, robust multimodal features, avoiding the problem of implicit information loss when manually designing features, and improving the accuracy of the evaluation model.
[0048] In the attachment status prediction step S106, a plurality of multimodal fusion features or the attachment status of the child in distress obtained from the single-round dialogue data are input into the attachment status assessment model of the child in distress based on multiple rounds of dialogue to predict the attachment status of the child in distress.
[0049] In some embodiments, multiple single-round multimodal understanding features or results extracted by the multimodal large model are fused at the feature level or result level, and the attachment status of children in distress is predicted by a deep learning model based on a sequence computing architecture. Deep learning models with a sequence computing architecture such as LSTM, GRU, and Transformer can capture long-term dependencies in multi-round conversations, thereby modeling the content of multi-round conversations and more accurately predicting the attachment status of children in distress.
[0050] By modeling the multi-round dialogue process using an attachment state assessment model based on a sequence calculation architecture on the basis of the multi-modal features extracted in each dialogue turn, the accuracy and robustness of attachment assessment are improved, and the problem of ignoring the influence of the interaction process in the data collection process in existing automated assessment and detection methods is solved.
[0051] Figure 3 The framework diagram of the attachment state prediction method based on multi-round dialogue in the present invention is shown. As Figure 3 shown, first, the dialogue turns are segmented, and then multi-modal fusion features or prediction results are obtained through a large language model. Finally, the multi-modal fusion features or prediction results are input into an attachment assessment model based on multi-round dialogue features to predict the attachment state of children.
[0052] Figure 4 The flowchart of some other embodiments of the attachment state prediction method for children in difficult situations based on the multi-modal large model of the present invention is shown. As Figure 4 shown, the method includes:
[0053] (1) Use an automated dialogue turn segmentation method, such as a speaker recognition method based on voiceprint segmentation and clustering, to segment the turns of the doctor's dialogue with children in difficult situations. And sequentially input the single-round dialogue information into a multi-modal large language feature extractor.
[0054] (2) Process the audio-visual data in the single-round dialogue through a pre-trained model to extract video-expression, video-action, audio-acoustic, and audio-language information.
[0055] (3) For video-expression, video-action, and audio-acoustics, use a pre-trained deep recognition model. Take the video or audio of the current round of conversation as input, and extract the output of the hidden layer of the pre-trained model as the deep feature. Then project this deep feature into the word embedding space used by the large language model. For the video-expression channel, the pre-trained model that can be used is SynFace. For the video-action channel, the pre-trained model that can be used is C3D of OpenMM. For the audio-acoustics channel, the pre-trained model that can be used is HuBERT. The method of projecting the deep feature into the word embedding space used by the large language model can be a neural network composed of Q-Former + multi-layer perceptron. This method can be supervised and trained using video / audio-description word (I, t) data with labels. Where I represents the video or audio input in the data, and t is the text description that matches this segment of the model. In the present invention, the content of the used (I, t) should be about the attachment of children in difficult situations, so as to constrain the model to project the input information into the word embedding area related to attachment. Input the video / audio I into the used pre-trained model, and input it into the large language model through the Q-Former + multi-layer perceptron network to generate text. Then update the parameters of the Q-Former + multi-layer perceptron network through the text generation error L txt-gen to update the parameters of the Q-Former + multi-layer perceptron network. The update formula is:
[0056] argmin Θ L txt-gen (LLM(P, Θ(PT_model(I))), t)
[0057] where Θ represents the neural network composed of Q-Former and multi-layer perceptron, LLM represents the used large language model, P is the prompt word used for training, and PT_model represents the pre-trained model used for feature extraction in the current channel.
[0058] (4) For the audio-language channel, use a pre-trained speech-to-text conversion (also known as speech recognition) model to extract the text information in the current round of conversation and input the text into the large language model. The speech recognition model that can be used is Whisper.
[0059] (5) The channel features projected into the word embedding space obtained from video-expression, video-action, and audio-acoustic channels; the text features extracted from the audio-language channel; and the metadata of the current round of conversation are used to input a prompt to the large language model asking it to predict the attachment status, extract the hidden layer features of the large language model as the multi-modal data features of the current round, and extract the output of the large language model as the prediction result of the current round. Among them, the large language model used can be GPT-3; the input conversation metadata can include the current speaking role, such as doctor, distressed child, and can also include information data such as the number of turns and duration of the current conversation; the prompt used can be: Attachment status refers to the deep dependence and connection that an individual has on the primary caregiver emotionally and psychologically, mainly divided into four types: secure, avoidant, anxious, and disorganized. Please act as a professional psychologist and, based on the conversation process between the distressed child and the psychologist regarding the attachment relationship provided below, currently <child / doctor speaking>, where the facial expression features of the child <speaking / listening> are <insert the projected facial expression data>, the movement features are <insert the projected movement feature data>, the acoustic features of the <child / doctor> speaking are <insert the projected acoustic features>, and the speaking content is <insert the text of the speech-to-text transcription>, please determine which type of attachment the child belongs to among secure, avoidant, anxious, and disorganized? And please explain the reason for your judgment.
[0060] (6) Input the features or results extracted by the feature extractor based on the multi-modal large prediction model into the assessment model for the attachment status of distressed children based on multi-round conversations (not shown in the figure). This method can be a feature-level fusion method based on the extracted multi-modal features or a decision-level fusion method based on the results output by the large model; the assessment model used can be a neural network with a sequential computing architecture, such as LSTM, GRU, etc.
[0061] The present invention first introduces a multi-modal feature extraction and fusion method based on pre-trained model multi-modal feature extraction - large language model multi-modal fusion understanding in the field of attachment automated assessment. In the video modality, pre-trained facial expression recognition and body movement recognition are used, and in the audio modality, pre-trained para-language and speech recognition models are used to extract the deep features of facial expressions, body movements, para-language, and speech respectively. The deep features are projected into the word embedding space of the large language model through a channel projector and multi-modal fusion understanding is carried out through the large language model, which is further used for attachment status assessment. When processing single-channel deep features, by introducing the constraint of attachment-related annotation data, the deep features are projected into the features in the word embedding space of the large language model, making the projected features more relevant to the attachment status and improving the effect of the large prediction model for multi-modal data fusion understanding.
[0062] Figure 5 The structural schematic diagram of some embodiments of the device for predicting the attachment status of children in distress based on a multimodal large model of the present invention is shown.
[0063] As Figure 5 shown, the device includes:
[0064] A dialogue turn segmentation module 100, configured to segment the dialogue turns between children in distress and doctors through an automated recognition method;
[0065] A feature extraction module 200, configured to extract expression, action, acoustic, and language channels from single-turn dialogue data, and perform data fusion through a large language model to obtain multimodal fusion features based on single-turn dialogue or the attachment status of children in distress;
[0066] An attachment status prediction module 300, configured to input multiple multimodal fusion features or the attachment status of children in distress obtained from single-turn dialogue data into an evaluation model for the attachment status of children in distress based on multi-turn dialogue to predict the attachment status of children in distress.
[0067] For the specific limitations of a device for predicting the attachment status of children in distress based on a multimodal large model, reference can be made to the limitations of a method for predicting the attachment status of children in distress based on a multimodal large model in the above text, which will not be elaborated here. Each module in the above device for predicting the attachment status of children in distress based on a multimodal large model can be implemented in whole or in part through software, hardware, and their combination. The above modules can be embedded in the processor of a computer device in hardware form or be independent of it, or can be stored in the memory of a computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.
[0068] The present invention also provides a computer device, which can be a terminal, and its internal structure diagram can be as Figure 6As shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements the above-mentioned method for predicting the attachment status of children in distress based on a multimodal large model. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc. Those skilled in the art can understand, Figure 6 The structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device to which the solution of the present invention is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0069] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned method for predicting the attachment status of children in distress based on a multimodal large model.
[0070] Those of ordinary skill in the art can understand that to implement all or part of the processes in the above method embodiments, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus memory bus (Rambus), direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0071] So far, the embodiments of the present invention have been described in detail. To avoid obscuring the concept of the present invention, some details well known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.
[0072] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are only for illustration and not for limiting the scope of the present invention. Those skilled in the art should understand that the above embodiments can be modified or some technical features can be equivalently replaced without departing from the scope and spirit of the present invention. The scope of the present invention is defined by the appended claims.
Claims
1. A method for predicting the attachment status of children in difficult situations based on a multimodal large model, including: Segmenting the dialogue turns between children in difficult situations and doctors through an automated recognition method; Taking expressions, actions, and acoustic channels extracted from single-turn dialogue data as inputs, using the output of the hidden layer of a pre-trained model as deep features, and projecting the deep features into the word embedding space used by a large language model through a neural network composed of Q-Former and a multi-layer perceptron to form word embedding space data; obtaining text features in the current turn of the dialogue through a speech recognition model; inputting the word embedding space data, text features, metadata of the current turn of the dialogue, and a prompt word for asking the large language model to predict the attachment status into the large language model, extracting the hidden layer features of the large language model as the multimodal fusion features of a single turn of dialogue, and outputting the result as the attachment status of the child in difficult situations in a single turn of dialogue; Inputting multiple of the above multimodal fusion features or the attachment status of children in difficult situations into an evaluation model for the attachment status of children in difficult situations based on multi-turn dialogue to predict the attachment status of children in difficult situations; The training method of the neural network is: Use (I, t) with labeled data for supervised training. The content of (I, t) is about the attachment of children in distress. Project the input information into the word embedding area related to attachment. Input I into the pre-trained model used, and then input it into the large language model through the neural network to generate text. Generate the error L through the text txt-gen Update the parameters of the neural network and update the formula: argmin Θ L txt-gen (LLM(P, Θ(PT_model(I))), t) I is the video or audio input, t is the text description, Θ is the neural network, LLM is the large language model, P is the prompt word, and PT_model is the pre-trained model used to extract features of the current channel; The segmentation of the dialogue turns between children in difficult situations and doctors includes: Segmenting the dialogue turns between doctors and children in difficult situations through a speaker recognition method based on voiceprint segmentation and clustering; The step of inputting multiple of the above multimodal fusion features or the attachment status of children in difficult situations into an evaluation model for the attachment status of children in difficult situations based on multi-turn dialogue to predict the attachment status of children in difficult situations includes: Performing feature-level or result-level fusion on multiple single-round multimodal understanding features or results extracted by the multimodal large model, and predicting the attachment status of children in difficult situations through a deep learning model based on a sequence calculation architecture.
2. The method for predicting the attachment status of children in difficult situations based on a multimodal large model according to claim 1, wherein The construction method of the evaluation model for the attachment status of children in difficult situations based on multi-turn dialogue is: Capturing the long-term dependencies in multi-turn dialogue through an attachment status evaluation model based on a sequence calculation architecture, so as to model the content of multi-turn dialogue.
3. The method for predicting the attachment status of children in difficult situations based on a multimodal large model according to claim 1, wherein The pre-trained model includes: For the expression channel, SynFace is used; For the action channel, C3D of OpenMM is used; For the acoustic channel, HuBERT is used.
4. The method for predicting the attachment status of children in difficult situations based on a multimodal large model according to claim 1, wherein The speech recognition model is Whisper.
5. A device for predicting the attachment status of children in difficult situations based on a multimodal large model, including: A dialogue turn segmentation module for segmenting the dialogue turns between children in difficult situations and doctors through an automated recognition method; A feature extraction module, which is used to extract expressions, actions, and acoustic channels from single-round dialogue data as inputs, and use the output of the hidden layer of a pre-trained model as deep features. Through a neural network composed of a Q-Former and a multi-layer perceptron, the deep features are projected into the word embedding space used by a large language model to form word embedding space data; obtain text features in the current round of dialogue through a speech recognition model; input the word embedding space data, text features, metadata of the current round of dialogue, and a prompt word for asking the large language model to predict the attachment status into the large language model, extract the hidden layer features of the large language model as the multi-modal fusion features of the single-round dialogue, and output the result as the attachment status of the child in distress in the single-round dialogue; An attachment status prediction module, which is used to input multiple of the multi-modal fusion features or the attachment status of the child in distress into an evaluation model for the attachment status of the child in distress based on multi-round dialogue to predict the attachment status of the child in distress; The training method of the neural network is as follows: Supervised training is carried out using the data with labels (I, t), where the content of (I, t) is about the attachment of children in difficult situations. The input information is projected into the word embedding area related to attachment. I is input into the pre-trained model used and then input into the large language model through the neural network to generate text; through the text generation error L txt-gen Update the parameters of the neural network, and the update formula is: argmin Θ L txt-gen (LLM(P, Θ(PT_model(I))), t) I is the video or audio input, t is the text description, Θ is the neural network, LLM is the large language model, P is the prompt word, and PT_model is the pre-trained model used to extract features of the current channel; The segmentation of the dialogue rounds between the child in distress and the doctor includes: Using a speaker recognition method based on voiceprint segmentation clustering to segment the dialogue rounds between the doctor and the child in distress; The step of inputting multiple of the multi-modal fusion features or the attachment status of the child in distress into an evaluation model for the attachment status of the child in distress based on multi-round dialogue to predict the attachment status of the child in distress includes: Performing feature-level or result-level fusion on multiple single-round multi-modal understanding features or results extracted by the multi-modal large model, and predicting the attachment status of the child in distress through a deep learning model based on a sequence calculation architecture.
6. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Multi-modal English speech ability assessment method
CN116862287A