A Multimodal Psychological State Evaluation System and Method for Prisoners Based on Incremental Learning

By adopting a multimodal depth model based on incremental learning in the psychological state assessment of prisoners, combining the fusion of video, text and scalar data and Bayesian incremental learning, the problems of low evaluation frequency, strong subjectivity and insufficient fusion of multimodal data in the existing technology are solved, and dynamic and accurate assessment of the psychological state of prisoners is achieved.

CN118366653BActive Publication Date: 2025-05-27HANGZHOU HUA TING TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410415918.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-08
Publication Date
2025-05-27
Estimated Expiration
2044-04-08

AI Technical Summary

Technical Problem

The prior art has problems in the assessment of psychological state of prisoners with low evaluation frequency, strong subjectivity and narrow coverage, and the fusion strategy of multimodal data is simple, and the interaction and synergistic relationship between different modes cannot be fully explored.

Method used

Using a multimodal depth model based on incremental learning, the fusion features are extracted using a specific encoder by acquiring and preprocessing video, text and scalar data, and an attention mechanism is introduced into the feature fusion module for fusion, and the fusion features are finally input into the Bayesian incremental learning module for dynamic evaluation.

Benefits of technology

The dynamic assessment of the psychological status of prisoners is achieved, efficient and reliable decision-making support is provided, the accuracy and comprehensiveness of the assessment is improved, and the model's adaptability to different individuals and environments is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118366653B_ABST
    Figure CN118366653B_ABST
Patent Text Reader

Abstract

The present application discloses a multi-modal psychological state assessment system and method for prisoners based on incremental learning. By acquiring multi-modal data of prisoners, including video, text, and scalar data, and preprocessing them to obtain features. The feature fusion module combines the attention mechanism for fusion to generate comprehensive features, and then continuously trains and iterates through the Bayesian incremental learning module to dynamically evaluate the psychological state. This system can deeply analyze, provide accurate information, dynamically track psychological changes, provide a basis for real-time decision-making, and continuously optimize the model to promote the progress of the assessment model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence. Specifically, the present application relates to a multimodal prisoner psychological state assessment system and method based on incremental learning. Background Art

[0002] Traditional psychological assessment methods mainly rely on manual regular testing and observation, which have defects such as low assessment frequency, strong subjectivity, and narrow coverage, making it difficult to grasp the psychological and behavioral changes of prisoners in a timely and comprehensive manner.

[0003] Some research works have attempted to extract facial, posture, voice and other modal features by analyzing the behavior videos of prisoners, and input them into the trained evaluation model to obtain mental health scores. However, these methods only consider video data and ignore other important information such as the prisoners' text communication and psychological test results. Therefore, the comprehensiveness and accuracy of the evaluation results are limited. Some researchers have also tried to integrate multimodal data such as video, audio, and text for behavioral analysis, and introduced attention mechanisms to improve the model's discrimination ability. However, the existing multimodal fusion strategy is relatively simple and fails to fully explore the interaction and synergy between different modalities. The fusion effect needs to be improved.

[0004] In summary, the prior art has the following obvious deficiencies:

[0005] There is a lack of full utilization and integration of multi-source heterogeneous data (such as videos, texts, test data, etc.), which limits the comprehensiveness and accuracy of the evaluation results; the evaluation models are usually static and unchanging, making it difficult to adapt to the complex and changeable psychological behaviors of prisoners, and the evaluation results lack dynamism and individualization. Summary of the invention

[0006] This application proposes a multimodal psychological state assessment system and method for prisoners based on incremental learning. This solution can continuously assess and adapt to the dynamic psychological and behavioral changes of prisoners through the close combination of incremental learning framework and multimodal deep model, and provide efficient and reliable decision support for psychological intervention. The technical solution is as follows:

[0007] According to one aspect of the present application, a method for multimodal psychological assessment of prison personnel based on incremental learning includes: obtaining multimodal data of prisoners, the multimodal data including video data, text data and scalar data; preprocessing the multimodal data to obtain preprocessed multimodal data, the preprocessed multimodal data including preprocessed video data features, preprocessed text data features and preprocessed scalar data features; sending the preprocessed multimodal data to corresponding encoders respectively, each type of data is processed by a specific encoder to extract high-level multimodal data features; fusing high-level features from different encoders in a feature fusion module, introducing an attention mechanism in the fusion process to strengthen the model's attention to key information, and obtaining comprehensive fusion features; sending the comprehensive fusion features as input to a Bayesian incremental learning module, and the Bayesian incremental learning module adapts to new data and situations through continuous training iterations, thereby realizing dynamic assessment of the psychological state of prisoners.

[0008] Collect surveillance videos, communication texts, vital signs data, etc. of prisoners and perform preprocessing. Extract frames and unify the size of video data; segment and add tags to text data; and normalize scalar data. Feature extraction: Use different encoders to extract features of each modality. Video features are extracted using C3D convolutional networks, and feature vectors are obtained through multi-layer convolution, pooling, and full connection; text features are extracted using the BERT model, and feature vectors are obtained through tag embedding, segment embedding, position embedding, and Transformer encoder; scalar features are extracted using multi-layer perceptron (MLP), and feature vectors are obtained through forward propagation and back propagation. Feature fusion: Input the extracted multimodal features into the fusion module and introduce the attention mechanism. The attention weights of each modality are calculated through a feedforward neural network, and the underlying features are weighted summed to obtain the fused multimodal representation. Incremental learning: Input the fused features into the Bayesian incremental learning module and continuously train the iterative model. Calculate the likelihood probability, update the posterior probability of the parameters, find the optimal parameters, and update the forgetting factor, and finally output the evaluation results.

[0009] According to one aspect of the present application, a multimodal prison personnel psychological assessment system based on incremental learning includes a data acquisition module for acquiring multimodal data of prisoners, wherein the multimodal data includes video data, text data and scalar data; a preprocessing module for preprocessing the multimodal data to obtain preprocessed multimodal data, wherein the preprocessed multimodal data includes preprocessed video data features, preprocessed text data features and preprocessed scalar data features; an advanced multimodal data feature extraction module for sending the preprocessed multimodal data to corresponding encoders respectively, wherein each type of data is processed by a specific encoder to extract advanced multimodal data features; a comprehensive fusion feature module for fusing advanced features from different encoders in a feature fusion module, wherein an attention mechanism is introduced in the fusion process to strengthen the model's attention to key information, and obtains comprehensive fusion features; and a comprehensive processing module for sending the comprehensive fusion features as input to a Bayesian incremental learning module, wherein the Bayesian incremental learning module adapts to new data and situations through continuous training iterations, thereby realizing dynamic assessment of the psychological state of prisoners.

[0010] According to one aspect of the present application, an electronic device includes: a processor and a memory, wherein the memory stores a computer program that can be called by the processor; the processor executes the multimodal method for assessing the psychological state of prisoners based on incremental learning in the background by calling the computer program stored in the memory.

[0011] According to one aspect of the present application, a computer-readable storage medium stores a rewritable computer program thereon; when the computer program is run on a computer device, the computer device executes the multimodal method for assessing the psychological state of prisoners based on incremental learning in the background.

[0012] The beneficial effects of the technical solution provided in this application are: by integrating multiple modal data such as video, text, test scores, etc., a comprehensive method is constructed to analyze the psychological state of prisoners. This fusion of multimodal data can conduct in-depth analysis from multiple angles and multiple levels, and can provide more accurate and comprehensive information than a single data source. The use of deep learning methods can better explore behavioral patterns, thereby better understanding and explaining the psychological state of prisoners. After introducing the incremental learning strategy, the system can update the model in time to cope with the arrival of new data. This method can dynamically track the psychological changes of prisoners, discover potential problems in a timely manner, and provide near real-time decision-making basis for psychological intervention. Incremental learning enables the system to continuously learn from new cases and improve the model's adaptability to different individuals and environments. The end-to-end training method allows the model to automatically extract the most effective feature representation, thereby better reflecting the psychological state of prisoners. In addition, incremental training also facilitates continuous iterative optimization of the system and promotes the continuous improvement of the evaluation model. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can obtain other drawings based on these drawings without creative work.

[0014] Figure 1 is a flow chart of the method involved in this application;

[0015] Figure 2 is a flow chart of S2 according to the method involved in this application;

[0016] Figure 3 is a flow chart of S3 according to the method involved in this application;

[0017] Figure 4 is a flow chart of S4 according to the method involved in this application;

[0018] Figure 5 is a flow chart of S5 according to the method involved in this application;

[0019] Figure 6 It is a system diagram involved in this application;

[0020] Figure 7 is a diagram of an electronic device according to the present application;

[0021] Figure 8 It is a schematic diagram of the structure of the computer-readable storage medium involved in this application. DETAILED DESCRIPTION

[0022] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present application.

[0023] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present disclosure refers to the presence of the features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0024] Example 1

[0025] The present application embodiment provides a multimodal method for assessing the psychological state of prisoners based on incremental learning, such as Figure 1 As shown, including:

[0026] S1: Obtain multimodal data of prisoners, including structured data such as monitoring data, communication records, physical data, psychological test results, etc. In prison scenarios, single-modal data often have one-sidedness and limitations, and it is difficult to accurately reflect the actual psychological condition of prisoners. The psychological state assessment of prisoners needs to rely on multi-faceted data support. Collecting multimodal data such as video surveillance, daily communication, and physiological signs can provide a basis for the subsequent construction of a comprehensive and dynamic psychological portrait of prisoners. By integrating multiple heterogeneous data sources, prisoners can be portrayed in three dimensions from multiple dimensions such as behavior, language, and physiology, greatly improving the accuracy and comprehensiveness of psychological assessments.

[0027] S2: Preprocessing the multimodal data to obtain preprocessed multimodal data, wherein the preprocessed multimodal data includes preprocessed video data features, preprocessed text data features, and preprocessed scalar data features, including frame extraction and unified input of video data, word segmentation and word embedding of text data, and normalization of scalar data.

[0028] like Figure 2In an exemplary embodiment, the obtained data is preprocessed, frame extraction is performed on the video data, unified input is performed, and text data is segmented and marked. The scalar data is normalized. S21. The video data is preprocessed, and the continuous video stream is sampled at fixed intervals and converted into a series of static image frames. All video frames are adjusted to the same size to ensure that the data received by the input layer of the network has a consistent shape. S22. The text data is preprocessed, the original communication text is segmented, the continuous character stream is converted into a series of independent semantic units, and special tags CLS and SEP are added. S23. The scalar data is preprocessed, and the numerical value of each feature dimension is normalized to a fixed range.

[0029] Through this step, targeted data preprocessing operations are performed to remove data redundancy and noise, convert different types of data into a processing format that the model can understand, and improve the learning efficiency and effect of the model. For example, frame extraction can capture key behaviors in the video, word segmentation and word embedding can reveal key information in the communication text and capture the implicit semantic relationship between words, and normalization can eliminate the dimensional differences of different signs. Preprocessed data is more suitable for input into deep learning models for feature learning and fusion.

[0030] S3: The preprocessed multimodal data are fed into corresponding encoders respectively, and each type of data is processed by a specific encoder to extract high-level multimodal data features;

[0031] like Figure 3 As shown, the preprocessed video data features are extracted using the C3D model, the preprocessed text data features are extracted using the Transformer-based BERT model, and the preprocessed scalar data features are extracted using a multi-layer perceptron.

[0032] S31. Extracting video data features using the C3D model; including:

[0033] S311, input the preprocessed video frame data sequence (N, D, H, W, C) into the convolutional neural network,

[0034] Where (N, D, H, W, C) is the input vector of the C3D model, N is the number of video clips processed at a time, C is the number of channels, usually 3, D is the number of consecutive frames in the video clip, H is the height of the frame, usually 112 or 128, W is the width of the frame, usually 112 or 128. Here N is 64, D is 16, H and W are 112, and C is 3.

[0035] S312, first input the first convolution block and perform convolution through 64 3*3*3 convolution kernels, with a step size of 1*1*1 and padding of 1*1*1. After the convolution is completed, it is activated by the ReLU function. After activation, a 2*2*2 pooling kernel is selected with a step size of 2*2*2. The convolution image is pooled using the maximum pooling method, and the data becomes (64, 8, 56, 56, 64).

[0036] S313, sequentially pass through the second to fifth convolution blocks, each block contains 128, 256, 512, 512 3*3*3 convolution kernels with a step size of 1*1*1 and a padding of 1*1*1. After the convolution is completed, it is activated by the ReLU function. After the activation is completed, a 2*2*2 pooling kernel is selected with a step size of 2*2*2. The convolution image is pooled using the maximum pooling method. The output sizes are (64, 4, 28, 28, 128), (64, 2, 14, 14, 256), (64, 1, 7, 7, 512), (64, 1, 4, 4, 512).

[0037] S314, flatten the obtained result and pass it through two fully connected layers, still using the ReLU function for activation to finally output the video data features. Add an additional global average pooling layer after the second fully connected layer to average the features in the time dimension to obtain a feature vector of size (1, 4096), which represents the features of the entire video clip.

[0038] S32. Use the BERT model to extract text data features; including:

[0039] S321, performing tag embedding, segment embedding, and position embedding on the preprocessed text information, and inputting the preprocessed and embedded tag sequence to be mapped into an embedding vector sequence E through an embedding layer;

[0040] E={e 1 ,e 2 ,…,e n},e i =E w x i +E p i+E s ;

[0041] Where: E w ∈R d×|V| is the token embedding matrix, |V| is the vocabulary size, E p ∈R d×n is the position embedding matrix, n is the sequence length, E s ∈R d is the sentence type embedding vector (used to distinguish sentence pairs), e i ∈Rd is the embedding vector of the i-th position

[0042] S322, use L-layer Transformer encoder to encode the embedded vector sequence E

[0043] H 0 =EH l = TransformerEncoder(H l-1 ),l=1,2,…,L;

[0044] in is the output of the encoder at layer l, is the Transformer encoder function, including multi-head self-attention and feedforward neural network.

[0045] Multi-head self-attention mechanism layer: Multi-head self-attention takes the output H of the previous layer l-1 Transformed into query Q, key K and value V, and then calculate the attention output

[0046] Q=H l-1 W Q

[0047] K=H l-1 W K

[0048] V=H l-1 W V

[0049] By calculating each head i The final output is MultiHead (H l-1 )

[0050]

[0051] MultiHead(H l-1 )=[head 1 ;head 2 ;…;head h ]W O

[0052] in is the projection matrix of query, key, and value, d k = d / h is the dimension of each attention head, h is the number of attention heads, The projection matrix for the multi-head attention output

[0053] Feedforward neural network layer: The feedforward neural network performs nonlinear transformation on the output of multi-head self-attention

[0054] FFN(x)=max(0,xW1 +b 1 )W 2 +b 2

[0055] in are the weights and biases of the first layer of feedforward network, b 2 ∈R d is the weight and bias of the second layer feedforward network, d ff is the hidden layer dimension of the feed-forward network.

[0056] S323: Output the high-level text data features extracted by the transform encoder.

[0057] S33, after preprocessing, the scalar data features are extracted using a multi-layer perceptron, by learning a function f MLP :R d →R k Map the original features into k-dimensional high-level feature representations, including:

[0058] S331, the input layer converts the original feature vector x i Passed to the first hidden layer The first hidden layer receives the output of the input layer and performs an affine transformation and nonlinear activation on it:

[0059]

[0060] in is the pre-activation value of the first hidden layer, is the weight matrix of the first hidden layer, is the bias vector of the first hidden layer, is the activation value of the first hidden layer, and σ(x) is a nonlinear activation function, such as ReLU, sigmoid, tanh, etc.

[0061] The next l-1 hidden layers transform the output of the previous layer in turn:

[0062]

[0063] in is the pre-activation value of the lth hidden layer, is the weight matrix of the lth hidden layer, is the bias vector of the lth hidden layer, is the activation value of the lth hidden layer.

[0064] The output layer takes the output of the last hidden layer The feature representation mapped to k dimensions is thus:

[0065]

[0066] where f i ∈R k is the k-dimensional feature representation extracted from the i-th data point, is the weight matrix of the output layer, b L+1 ∈R k is the bias vector of the output layer.

[0067] S332, back propagation algorithm calculates the gradient of loss function and model parameters, including:

[0068] Using mean square error (MSE) as the loss function:

[0069]

[0070] Where g(x) is the reconstruction function that represents the feature f i Map back to the original input space.

[0071] The gradient of the loss function to each layer parameter is calculated by back propagation algorithm

[0072] S333. Use the gradient descent algorithm to update the weights and biases of the model to minimize the loss function:

[0073]

[0074] Where η is the learning rate.

[0075] S334. Calculate the output of the last hidden layer of the MLP through forward propagation and use it as the feature representation of the data.

[0076] It should be noted that scalar data refers to data with a single value, which is one-dimensional and does not contain other dimensions besides direction or size. In the context of the multimodal prison personnel psychological state assessment system, scalar data may include but is not limited to the following types: Physiological indicators: such as heart rate, blood pressure, body temperature, respiratory rate, etc., which can reflect the physiological state of prisoners and may be related to their psychological state. Behavior counts: For example, the number of steps, the number of hand movements, or the frequency of other specific behaviors in a certain period of time. Psychological test scores: Prisoners may need to undergo regular mental health assessments, and the scores in the test results can be input into the system as scalar data. Time series data: such as data obtained from continuous monitoring, such as changes in sleep quality over time, daily activity patterns, etc. Questionnaire survey results: Data collected through questionnaires, such as emotional state, stress level, etc., are usually expressed in the form of scores. In this system, scalar data is preprocessed (such as normalization) and features are extracted using a multi-layer perceptron (MLP).

[0077] The C3D model can simultaneously extract spatial and temporal features from videos, capturing the subtle expressions, body movements, and other behavioral information of prisoners; the BERT model can model the contextual information of texts, learning implicit emotional patterns and psychological tendencies; the MLP model can fit the nonlinear relationships in physiological sign data, reflecting the physical and mental changes of prisoners. The comprehensive use of these advanced models can maximize the mining of psychological characteristics contained in multimodal data, forming a more accurate and comprehensive psychological portrait of prisoners.

[0078] S4: In the feature fusion module, high-level features from different encoders are fused. The attention mechanism is introduced in the fusion process to strengthen the model's attention to key information and obtain comprehensive fusion features.

[0079] like Figure 4 In an exemplary embodiment, S4, the extracted features are fused in a feature fusion module, and an attention mechanism is introduced into the fusion process. It includes:

[0080] S41, extracting the video feature f v (X v ) Text feature f t (X t ) Scalar feature f s (X s ) as the input of the fusion module:

[0081] f v (X v )=σ(W v *X v +b v )

[0082] ft (X t )=σ(W t *X t +b t )

[0083] f s (X s )=σ(W s *X s +b s )

[0084] Where W v ,W t ,W s are the weight matrices of video, text, and scalar features, respectively, and b v ,b t ,b s is the corresponding bias vector, σ is the nonlinear activation function, X v , X t , X s are the original features of video, text, and scalar features respectively.

[0085] S42, calculating the attention weight of each modality video attention weight, including:

[0086] Video attention weight:

[0087] Text attention weight:

[0088] Scalar attention weights:

[0089] where v v ,v t ,v s is the attention query vector, W v ′,W t ′,W s ′ is the attention weight matrix, b′ v ,b t ′,b s ′ is the attention bias vector, and the softmax function is used to normalize the attention weights so that their sum is 1.

[0090] S43, feature fusion, including:

[0091] Fusion of the first layer of features: The underlying features of each modality are weighted and summed according to the attention weights to obtain the first layer of fusion features h1:

[0092] h 1 =α v *f v (X v)+α t *f t (X t )+α s *f s (X s )

[0093] Extract high-level semantic features of each modality Video semantic features, including:

[0094] Video semantic features: g v (f v (X v ))=σ(U v *f v (X v )+c v )

[0095] Text semantic features: g t (f t (X t ))=σ(U t *f t (X t )+c t )

[0096] Scalar semantic feature: g s (f s (X s ))=σ(U s *f s (X s )+c s )

[0097] Among them U v ,U t ,U s is the weight matrix of high-level semantic features, c v ,c t ,c s is the corresponding bias vector.

[0098] Fusion of the second layer features: Fusion of the first layer features h1 with the high-level semantic features g of each modality v ,g t ,g s According to the weight matrix Perform weighted summation and then undergo nonlinear transformation The second layer fusion feature h2 is obtained, including:

[0099]

[0100] S44, obtain the comprehensive fusion feature y = σ (W o *h 2 +b o )

[0101] Multimodal feature fusion gives full play to the synergistic effect of data of different modalities, and builds a more comprehensive and in-depth representation of the psychological state of prisoners by capturing psychological information in multiple dimensions such as behavior, speech, and physiology. In text processing, the attention mechanism can help the model focus on the most relevant parts of the input text, thereby extracting the most useful features for the current task. For example, when processing an interview record of a prisoner, certain sentences or words may be more reflective of their psychological state and potential risk factors. The attention mechanism can automatically identify and give higher weights to these key parts, so that the model can generate a more refined and targeted representation.

[0102] S5: The comprehensive fusion features are sent as input to the Bayesian incremental learning module. The Bayesian incremental learning module adapts to new data and situations through continuous training iterations, thereby achieving dynamic assessment of the psychological state of prisoners.

[0103] like Figure 5 As shown, in an exemplary embodiment, S5, the fused features are used as inputs of the Bayesian incremental learning module, and the continuous training iterative model includes:

[0104] S51: Assuming that the data points are independent of each other, it should be noted that the data points represent comprehensive fusion features. The joint probability can be decomposed into the product of marginal probabilities to calculate the likelihood probability of the current model on the new data:

[0105]

[0106] Among them, D old =(x 1 ,y 1 ),(x 2 ,y 2 ),...,(x n ,y n ) is an existing data set, D new =(x 1 ′,y 1 ′),(x 2 ′,y 2 ′),...,(x m ′,y m ′) is the newly arrived data set, θ t are the current model parameters.

[0107] S52 calculates the new posterior probability:

[0108] p(θ|D old ,D new )∝p(D new |θ) λ(t)*p(θ|D old )

[0109] where p(θ|D old ) is the posterior probability of the current model based on old data, and λ(t) is the forgetting factor.

[0110] S53 maximizes the posterior probability, obtains the new model parameters, takes the logarithm of the posterior probability, and converts it into the maximum

[0111] The problem of minimizing the negative log-likelihood and negative log-prior:

[0112]

[0113] S54 uses stochastic gradient descent. Solve the above minimization problem:

[0114]

[0115] Where η is the learning rate, It means to find the gradient of θ.

[0116] S55 Update Forget Factor

[0117] λ(t+1)=f(λ(t),t)

[0118] Where f(·) is a predefined decreasing function. There is λ(t+1)=λ(t)*e (-t / s) .

[0119] S56 repeats steps S51-S55 until the convergence condition is met or the maximum number of iterations is reached.

[0120] The psychological state of prisoners is dynamically evolving, and new data will be continuously generated, which may contain the latest changes in the psychological characteristics of prisoners. Using incremental learning to continuously optimize the model can make the evaluation system constantly adapt to the latest distribution of the psychological state of prisoners, while avoiding catastrophic forgetting, and retaining the memory of previously stable psychological characteristics while learning new knowledge.

[0121] The Bayesian incremental learning module achieves a balance between model performance and update speed by introducing a forgetting factor, allowing the evaluation system to continue learning from the continuously accumulated prison data, and the characterization of the psychological state of prisoners can be updated in real time. Compared with traditional static evaluation methods, this method can keep up with the dynamic evolution of the psychological state of prisoners, and the evaluation results always maintain a high degree of timeliness and accuracy. At the same time, Bayesian learning avoids catastrophic forgetting, and can form long-term memory for some stable psychological characteristics of prisoners, which will not be forgotten by the impact of new data.

[0122] S6: Output the evaluation results to achieve dynamic evaluation of the psychological state of prisoners.

[0123] The model-predicted results of the inmates' psychological state assessment are clearly presented to prison administrators or relevant experts, and an interpretable and operational analysis report is generated to provide practical suggestions and basis for prison management decisions.

[0124] The present invention proposes a multimodal psychological state assessment method for prisoners based on incremental learning, which effectively improves the defects of traditional psychological assessment methods such as strong subjectivity, low assessment frequency, narrow coverage, and the shortcomings of existing multimodal fusion strategies such as simplicity and static assessment models.

[0125] This invention overcomes the one-sidedness and limitations of single modality data by introducing multi-source heterogeneous data such as video, text, and physiological data, and adopts advanced deep learning models for feature extraction and fusion. It achieves a comprehensive characterization of prisoners' behavior, language, physiology and other dimensions, greatly improving the accuracy and comprehensiveness of psychological assessment.

[0126] At the same time, the invention also designed an incremental learning strategy based on Bayesian inference, which gives the evaluation model the ability to adjust parameters according to the real-time psychological dynamic changes of prisoners, and introduces forgetting factors to adapt to the complex and changeable psychological behavior patterns of prisoners, thereby enhancing the dynamics and individualization of the evaluation results.

[0127] Example 2

[0128] This embodiment is carried out based on the embodiment 1, and the parts that are the same as those in the embodiment are not further elaborated here.

[0129] like Figure 6 As shown, a multimodal prisoner psychological state assessment system based on incremental learning, the system includes:

[0130] A data acquisition module, used to acquire multimodal data of prisoners, wherein the multimodal data includes video data, text data and scalar data;

[0131] A preprocessing module, used for preprocessing the multimodal data to obtain preprocessed multimodal data, wherein the preprocessed multimodal data includes preprocessed video data features, preprocessed text data features and preprocessed scalar data features;

[0132] The advanced multimodal data feature extraction module is used to send the pre-processed multimodal data to the corresponding encoders respectively, and each type of data is processed by a specific encoder to extract the advanced multimodal data features;

[0133] Comprehensive fusion feature module, which is used to fuse high-level features from different encoders in the feature fusion module. The attention mechanism is introduced in the fusion process to strengthen the model's attention to key information and obtain comprehensive fusion features;

[0134] The comprehensive processing module is used to send the comprehensive fusion features as input to the Bayesian incremental learning module. The Bayesian incremental learning module adapts to new data and situations through continuous training iterations, thereby realizing dynamic assessment of the psychological state of prisoners.

[0135] Example 3

[0136] A schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 7 As shown, according to another aspect of the present application, an electronic device 100 is also provided. The electronic device 100 may include one or more processors and one or more memories. The memories store computer readable codes, and when the computer readable codes are executed by one or more processors, the multimodal prisoner psychological state assessment method based on incremental learning as described above can be implemented.

[0137] The method or device according to the embodiment of the present application can also be used by Figure 7 The electronic device architecture shown in FIG. Figure 7 As shown, the electronic device 100 may include a bus 101, one or more CPUs 102, a ROM 103, a RAM 104, a communication port 105 connected to a network, an input / output component 106, a hard disk 107, etc. The storage device in the electronic device 100, such as the ROM 103 or the hard disk 107, may store the implementation of the multimodal prisoner psychological state assessment method based on incremental learning provided in the present application. The implementation of a multimodal method for assessing the psychological state of prisoners based on incremental learning may include the following steps: obtaining multimodal data of prisoners, wherein the multimodal data includes video data, text data and scalar data; preprocessing the multimodal data to obtain preprocessed multimodal data, wherein the preprocessed multimodal data includes preprocessed video data features, preprocessed text data features and preprocessed scalar data features; sending the preprocessed multimodal data to corresponding encoders respectively, wherein each type of data is processed by a specific encoder to extract high-level multimodal data features; fusing high-level features from different encoders in a feature fusion module, introducing an attention mechanism in the fusion process to strengthen the model's attention to key information, and obtaining comprehensive fusion features; sending the comprehensive fusion features as input to a Bayesian incremental learning module, wherein the Bayesian incremental learning module adapts to new data and situations through continuous training iterations, thereby achieving dynamic assessment of the psychological state of prisoners.

[0138] Furthermore, the electronic device 100 may further include a user interface 108. Figure 7 The architecture shown is only exemplary and can be omitted according to actual needs when implementing different devices. Figure 7 One or more components of an electronic device are shown.

[0139] Example 4

[0140] Figure 8 Schematic diagram of a computer-readable storage medium structure provided by an embodiment of the present application. Figure 8 As shown, a computer-readable storage medium 200 according to an embodiment of the present application is shown. Computer-readable instructions are stored on the computer-readable storage medium 200. When the computer-readable instructions are executed by the processor, the multimodal method for assessing the psychological state of prisoners based on incremental learning according to the embodiment of the present application described with reference to the above figures can be implemented. The computer-readable storage medium 200 includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include random access memory (RAM) and cache memory (cache), etc. Non-volatile memory may include read-only memory (ROM), hard disk, flash memory, etc.

[0141] In addition, according to the implementation of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium, the non-transitory machine-readable storage medium stores machine-readable instructions, the machine-readable instructions can be run by a processor to execute instructions corresponding to the method steps provided by the present application, and when the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.

[0142] The methods, apparatuses, and devices of the present application may be implemented in many ways. For example, the methods, apparatuses, and devices of the present application may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present application are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present application may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers recording media storing programs for executing the method according to the present application.

[0143] In addition, the parts of the above-mentioned technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive redundancy.

[0144] The specific implementation modes described above further describe the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only a specific implementation mode of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present embodiment should be included in the protection scope of the present embodiment.

[0145] The above preset parameters or preset thresholds are all set by technicians in this field according to actual conditions or obtained through large-scale data simulation.

[0146] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A multimodal method for assessing the psychological state of prisoners based on incremental learning, characterized in that: The method comprises the following steps: S1. Acquire multimodal data of prisoners, where the multimodal data includes video data, text data, and scalar data; S2, preprocessing the multimodal data to obtain preprocessed multimodal data, wherein the preprocessed multimodal data includes preprocessed video data features, preprocessed text data features, and preprocessed scalar data features; S3, sending the preprocessed multimodal data to the corresponding encoders respectively, and each type of data is processed by a specific encoder to extract high-level multimodal data features; S4. In the feature fusion module, high-level features from different encoders are fused. The attention mechanism is introduced in the fusion process to strengthen the model's attention to key information and obtain comprehensive fusion features. S5. Send the comprehensive fusion features as input to the Bayesian incremental learning module. The Bayesian incremental learning module adapts to new data and situations through continuous training iterations, thereby achieving a dynamic assessment of the psychological state of prisoners. The attention mechanism is introduced in the fusion process to strengthen the model's attention to key information and obtain comprehensive fusion features, including: S41, extracting video features , text features , scalar features As the input of the feature fusion module: in are the weight matrices of video, text, and scalar features respectively, is the corresponding bias vector, is a nonlinear activation function, , , are the original features of video, text, and scalar features respectively; S42, calculating the attention weight of each modality, including: Video attention weight: ; Text attention weight: ; Scalar attention weights: ; in is the attention query vector, is the attention weight matrix, is the attention bias vector, and the softmax function is used to normalize the attention weights so that their sum is 1; S43, feature fusion, including: Fusion of the first layer of features: The underlying features of each modality are weighted and summed according to the attention weights to obtain the first layer of fusion features h1: Extract high-level semantic features of each modality, including: Video semantic features: Text semantic features: Scalar semantic features: in is the weight matrix of high-level semantic features, is the corresponding bias vector; Fusion of the second layer features: Fusion of the first layer features h1 with the high-level semantic features of each modality According to the weight matrix Perform weighted summation and then undergo nonlinear transformation The second layer fusion feature h2 is obtained, including: S44. Obtain comprehensive fusion features .

2. According to claim 1, a multimodal method for assessing the psychological state of prisoners based on incremental learning is characterized in that: The preprocessed multimodal data are respectively sent to the corresponding encoders, including extracting the preprocessed video data features using the C3D model, extracting the preprocessed text data features using the Transformer-based BERT model, and extracting the preprocessed scalar data features using a multi-layer perceptron.

3. According to claim 2, a multimodal method for assessing the psychological state of prisoners based on incremental learning is characterized in that: The steps for extracting features of preprocessed text data using the Transformer-based BERT model include: The preprocessed text information is embedded in tags, segments, and positions; the input preprocessed and embedded tag sequence is mapped into an embedding vector sequence E through the embedding layer; Use L-layer Transformer encoder to encode the embedding vector sequence E; ; in For the The output of the layer encoder, is the Transformer encoder function, including multi-head self-attention and feedforward neural network; Multi-head self-attention mechanism layer: Multi-head self-attention takes the output of the previous layer Transformed into query Q, key K and value V, and then calculate the attention output By calculating each The final output is multi-head self-attention in is the projection matrix of query, key, and value, is the dimension of each attention head, h is the number of attention heads, is the projection matrix of the multi-head attention output; Feedforward neural network layer: The feedforward neural network performs nonlinear transformation on the output of multi-head self-attention in are the weights and biases of the first layer of feedforward network, are the weights and biases of the second layer feedforward network, is the hidden layer dimension of the feedforward network; The output is passed through a transform encoder to extract high-level text data features.

4. According to claim 2, a multimodal method for assessing the psychological state of prisoners based on incremental learning is characterized in that: The preprocessed scalar data features are extracted using a multi-layer perceptron, including: The back-propagation algorithm calculates the gradient of the loss function and the model parameters, including: Using mean square error (MSE) as the loss function: in Represent the feature as the reconstruction function Map back to the original input space; The gradient of the loss function to each layer parameter is calculated by back propagation algorithm S333. Use the gradient descent algorithm to update the weights and biases of the model to minimize the loss function: in is the learning rate.

5. According to claim 1, a multimodal method for assessing the psychological state of prisoners based on incremental learning is characterized in that: The comprehensive fusion features are sent as input to the Bayesian incremental learning module for continuous training iterations, including: Decompose the joint probability into the product of marginal probabilities and calculate the likelihood probability of the current model on the new data: in, is the newly arrived data set, is the current model parameter; S52 calculates the new posterior probability: in is the posterior probability of the current model based on old data For existing data sets; S53 maximizes the posterior probability, obtains the new model parameters, takes the logarithm of the posterior probability, and transforms it into a problem of minimizing the negative log-likelihood and negative log-prior: S54 uses stochastic gradient descent to solve the above minimization problem: in is the learning rate, Express Find the gradient, For the forgetting factor; S55 Update Forget Factor in ,have ; S56 repeats steps S51-S55 until the convergence condition is met or the maximum number of iterations is reached.

6. A multimodal prisoner psychological state assessment system based on incremental learning, the system is implemented based on the method described in any one of claims 1 to 5, characterized in that: A data acquisition module, used to acquire multimodal data of prisoners, wherein the multimodal data includes video data, text data and scalar data; A preprocessing module, used for preprocessing the multimodal data to obtain preprocessed multimodal data, wherein the preprocessed multimodal data includes preprocessed video data features, preprocessed text data features and preprocessed scalar data features; The advanced multimodal data feature extraction module is used to send the pre-processed multimodal data to the corresponding encoders respectively, and each type of data is processed by a specific encoder to extract the advanced multimodal data features; Comprehensive fusion features are used to fuse high-level features from different encoders in the feature fusion module. The attention mechanism is introduced in the fusion process to strengthen the model's attention to key information and obtain comprehensive fusion features. The comprehensive processing module is used to send the comprehensive fusion features as input to the Bayesian incremental learning module. The Bayesian incremental learning module adapts to new data and situations through continuous training iterations, thereby realizing dynamic assessment of the psychological state of prisoners.

7. An electronic device, characterized in that: include: processor and memory, wherein The memory stores a computer program that can be called by the processor; The processor executes the multimodal prisoner psychological state assessment method based on incremental learning as described in any one of claims 1 to 5 in the background by calling the computer program stored in the memory.

8. A computer-readable storage medium, characterized in that: A rewritable computer program is stored thereon; When the computer program is executed on a computer device, the computer device executes any one of the incremental learning-based multimodal prisoner psychological state assessment methods of claims 1-5 in the background.

Citation Information

Patent Citations

  • Detainee emotion recognition method for multi-modal feature fusion based on Transformer, equipment, and medium

    CN113822192A