Artificial intelligence-based employee psychological portraying method and system

By collecting 3D point cloud images and audio data from sandbox games, and utilizing hierarchical feature encoding technology and LSTM models, the problem of the inability of existing technologies to deeply explore subconscious psychological characteristics has been solved, enabling automated and dynamic prediction of employee psychological profiles.

CN121075652APending Publication Date: 2025-12-05XICHANG COLLEGE +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511250756.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing technologies struggle to fully capture the subconscious psychological characteristics in sandplay games, lack in-depth mining and dynamic analysis of multimodal data, and are unable to construct accurate employee psychological profiles.

Method used

By collecting 3D point cloud image data and audio data during the sandbox game process, hierarchical feature coding technology (open coding, main axis coding, core coding) is used to deeply mine multimodal data, construct a time series prediction model based on LSTM, and generate employee psychological profiles.

Benefits of technology

It enables the automatic extraction and dynamic prediction of employees' subconscious psychological characteristics, providing objective and comprehensive psychological health assessment and risk warning support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121075652A_ABST
    Figure CN121075652A_ABST
Patent Text Reader

Abstract

The invention discloses an employee psychological portraying method and system based on artificial intelligence, and relates to artificial intelligence. The method comprises the following steps: collecting multi-modal data including image data and audio data; performing hierarchical feature coding on the multi-modal data to obtain a multi-level psychological feature vector set; wherein the hierarchical feature coding comprises open coding, main shaft coding and core coding; according to the multi-level psychological feature vector set, constructing a time sequence prediction model based on LSTM; and generating an employee psychological portrait by using the time sequence prediction model. In order to solve the problem that subconscious psychological features are difficult to capture in psychological assessment of employees in the prior art, three-dimensional point cloud image data and audio data in a sand table game process are collected, and a hierarchical feature coding technology (including open coding, spindle coding and core coding) is used for carrying out deep mining and the like on multi-modal data; and subconsciousness psychological projection characteristics are fully excavated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence, in particular to a method and system for employee psychological portrait based on artificial intelligence. BACKGROUND

[0002] Subconscious psychological characteristics refer to the psychological activities and contents below the individual's consciousness threshold, including suppressed emotions, implicit cognitive patterns, and deep personality structures. These characteristics are often the fundamental factors driving behavior and decision-making. Psychological research shows that 95% of human cognitive activities occur at the subconscious level, but traditional assessment methods cannot access this area. The main reasons include: (1) direct questioning cannot obtain subconscious content, and the subject may not be aware of the existence of these characteristics; (2) defense mechanisms make individuals tend to hide or beautify their true psychological state; (3) the limitations of language expression make it difficult to accurately convey complex inner experiences; (4) single modality data cannot fully reflect the multidimensionality of psychological activities.

[0003] Sandplay therapy, as a projective psychological technique, allows individuals to freely place miniature models in the sand to express their inner world, bypassing the conscious level of defense and directly presenting subconscious content. Based on Jung's analytical psychology theory, sandplay works are considered symbolic expressions of psychological content, with spatial layout, object selection, and placement order all containing rich psychological information. However, traditional sandplay analysis mainly relies on the subjective interpretation of therapists, lacks objective quantitative standards, and ignores dynamic information and multi-modal characteristics during the game process.

[0004] Existing computer-aided psychological assessment techniques, although introducing machine learning methods, still have the following limitations: (1) Most systems only analyze the final static images of the sand tray, ignoring the temporal information and dynamic changes of the placement process; (2) Lack of synchronous analysis of speech data during the game process, unable to capture the relevance of language and behavior; (3) Feature extraction remains at the surface level, failing to deeply explore the internal relationship between cross-modal features; (4) Lack of a systematic feature encoding framework, making it difficult to build a multi-level psychological feature representation; (5) Unable to realize dynamic prediction and trend analysis of psychological state.

[0005] In particular, in the extraction of subconscious characteristics, existing technologies fail to fully utilize the projective nature of sandplay: the distribution of objects in different spatial regions reflects the allocation pattern of psychological energy; the spatial relationship between objects implies inner conflict or integration state; the placement order embodies the priority of psychological construction; symbolic vocabulary in language narratives reveals deep psychological content; subtle changes in sound characteristics reflect the underlying emotional flow. These multi-modal, multi-level information need to be effectively extracted and integrated through more refined and systematic analysis methods.

[0006] Therefore, by deeply mining the mapping relationship between behavior performance and inner activity, a more accurate and in-depth employee psychological portrait is constructed, thereby providing scientific support for enterprise human resource management and employee psychological health maintenance. SUMMARY

[0007] In view of the fact that the existing employee psychological assessment is difficult to capture the subconscious psychological characteristics, the present application provides an employee psychological portrait method and system based on artificial intelligence, which fully mines the subconscious psychological projection characteristics by collecting three-dimensional point cloud image data and audio data in the sand table game process, and using hierarchical feature coding technology (including open coding, main axis coding, and core coding) to deeply mine the multi-modal data.

[0008] One aspect of the present application provides an employee psychological portrait method based on artificial intelligence, comprising: collecting multi-modal data including image data and audio data; performing hierarchical feature coding on the multi-modal data to obtain a multi-level psychological feature vector set; wherein the hierarchical feature coding includes open coding, main axis coding and core coding; constructing a time series prediction model based on LSTM according to the multi-level psychological feature vector set; and generating an employee psychological portrait using the time series prediction model.

[0009] Further, the image data includes three-dimensional point cloud data of the sand table game scene, object recognition results and spatial position relationship matrix; the audio data includes audio text transcription content, acoustic feature parameters and emotion annotation results; the acoustic feature parameters include pitch, speech rate and pause duration;

[0010] The object recognition result refers to a data set obtained by automatically detecting and classifying various objects in the sand table game scene through computer vision algorithms (such as YOLOv5), including object class labels (such as people, animals, buildings, plants, vehicles, etc.), three-dimensional bounding box coordinates, object IDs and detection confidence. The result analyzes the types of objects selected by the employee in the sand table game and their psychological symbolic meanings.

[0011] The spatial position relationship matrix refers to a mathematical representation describing the spatial relationship between objects in the sand table, which is specifically represented as a three-dimensional relationship tensor in the present application wherein k is the number of objects. The matrix contains three layers of information: the first layer is a normalized inter-object horizontal distance matrix; the second layer is a buried relationship matrix, which records the degree of object coverage by sand (0 indicates no burial, 0.5 indicates partial burial, and 1 indicates deep burial); the third layer is a stacking relationship matrix, which identifies the vertical contact relationship between objects.

[0012] Acoustic feature: refers to the physical acoustic index reflecting the speaker's psychological state extracted from the audio signal, including: fundamental frequency (F0, reflecting the pitch), short-time energy (reflecting the sound intensity), mel-frequency cepstral coefficient (MFCC, capturing the timbre feature), speech rate (the number of words per unit time), pause duration (the duration of the silent segment), decision duration (the interval between adjacent objects), emotional change feature (standard deviation of speech rate), and introspection degree feature (proportion of silent duration).

[0013] Emotion annotation result: refers to the 7-dimensional psychological state probability distribution vector output by the trained multi-label classification model after emotion recognition of the audio, including the probability values (value range 0-1) of seven basic emotional states of anxiety, depression, anger, fear, sadness, happiness and calm. The model is trained based on a large-scale psychological counseling corpus and can identify the emotional changes and psychological states of the speaker during the sandplay process.

[0014] Further, hierarchical feature encoding is performed on the multi-modal data, including: performing open encoding to extract features from the multi-modal data to obtain primary features; performing principal axis encoding to construct the association between the primary features and generate associated features; performing core encoding to cluster and reduce the dimension of the associated features to obtain a set of multi-level psychological feature vectors; wherein the open encoding: refers to the first stage of hierarchical feature encoding, which is a process of preliminary feature extraction on the original multi-modal data. In this application, open encoding independently processes the image and audio data of the sandplay game through a deep learning algorithm.

[0015] Principal axis encoding: refers to the second stage of hierarchical feature encoding, which is a process of mining the internal association between different modal features. In this application, the principal axis encoding discovers the consistency pattern of behavior and language by calculating the cross-modal correlation.

[0016] Core encoding: refers to the third stage of hierarchical feature encoding, which is a process of high-level abstraction and structured organization of associated features. In this application, core encoding extracts core psychological features through a machine learning algorithm.

[0017] Further, the open encoding is performed to extract features from the multi-modal data to obtain primary features, including: using a three-dimensional point cloud processing algorithm to extract spatial features from the image data of the sandplay game scene to obtain an image feature vector; using natural language processing technology to extract semantic and acoustic features from the audio data to obtain an audio feature vector; and splicing the image feature vector and the audio feature vector to obtain the primary features.

[0018] Further, the image data of the sand table game scene is processed by a three-dimensional point cloud processing algorithm to extract spatial features and obtain an image feature vector, including: collecting three-dimensional point cloud data of the sand table game by using a depth camera, and dividing the point cloud data into nine regions based on sand table boundary coordinates; wherein the nine regions include a central region, four corner regions, and four edge regions; local features of each region are extracted using PointNet++ respectively, and are spliced to form a partition feature vector to obtain the psychological projection differences of different regions; game objects in the sand table are identified and the corresponding placement timing is recorded to construct a timing encoding vector T representing the placement order of the objects, and the object orientation angle θ is extracted to generate a composite feature vector containing the object category, timing, and orientation; a multi-level position matrix is constructed according to the spatial relationship of the sand table, the first layer calculates the horizontal distance between the objects and normalizes according to the size of the sand table; the second layer identifies the degree of coverage of the objects by the sand, whether partially or completely; the third layer detects the vertical contact between the objects to generate a three-dimensional relationship tensor R; the partition feature vector, the composite feature vector, and the three-dimensional relationship tensor R are spliced to obtain the image feature vector.

[0019] The central region, the four corner regions, and the four edge regions refer to nine sub-regions obtained by dividing the sand table game space in a nine-square grid manner. The central region is located at the center of the sand table and represents the core self in the psychological projection theory; the four corner regions include the upper left corner (past), the upper right corner (future), the lower left corner (subconscious), and the lower right corner (reality), reflecting the psychological projection of time and consciousness; the four edge regions include the upper edge (ideal), the lower edge (foundation), the left edge (intrinsic), and the right edge (extrinsic), embodying the relationship between the individual and the external world. This division method is based on sand table game psychology theory, and the placement of objects in different regions has specific psychological symbolic meanings.

[0020] Local features refer to high-dimensional feature vectors extracted from the point cloud data of each region using the PointNet++ deep learning network. Specifically, they include the point cloud density, height distribution, shape contour, and object aggregation degree in the region, as well as geometric and topological features. PointNet++ can capture the local point cloud patterns of each region through multi-level feature aggregation, outputting a feature vector with a dimension of 256. These features reflect the creative tendencies and spatial usage patterns of employees in different psychological projection regions.

[0021] Timing encoding vector T refers to a one-dimensional vector recording the time sequence of object placement in the sand table game, where n is the total number of objects. The i-th element T[i] in the vector represents the placement timestamp or sequence number of the i-th object. For example, T = [1, 3, 2, 5, 4] represents the placement order of five objects. This vector analyzes the thought process, decision-making pattern, and psychological construction process of employees, and the placement order often reflects the priority of the subconscious and the allocation of psychological energy.

[0022] Three-dimensional relationship tensor R: refers to a three-order tensor describing the spatial relationship between objects in the sand table, with dimensions where k is the number of objects. The tensor contains three relationship matrices: R[ :, :, 0] is the horizontal distance matrix, and the element R[i, j, 0] represents the normalized Euclidean distance between objects i and j; R[ :, :, 1] is the buried relationship matrix, and the element R[i, j, 1] represents the degree of object j being buried by object i (0 = not buried, 0.5 = partially buried, 1 = completely buried); R[ :, :, 2] is the stacking relationship matrix, and the element R[i, j, 2] represents whether object i is above object j (0 = no contact, 1 = vertical contact). The tensor comprehensively depicts the three-dimensional spatial relationship between objects.

[0023] Further, natural language processing techniques are used to extract semantic and acoustic features from audio data to obtain an audio feature vector, including: performing speech recognition on the audio data during the sand table game process and converting it into text data, using a BERT model to encode the text data, extracting context features related to object operations through an attention mechanism, and generating narrative semantic vectors; performing part-of-speech tagging and dependency syntax analysis on the text data, identifying anthropomorphism, symbolism, and emotional projection words, calculating the TF-IDF weights of each word, and selecting the top % words by weight to construct a psychological projection feature vector; performing frame processing on the audio signal to extract game process acoustic feature vectors; aligning the audio time axis with the object operation time axis through a dynamic time warping algorithm, calculating the timing correlation coefficient of language expression and behavior action, and generating a timing alignment vector; inputting the game process acoustic feature vector into a multi-label classification model trained based on a large psychological counseling corpus, and outputting a 7-dimensional psychological state probability distribution vector; using a feature concatenation method to concatenate the narrative semantic vector, the psychological projection feature vector, the game process acoustic feature vector, the timing alignment vector, and the psychological state probability distribution vector, and performing feature fusion through a fully connected layer to obtain an audio feature vector.

[0024] wherein the narrative semantic vector: refers to a high-dimensional vector representation (usually 768 dimensions) obtained by deep semantic encoding of the speech transcription text of the employee during the sand table game process through the BERT pre-trained language model. The vector captures the context semantic information related to object operations through the self-attention mechanism of BERT, such as the protection awareness and safety needs implied in the sentence "I put the small house in the corner to protect it", and can understand the psychological motivation behind language expression.

[0025] Projective vocabulary: refers to a specific type of vocabulary in sandplay narration that has psychological projection significance, including: anthropomorphism vocabulary (such as "the tree is crying", "the house is lonely"), which gives inanimate objects human characteristics; symbolic vocabulary (such as "this is my castle", "the fence represents protection"), which expresses metaphorical and symbolic meanings; emotional vocabulary (such as "fear", "warmth", "suppression"), which directly reflects emotional states. These words are linguistic markers of psychological projection, reflecting the speaker's unconscious psychological content.

[0026] Psychological projection feature vector: refers to a quantitative feature representation based on projective vocabulary. Specifically, by calculating the TF-IDF (Term Frequency-Inverse Document Frequency) weight of each projective vocabulary, selecting the top 20% of high-value words, and grouping their TF-IDF values into a sparse vector, the vector quantifies the intensity of psychological projection and the distribution of projection content in the sandplay game. The dimension is equal to the number of selected projective vocabularies.

[0027] Dynamic Time Warping (DTW): refers to an algorithm for measuring the similarity of two time series, which can handle the problem of nonlinear alignment on the time axis. In this application, DTW performs flexible matching between the audio time axis (language expression) and the object operation time axis (behavior action), allowing for changes in speech rate and pauses, and finding the optimal time alignment path to accurately analyze the timing correspondence between speech and behavior.

[0028] Time alignment vector: refers to a quantitative vector calculated by the DTW algorithm, containing the timing correlation indicators between language expression and object operation. Specifically, it includes: cumulative distance of alignment path (reflecting overall synchronization degree), local alignment coefficient (each object operation corresponding language segment correlation), time offset (language relative to behavior advance or lag degree) etc. The vector dimension is usually 3 times the number of objects, fully characterizing the consistency of speech and action.

[0029] Multi-label classification model trained based on large-scale psychological counseling corpus: refers to a deep learning model trained using a large-scale psychological counseling dialogue audio dataset (containing tens of thousands of hours of labeled data). The model uses a CNN-LSTM architecture, with acoustic feature sequences as input, convolutional layers to extract local patterns, LSTM layers to model temporal dependencies, and finally a sigmoid-activated fully connected layer to output the probabilities of multiple emotion labels. The model's emotion recognition accuracy in the psychological counseling scenario is above 85%.

[0030] 7-dimensional psychological state probability distribution vector: refers to the vector containing the probability values of 7 basic emotional states output by the multi-label classification model, that is, [P_ anxiety, P_ depression, P_ anger, P_ fear, P_ sadness, P_ happiness, P_ calmness], each probability value ranges from [0, 1] and is independent of each other (multiple emotions can exist at the same time). For example, [0.7, 0.3, 0.1, 0.2, 0.4, 0.2, 0.1] indicates that the employee is mainly anxious (70%) and accompanied by a certain degree of depression (30%) and sadness (40%) at that moment.

[0031] Further, the audio signal is subjected to frame processing, and a game process acoustic feature vector is extracted, including: extracting basic acoustic parameters including fundamental frequency, energy, and mel-frequency cepstral coefficient with a preset time step as a window; calculating the time interval between adjacent object placement actions as a decision duration feature; calculating the standard deviation of speech rate to obtain an emotional change feature by using a sliding window method; marking the silent segment and calculating the duration ratio to obtain the introspection degree feature using a speech activity detection algorithm; combining the basic acoustic parameters, the decision duration feature, the emotional change feature, and the introspection degree feature to generate a game process acoustic feature vector.

[0032] Among them, the decision duration feature: refers to the time interval sequence between adjacent two object placement actions in the sand table game process. The specific calculation method is: let the placement time of the i-th object be , then the decision duration . This feature forms a duration vector , where n is the total number of objects. The decision duration reflects the depth of thinking, the degree of decision hesitation and the complexity of cognitive processing of the employee. Longer decision duration may indicate inner conflict or deep thought, and shorter decision duration may indicate impulsive or clear thinking.

[0033] Emotional change feature: refers to a numerical feature that quantifies the degree of change in the emotional state of the employee by analyzing the fluctuation of speech rate. Specifically, the speech rate (number of words per second) in each window is calculated using the sliding window method (window size of 10 seconds, step size of 5 seconds), and then the standard deviation σ_ speech rate of all window speech rate values is calculated. The larger the feature value, the more intense the speech rate fluctuation, indicating that the emotional ups and downs are large or the psychological state is unstable; the smaller the feature value, the more stable the speech rate, indicating that the emotion is relatively stable. This is a scalar feature that can effectively capture the dynamic changes in emotion during the game process.

[0034] The introspection degree feature refers to an index quantifying the degree of silent thinking of the employee during the sand table game. The voice activity detection (VAD) algorithm is used to identify the silent segments in the audio (segments with energy lower than a threshold and lasting more than 0.3 seconds), and the proportion of the total silent time length to the entire game time length is calculated: introspection degree = Σ silent segment time length / total game time length. The ratio ranges from 0 to 1, and a higher value (such as >0.4) indicates that the employee tends to introspect and think deeply, and a lower value indicates that the employee tends to talk while doing, with a high degree of externalization of thought.

[0035] The game process acoustic feature vector refers to a multi-dimensional vector representation that comprehensively reflects the sound features during the entire sand table game process. The vector is composed of four parts: (1) statistical values of basic acoustic parameters, including mean and variance of fundamental frequency (2 dimensions), mean and variance of energy (2 dimensions), and mean vector of MFCC coefficients (13 dimensions); (2) statistical values of decision time length features, including average decision time length, maximum / minimum decision time length, and decision time length standard deviation (4 dimensions); (3) emotion change feature (1 dimension); (4) introspection degree feature (1 dimension). Finally, a 23-dimensional feature vector is formed, which fully characterizes the sound performance and psychological rhythm of the employee during the game process.

[0036] Further, the principal axis encoding is performed to build the association between the primary features, and the associated features are generated, including: calculating the cross-correlation coefficient between the composite feature vector in the image feature vector and the time sequence alignment vector in the audio feature vector, building a cross-modal time sequence association matrix to capture the consistency of object operation behavior and language expression; calculating the correlation between the partition feature vector in the image feature vector and the psychological projection feature vector in the audio feature vector through cosine similarity, generating a space-semantic association vector to reflect the correspondence between the sand table space layout and the psychological projection content; based on the three-dimensional relationship tensor R and the psychological state probability distribution vector, using a graph neural network to build a mapping network of object relationship and psychological state, extracting the associated features of object delivery mode and emotional expression; performing dynamic time warping on the decision time length features in the game process acoustic feature vector and the time sequence encoding vector T in the composite feature vector, generating behavior decision association features to reflect the internal relationship between the thinking process and the operation behavior; using a multi-head attention mechanism to weight and fuse the cross-modal time sequence association matrix, the space-semantic association vector, the associated features of the object delivery mode and the emotional expression, and the behavior decision association features, to obtain the associated features.

[0037] Among them, the cross-modal timing correlation matrix: refers to the correlation coefficient matrix quantifying the timing consistency between object operation behavior and language expression. The specific calculation method is: the cross-correlation analysis is performed on the composite feature vector (containing the category, placement timing and orientation information of each object) and the timing alignment vector (containing the relevance of each object operation corresponding language segment), to generate an n x n matrix M, wherein M[i, j] represents the timing correlation strength of the i-th object operation and the j-th language segment, and the value range is [-1, 1]. The high value on the diagonal of the matrix indicates consistency of speech and action, and the high value on the non-diagonal line may indicate anticipatory expression or retrospective explanation.

[0038] Space-semantic correlation vector: refers to the vector representation reflecting the correlation between the layout of objects in different regions of the sand table and the content of the psychological projection language. By calculating the cosine similarity between the 9 partition feature vectors (local point cloud features of each region) and the psychological projection feature vector (TF-IDF vector based on projection vocabulary), a 9-dimensional vector is obtained , wherein, represents the correlation strength of the i-th region and the psychological projection content. For example, if the central region similarity is high, it may indicate that there is more self-related projection content.

[0039] Object delivery mode: refers to the association pattern between object spatial relationship and emotional state learned by the graph neural network. Specifically, the object is taken as the graph node, the three-dimensional relationship tensor R is taken as the edge weight, and the psychological state probability distribution vector is taken as the node feature. The high-order interaction features extracted by multi-layer graph convolution operation. This mode can identify the corresponding relationship between specific object combination modes (such as enclosure, dispersion, and stacking) and specific emotional states (such as defense, openness, and depression), forming a "space configuration-psychological state" mapping mode.

[0040] Behavior decision correlation feature: refers to the feature vector obtained by correlating the decision thinking time and the actual operation sequence. By aligning the decision time series with the timing encoding vector T through the DTW algorithm, the following is calculated: (1) decision complexity index, reflecting the correlation between thinking time and operation complexity; (2) decision rhythm feature, indicating the alternating mode of fast and slow decisions; (3) decision consistency score, measuring whether the decision time of similar objects is similar. These features comprehensively reflect the internal relationship between the cognitive processing process and the behavior performance of employees.

[0041] Further, the execution core encodes, clusters and reduces dimensions of the associated features to obtain a multi-level psychological feature vector set, including: performing principal component analysis on the cross-modal time sequence association matrix in the associated features to extract the first K principal components as behavior consistency features to constitute a first-level psychological feature vector; inputting the space-semantic association vector and the object delivery mode and the associated features of emotional performance into a variational autoencoder to extract a converged mean vector μ as a projection mode feature to constitute a second-level psychological feature vector; applying a hierarchical clustering algorithm to the behavior decision association features, dividing N clustering clusters according to the decision mode, calculating the center vector of each clustering cluster as a decision style feature to constitute a third-level psychological feature vector; using a t-SNE algorithm to jointly reduce dimensions of the first-level psychological feature vector, the second-level psychological feature vector and the third-level psychological feature vector, and mapping the features to the same representation space; assigning a weight coefficient to each level through an adaptive weight; structurally organizing the weighted feature vectors of each level according to a psychological theory framework to form a multi-level psychological feature vector set including a behavior layer, a cognitive layer and a personality layer.

[0042] Among them, the variational autoencoder (Variational Auto encoder, VAE): refers to a deep learning network structure based on a probabilistic generative model, which extracts the latent features of psychological projection in the application. VAE is composed of an encoder and a decoder, the encoder maps the input space-semantic association vector and object relationship features to the probability distribution (mean μ and variance σ) of the latent space, and samples the latent variable z through the reparameterization technique, and the decoder reconstructs z back to the original feature space. When training, the reconstruction error and the KL divergence are optimized at the same time, so that the latent space has good continuity and interpretability, and can capture the essential mode of psychological projection.

[0043] The converged mean vector μ: refers to the mean vector parameter of the latent space output by the encoder network after the training of the variational autoencoder converges. In VAE, the encoder maps the input features to a Gaussian distribution N(μ, σ²), where μ is the mean vector of the distribution, usually with a dimension of 32 or 64. Training convergence means that the loss function (reconstruction error + βKL divergence) no longer decreases significantly, and the μ vector at this time stably encodes the core information of the input features, removes noise and redundancy, and becomes a compressed representation of the projection mode. The vector has semantic continuity, and similar psychological projection modes are close in the latent space.

[0044] Decision style feature: refers to the typical decision-making mode characteristics of employees in the sand table game identified by the hierarchical clustering algorithm. The specific process is: using Ward linkage criterion for hierarchical clustering on behavioral decision correlation features (including decision-making time, decision-making rhythm, decision-making consistency, etc.), determining the optimal clustering number N (usually 3-5 classes) through the silhouette coefficient, each class represents a decision style, such as “cautious type” (long and stable decision-making time), “impulsive type” (rapid and variable decision-making), “system type” (consistent decision-making on similar objects), etc. Calculate the center vector of each cluster as the numerical representation of the decision style, and finally obtain N decision style feature vectors.

[0045] Joint dimension reduction: refers to the process of using t-SNE algorithm to map feature vectors from different levels and different dimensions to the same low-dimensional representation space. In this application, the first level (behavior consistency feature), the second level (projection pattern feature), and the third level (decision style feature) three groups of feature vectors are first standardized and spliced, and then the t-SNE algorithm is applied. By optimizing the KL divergence to maintain the local neighborhood structure in the high-dimensional space, the features are mapped to a unified 50-dimensional space. This joint dimension reduction ensures that the features of different levels are comparable and can be integrated in the same coordinate system, while preserving the relative relationship and topological structure of the features of each level.

[0046] Another aspect of the present application also provides an employee psychological portrait system based on artificial intelligence, comprising: a multi-modal data acquisition module for acquiring multi-modal data including image data and audio data; wherein the image data includes three-dimensional point cloud data of the sand table game scene, object recognition results and spatial position relationship matrix; the audio data includes audio text transcription content, acoustic feature parameters and emotion annotation results; a hierarchical feature encoding module for performing hierarchical feature encoding on the multi-modal data to obtain a set of multi-level psychological feature vectors;

[0047] The hierarchical feature encoding module comprises: an open coding submodule that extracts features from the multi-modal data to obtain primary features, including using a three-dimensional point cloud processing algorithm to extract spatial features from image data of the sand table game scene to obtain an image feature vector, using natural language processing technology to extract semantic and acoustic features from audio data to obtain an audio feature vector, and splicing the two to obtain the primary features; a main shaft coding submodule that constructs the association between the primary features, generates associated features by calculating a cross-modal time sequence association matrix, a space-semantic association vector, an object delivery mode and emotional performance associated feature, and a behavior decision associated feature, and using a multi-head attention mechanism for weighted fusion; a core coding submodule that clusters and reduces the dimensions of the associated features, extracts behavior consistency features to form a first-level psychological feature vector through principal component analysis, extracts projection mode features to form a second-level psychological feature vector through a variational autoencoder, extracts decision style features to form a third-level psychological feature vector through hierarchical clustering, and uses a t-SNE algorithm for joint dimension reduction to form a multi-level psychological feature vector set comprising a behavior layer, a cognitive layer and a personality layer according to a psychological theory framework;

[0048] A time sequence prediction model construction module constructs a time sequence prediction model based on LSTM according to the multi-level psychological feature vector set, processes the behavior layer, cognitive layer and personality layer feature sequences through a three-layer stacked LSTM network, and integrates the features through an attention mechanism and a gating fusion unit;

[0049] A psychological portrait generation module generates an employee psychological portrait using the time sequence prediction model, including predicting future psychological state trends, dividing psychological types, integrating multi-time scale psychological features, and outputting a structured psychological portrait document containing psychological health state scores, dominant psychological types, key psychological feature labels and future trend predictions.

[0050] Compared with the prior art, the application has the following advantages:

[0051] In view of the problems in the prior art that the employee psychological assessment mainly relies on traditional methods such as questionnaire survey and interview, and the strong subjectivity, low efficiency, difficulty in capturing subconscious psychological characteristics, lack of comprehensive analysis of behavior and language, and inability to realize dynamic prediction of psychological state, the present application provides an employee psychological portrait method based on artificial intelligence, which collects three-dimensional point cloud image data and audio data in the sand table game process, uses hierarchical feature coding technology (including open coding, main axis coding and core coding) to deeply mine the multi-modal data, and constructs a time series prediction model based on LSTM, which can realize: automatically extracting multi-dimensional psychological projection characteristics such as sand table space layout, object placement order and language expression; revealing the consistency of behavior and language and the potential psychological mode through cross-modal correlation analysis; constructing a three-dimensional psychological portrait based on the multi-level features of the behavior layer, the cognitive layer and the personality layer; using the time series model to predict the future psychological state change trend of the employee, and providing objective, comprehensive and dynamic psychological health assessment and risk warning support for enterprise human resource management. BRIEF DESCRIPTION OF DRAWINGS

[0052] The present application will be further described in the form of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting, and in these embodiments, the same numbers represent the same structures, wherein:

[0053] Figure 1 is an exemplary flowchart of an employee psychological portrait method based on artificial intelligence according to some embodiments of the present application;

[0054] Figure 2 is an exemplary flowchart of constructing primary features according to some embodiments of the present application;

[0055] Figure 3 is an exemplary flowchart of constructing associated features according to some embodiments of the present application;

[0056] Figure 4 is an exemplary flowchart of constructing a multi-level psychological feature vector set according to some embodiments of the present application. DETAILED DESCRIPTION

[0057] The method and system provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0058] As shown in Figure 1 , multi-modal data including image data and audio data are collected; hierarchical feature coding is performed on the multi-modal data to obtain a multi-level psychological feature vector set; wherein the hierarchical feature coding includes open coding, main axis coding and core coding; a time series prediction model based on LSTM is constructed according to the multi-level psychological feature vector set; and an employee psychological portrait is generated by using the time series prediction model.

[0059] S1: Construct a multi-source heterogeneous data collection and preprocessing system:

[0060] S11: Collect three-dimensional scene data through visual sensors, use a convolutional neural network target detection algorithm to identify and locate objects in the scene in real time, extract object identifiers ID, three-dimensional coordinates (x, y, z), and timestamps t, and construct an object spatiotemporal dataset , and calculate the Euclidean distance matrix between any two objects ;

[0061] S12: Collect acoustic signals through audio sensors, use an end-to-end deep learning model to convert audio streams into text sequences, and extract audio feature vectors , including fundamental frequency, formant, and mel-frequency cepstral coefficients, and use a pre-trained classifier to output class labels;

[0062] S2: Perform hierarchical feature encoding on multi-source data:

[0063] As shown in Figure 2 S21: First layer encoding - use a depth camera to collect RGB-D data of the sand table game at 30fps, convert the depth map to a three-dimensional point cloud through the camera intrinsic matrix K , use the RANSAC algorithm to detect the sand table plane and extract the coordinates of the four corner points 、 , construct a 3x3 grid partition function: , divide the point cloud P into 9 subsets ;

[0064] For each regional point cloud , use the farthest point sampling to select 1024 points as input, extract features through the Set Abstraction layer of PointNet++: first aggregate 64 neighboring points within a radius of 0.1m using Ball Query, extract local features through MLP (3, 64, 128), and then aggregate through max pooling to obtain 256-dimensional regional features , concatenate the 9 regional features in the order of spatial position: ;

[0065] Use YOLOv5 to detect objects in the sand table, obtain the 3D bounding box and class label of each object, and record the detection timestamp Construct a time series encoding (i is the placement order and k is the total number of objects); calculate the main direction of the object point cloud through PCA and extract the orientation angle where v is the first principal component; one-hot encoding (100 dimensions) of the object class, concatenation with the time-series encoding (1 dimension) and the orientation encoding (mapped to 2 dimensions) to form a 103-dimensional single-object feature, k objects form a composite feature matrix and ; ;

[0066] Compute the horizontal Euclidean distance between objects i and j , normalized as ( , the diagonal length of the sand table), construct the distance matrix ; compare the z-coordinate of the object bottom with the local sand surface height , compute the burial rate ( , the height of the object), when mark as deep burial (value 1), mark as partial burial (value 0.5), otherwise 0, construct the burial matrix ; detect the object pairs that intersect and , compute the barycentric projection determine the stacking relationship, construct the stacking matrix , merge as ;

[0067] Use the Whisper model for speech recognition on the audio to obtain a timestamped text sequence, input the text into the BERT-base-chinese model, extract the 768-dimensional output of the [CLS] token as the global semantic representation, and extract the hidden layer output corresponding to the position of the object operation-related vocabulary (through keyword matching "place", "put", etc.) through attention pooling to aggregate into an operation semantic vector ;

[0068] Use jieba for word segmentation and part-of-speech tagging, identify the modifier-head structure through dependency syntax analysis, filter out the projective expressions containing personification (such as "the little man stands"), symbolism (such as "representing power"), and emotionalization (such as "feeling lonely"), calculate the TF-IDF: , select the top-50 TF-IDF values of the words as features ;

[0069] Frame the audio with a frame length of 25ms and a frame shift of 10ms, extract 13-dimensional MFCC, fundamental frequency F0 (through autocorrelation method), and short-time energy E to form a 16-dimensional basic acoustic feature; identify the object operation time (through video synchronization marking), calculate the time interval between adjacent operations As the decision duration; calculate the speech rate v(t) = number of words / duration within a 5-second sliding window, and calculate the standard deviation. As an indicator of mood changes; WebRTC VAD was used to detect silent segments and to calculate the percentage of total silent time. splicing as ;

[0070] Constructing audio frame sequences and operation sequence Calculate the DTW distance matrix The optimal path is found through dynamic programming, and the correlation coefficient is calculated along the path. The path coordinates and correlation coefficients are encoded into a 32-dimensional time-series aligned vector. ;

[0071] Will Input a 7-layer CNN-LSTM network (pre-trained on 100,000 hours of psychological counseling audio), and the output layer uses sigmoid activation to generate a 7-dimensional probability vector F_emotion=[p_anxiety, p_depression, p_anger, p_fear, p_sadness, p_happiness, p_calm];

[0072] Will (768 dimensions) (50 dimensions) (256 dimensions) (flattened to 32T dimension) and (7-dimensional) Concatenate along the feature dimension, input to a 3-layer fully connected network FC (1081+32T, 512, 256, D_audio), use ReLU activation and 0.5 dropout, output 3D audio feature vector ;

[0073] Ultimately , , R and Concatenate them in a predefined order to form an initial feature set. ,in, .

[0074] In a specific embodiment of this application, the first-level coding still follows the process of conceptualizing and categorizing the original information. During this coding stage, 14,225 lines of original sentences and corresponding initial concepts were obtained from 24 cases. After merging synonymous initial concepts, a total of 271 concepts and 17 categories for the first-level coding of the organizational commitment structure were extracted, as detailed in Table 1.

[0075] Table 1 First-level coding

[0076]

[0077] As shown in Figure 3 S22: second layer encoding - separate feature vectors from primary features through index slicing operation, assuming the primary feature dimension is D, the first D1 dimensions are extracted as image feature vectors and the last D2 dimensions are extracted as audio feature vectors according to a preset dimension mapping table; wherein the image feature vectors are further sliced to obtain partition feature vectors (9x128 dimensions), composite feature vectors (Mx64 dimensions, M is the number of objects) and three-dimensional relationship tensors , the audio feature vectors are divided into narrative semantic vectors (768 dimensions), psychological projection feature vectors (50 dimensions), game process acoustic feature vectors (256 dimensions), time alignment vectors (Tx32 dimensions, T is the number of time steps) and psychological state probability distribution vectors (7 dimensions);

[0078] reshape the composite feature vector into a two-dimensional matrix , and reshape the time alignment vector into a matrix , and calculate the cross-correlation matrix through matrix multiplication: , and obtain the cross-modal time sequence correlation matrix (MxT) by normalizing the result, wherein each element represents the correlation strength of the ith object and the jth time step;

[0079] flatten the partition feature vector into a one-dimensional vector S (1152 dimensions), and denote the psychological projection feature vector as P (50 dimensions), and calculate the cosine similarity through vector dot product: , and repeat the calculation on 9 partitions to obtain a 9-dimensional similarity vector, and then map it to a 256-dimensional space through a fully connected layer to generate a space-semantic correlation vector;

[0080] input the three-dimensional relationship tensor R as an adjacency matrix into the graph convolution network, and input the psychological state probability distribution vector as the node feature, and perform 3-layer graph convolution operation: , wherein A is the adjacency matrix, D is the degree matrix, and W is the learnable parameter, and finally output the correlation feature of the object delivery mode and emotional performance in Mx128 dimensions;

[0081] extract a decision duration feature sequence D (T1 dimensions) from the game process acoustic feature vector, and extract a time sequence encoding vector T (M dimensions) from the composite feature vector, and construct a cost matrix , and solve the optimal alignment path through dynamic programming: , and encode the alignment path into a 128-dimensional behavior decision correlation feature;

[0082] set 8 attention heads, each with a dimension of 64, perform linear transformation on the input 4 types of correlation features to obtain Q, K and V matrices, and calculate the attention weight: The outputs of the 8 heads are concatenated and passed through a feedforward network, and finally output 512-dimensional association features.

[0083] In the specific embodiments of the present application, the secondary coding is the second step of coding analysis, through the generic analysis, the various connections between categories are found and established to form more systematic generalization of the main categories, and 16 main categories of the secondary coding of the organizational commitment structure are induced. In order to show the original evidence of the categories (concepts) generated by the primary coding, and the relevance between the categories of the primary coding embodied in the main categories formed by the secondary coding, Table 2 lists the quotations of the three categories (concepts) of "self-value", "work reputation" and "sense of honor" in the primary coding and the "self-value" main category formed by the secondary coding to which they belong.

[0084] Table 2 Secondary coding

[0085]

[0086] As Figure 4 shown, S23: third-level coding - principal component analysis is performed on the cross-modal temporal association matrix in the association features, the eigenvalues and eigenvectors of the covariance matrix are calculated, the first K principal components whose cumulative variance contribution rate reaches 85% are selected after the eigenvalues are arranged in descending order, and the original matrix is projected into a K-dimensional subspace to obtain behavior consistency features, which constitute the first-level psychological feature vector;

[0087] The space-semantics association vector and the object delivery mode and emotion performance association features are concatenated and input into a variational autoencoder, the input is compressed into a mean vector μ and a standard deviation vector σ through an encoder network, the latent variable is sampled using the reparameterization trick (where ε ~ N(0, 1)) and the reconstruction error and KL divergence loss are calculated after the decoder is reconstructed, the converged mean vector μ is extracted as the projection mode feature, which constitutes the second-level psychological feature vector;

[0088] The Euclidean distance matrix between samples is calculated for the behavior decision association features, hierarchical clustering is performed using the Ward linkage criterion, the optimal clustering number N is determined by the silhouette coefficient, and the arithmetic mean of the feature vectors in each clustering cluster is calculated to obtain N clustering centers, which are concatenated as decision style features, constituting the third-level psychological feature vector;

[0089] The three-level psychological feature vectors are standardized and concatenated, the perplexity is set to 30 and the learning rate is set to 200 using the t-SNE algorithm, the KL divergence is minimized through gradient descent, and the high-dimensional features are mapped to a 50-dimensional unified representation space after 1000 iterations;

[0090] An auto-attention network is constructed to calculate the importance scores of each level, and adaptive weight coefficients are obtained through softmax normalization , meet ;

[0091] According to the three-layer structure theory of psychology, the first-level features are mapped to the behavior layer vector, the second-level features are mapped to the cognitive layer vector, and the third-level features are mapped to the personality layer vector, and a multi-level psychological feature vector set is constructed by weighted combination Table 3 is the three-level encoding of the embodiment.

[0092] Table 3 Three-level encoding

[0093]

[0094] S3, according to the multi-level psychological feature vector set, a time series prediction model based on LSTM is constructed; S4, an employee psychological portrait is generated by using the time series prediction model, including:

[0095] The multi-level psychological feature vector set is arranged in time sequence, wherein each time step contains feature vectors of three levels of behavior layer, cognitive layer and personality layer, and an input sequence , wherein, : [F_ behavior ^ t, F_ cognition ^ t, F_ personality ^ t];

[0096] A three-layer stacked LSTM network is constructed, the first layer LSTM contains 256 hidden units to process the behavior layer feature sequence, the second layer LSTM contains 128 hidden units to process the cognitive layer feature sequence, and the third layer LSTM contains 64 hidden units to process the personality layer feature sequence, and the information flow is maintained between layers through residual connection;

[0097] An attention mechanism is added after each LSTM layer to calculate the time step weight , wherein, is the LSTM hidden state, and a hierarchical representation vector is obtained by weighted summation;

[0098] The outputs of the three-layer LSTM are adaptively integrated through a gated fusion unit, the gated weight g = σ (W_g[h_ behavior; h_ cognition; h_ personality] + b_g), and the fused feature h_ fusion = g⊙h_ behavior+ (1-g) ⊙ (h_ cognition+h_ personality);

[0099] A time series prediction head is constructed based on the fused feature, a fully connected layer is used to predict the psychological state change of the future k time steps, and an output prediction sequence is output, each y contains a psychological health risk score, an emotional stability index and a stress bearing capacity evaluation;

[0100] Based on the historical feature sequence and the prediction result, the employee is divided into five psychological types by clustering algorithm: adaptive type, achievement type, social type, thinking type and innovative type, and the membership probability of each type is calculated;

[0101] Integrating psychological features of multiple time scales, including short-term emotional fluctuation patterns (based on the last 7 days of data), medium-term behavioral tendencies (based on the last 30 days of data), and long-term personality traits (based on all historical data), a three-dimensional psychological state matrix is constructed;

[0102] A structured employee psychological portrait is generated, including: current mental health status score and risk level, dominant psychological type and secondary type distribution, key psychological feature labels (such as stress resistance, team collaboration tendency, and innovation thinking activity level), future psychological state change trend prediction and early warning indicators, and the output is a JSON format psychological portrait document.

[0103] The above describes the present application and its embodiments in a schematic manner, which is not restrictive, and the present application can be realized in other specific forms without departing from the spirit or essential characteristics of the present application. The embodiments shown in the drawings are only one of the embodiments of the present application, and the actual structure is not limited thereto. Therefore, if a person skilled in the art is inspired by it, without departing from the spirit of the present application, similar structural forms and embodiments can be designed without creative design, which should belong to the protection scope of the present application. In addition, the word "comprising" does not exclude other elements or steps, and the word "one" before the element does not exclude the inclusion of "multiple" elements. The words "first", "second" and the like are used to indicate names, not any specific order.

Claims

1. An employee psychological profiling method based on artificial intelligence, characterized in that, include: Acquire multimodal data, including image and audio data; Hierarchical feature encoding is performed on multimodal data to obtain a multi-level set of psychological feature vectors; among which, hierarchical feature encoding includes open encoding, principal axis encoding and core encoding; Based on the multi-level psychological feature vector set, a time series prediction model based on LSTM is constructed; Employee psychological profiles are generated using time-series prediction models.

2. The employee psychological profiling method based on artificial intelligence according to claim 1, characterized in that: The image data includes 3D point cloud data of the sandbox game scene, object recognition results, and spatial position relationship matrix; The audio data includes the textual transcription of the audio, acoustic feature parameters, and sentiment annotation results; Acoustic characteristic parameters include pitch, speech rate, and pause duration.

3. The employee psychological profiling method based on artificial intelligence according to claim 2, characterized in that: Perform hierarchical feature encoding on multimodal data, including: Open coding is performed to extract features from multimodal data, resulting in primary features; Perform principal axis encoding to build the relationships between primary features and generate associated features; The core encoding is performed to cluster and reduce the dimensionality of the associated features, resulting in a multi-level set of psychological feature vectors.

4. The employee psychological profiling method based on artificial intelligence according to claim 3, characterized in that: Open coding is performed to extract features from multimodal data, resulting in primary features, including: Spatial feature extraction is performed on the image data of the sandbox game scene using a 3D point cloud processing algorithm to obtain image feature vectors; Natural language processing techniques are used to extract semantic and acoustic features from audio data to obtain audio feature vectors. The image feature vector and the audio feature vector are concatenated to obtain the primary features.

5. The employee psychological profiling method based on artificial intelligence according to claim 4, characterized in that: The image feature vector is obtained, including: A depth camera was used to collect 3D point cloud data of a sandbox game. The point cloud data was divided into 9 regions based on the boundary coordinates of the sandbox. The 9 regions include a central region, four corner regions, and four side regions. PointNet++ is used to extract local features of each region, and the features are concatenated to form a partition feature vector to obtain the differences in psychological projection in different regions; Identify game objects in the sandbox and record their placement sequence. Construct a temporal encoding vector T to represent the object placement order. Simultaneously extract the object orientation angle θ and generate a composite feature vector containing object category, temporal sequence, and orientation. A multi-level position matrix is ​​constructed based on the spatial relationships of the sand table. The first layer calculates the horizontal distance between objects and normalizes it according to the size of the sand table. The second layer identifies the degree to which objects are partially or completely covered by sand. The third layer detects the vertical contact between objects and generates a three-dimensional relationship tensor R. The image feature vector is obtained by concatenating the partition feature vector, the composite feature vector, and the three-dimensional relation tensor R.

6. The employee psychological profiling method based on artificial intelligence according to claim 4, characterized in that: The audio feature vector is obtained, including: The audio data during the sandbox game is subjected to speech recognition and converted into text data. The text data is encoded using the BERT model, and contextual features related to object operations are extracted through the attention mechanism to generate narrative semantic vectors. Part-of-speech tagging and dependency parsing are performed on the text data to identify projective words containing anthropomorphism, symbolism and emotionality. The TF-IDF weight of each word is calculated and the top % of words with the highest weights are selected to construct psychological projection feature vectors. The audio signal is processed in frames to extract the acoustic feature vectors of the game process; The audio timeline is aligned with the object operation timeline using a dynamic time warping algorithm. The temporal correlation coefficient between language expression and behavioral action is calculated, and a temporal alignment vector is generated. The acoustic feature vector of the game process is input into a multi-label classification model trained on a large-scale psychological counseling database, and the output is a 7-dimensional psychological state probability distribution vector. The feature concatenation method is used to concatenate the narrative semantic vector, psychological projection feature vector, game process acoustic feature vector, temporal alignment vector, and psychological state probability distribution vector, and then perform feature fusion through a fully connected layer to obtain the audio feature vector.

7. The employee psychological profiling method based on artificial intelligence according to claim 6, characterized in that: Extract the acoustic feature vectors of the game process, including: The basic acoustic parameters, including fundamental frequency, energy, and Mel frequency cepstral coefficients, are extracted using a preset time step as a window. The time interval between the placement actions of adjacent objects is calculated as a decision duration feature; The standard deviation of speech rate was calculated using the sliding window method to obtain characteristics of emotional changes. The introspection level feature is obtained by using a speech activity detection algorithm to mark silent segments and statistically analyzing their duration percentage. By combining basic acoustic parameters, decision duration features, emotion change features, and introspection level features, an acoustic feature vector of the game process is generated.

8. The employee psychological profiling method based on artificial intelligence according to claim 6, characterized in that: Generate associated features, including: Calculate the cross-correlation coefficient between the composite feature vector in the image feature vector and the temporal alignment vector in the audio feature vector, construct a cross-modal temporal correlation matrix, and capture the consistency between object operation behavior and language expression; By calculating the correlation between the partition feature vector in the image feature vector and the psychological projection feature vector in the audio feature vector using cosine similarity, a spatial-semantic association vector is generated to reflect the correspondence between the spatial layout of the sand table and the psychological projection content. Based on the three-dimensional relation tensor R and the probability distribution vector of psychological state, a graph neural network is used to construct a mapping network between object relations and psychological state, and to extract the correlation features between object delivery patterns and emotional expression. The decision duration feature in the acoustic feature vector of the game process and the temporal coding vector T in the composite feature vector are dynamically time-warped to generate behavioral decision association features, so as to reflect the intrinsic connection between the thinking process and the operational behavior. A multi-head attention mechanism is used to weight and fuse cross-modal temporal correlation matrix, spatial-semantic correlation vector, correlation features of object delivery mode and emotional expression, and correlation features of behavioral decision-making to obtain correlation features.

9. The employee psychological profiling method based on artificial intelligence according to claim 8, characterized in that: The resulting multi-level psychological feature vector set includes: Principal component analysis was performed on the cross-modal temporal correlation matrix in the correlation features, and the top K principal components were extracted as behavioral consistency features to form the first-level psychological feature vector. The spatial-semantic association vector and the association features between object delivery patterns and emotional expression are input into the variational autoencoder. The converged mean vector μ is extracted as the projection pattern feature to form the second-level psychological feature vector. Hierarchical clustering algorithm is applied to the behavioral decision-making association features, which are divided into N clusters according to the decision-making pattern. The center vector of each cluster is calculated as the decision style feature, forming the third-level psychological feature vector. The t-SNE algorithm is used to perform joint dimensionality reduction on the first-level, second-level, and third-level psychological feature vectors, mapping the features to the same representation space. Each level is assigned a weight coefficient through adaptive weighting; The weighted feature vectors at each level are structured and organized according to the framework of psychological theory to form a multi-level set of psychological feature vectors that includes behavioral, cognitive, and personality layers.

10. An employee psychological profiling system based on artificial intelligence, characterized in that, include: The data acquisition module collects multimodal data, including image data and audio data. The image data includes 3D point cloud data of the sandbox game scene, object recognition results, and spatial position relationship matrix; the audio data includes the text transcription content of the audio, acoustic feature parameters, and sentiment annotation results. The feature encoding module performs hierarchical feature encoding on multimodal data to obtain a multi-level set of psychological feature vectors; The temporal prediction module constructs an LSTM-based temporal prediction model based on a multi-level psychological feature vector set. It processes the feature sequences of the behavioral layer, cognitive layer, and personality layer through a three-layer stacked LSTM network, and integrates the features through an attention mechanism and a gating fusion unit. The psychological profiling module uses a time-series prediction model to generate psychological profiles of employees.