Multi-modal data fusion-based AI (Artificial Intelligence) generation virtual avatar recognition degree evaluation method and device

Through the multimodal data fusion method, a virtual clone recognition prediction model is constructed using eye movement, fMRI and behavioral data, which solves the problem of single dimensions and subjective interference in virtual clone evaluation, and achieves a more accurate recognition evaluation.

CN120387137AActive Publication Date: 2025-07-29SHANGHAI INTERNATIONAL STUDIES UNIVERSITY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510504653.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-29
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In the prior art, the virtual clone evaluation method has a single dimension and is easily disturbed by subjective cognitive interference, making it difficult to objectively reflect the real psychological reaction of consumers.

Method used

The multimodal data fusion method was used to obtain the subject's eye movement data, fMRI data and behavioral data, and feature-weighted fusion was performed through late fusion strategies and cross-modal attention mechanisms to construct a clone recognition prediction model and output evaluation results.

Benefits of technology

The prediction accuracy and objectivity of virtual clone recognition are improved, and consumers' sense of identity and affinity for virtual clones can be accurately evaluated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387137A_ABST
    Figure CN120387137A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data fusion-based method and device for evaluating the recognition degree of a virtual duplicate generated by AI, relates to the technical field of artificial intelligence, and solves the problem that the virtual duplicate is difficult to evaluate objectively due to single dimension and subjective cognitive interference when the virtual duplicate is evaluated in the prior art. The method comprises the following steps: acquiring eye movement data, fMRI data and behavior data of a subject, preprocessing the eye movement data, the fMRI data and the behavior data, extracting key features, realizing dynamic weighted fusion of the key features by adopting a late fusion strategy and a cross-modal attention mechanism to obtain multi-modal features, and constructing and applying a duplicate recognition degree prediction model based on the multi-modal features. According to the method, complementary information among all modes can be fully mined, the method has important theoretical significance and practical value for improving the prediction precision and objectivity of the virtual avatar recognition degree, the virtual avatar recognition degree can be objectively and accurately evaluated, and the prediction effect is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and device for evaluating the recognition of an AI-generated virtual avatar based on multimodal data fusion. Background Art

[0002] With the rapid development of artificial intelligence, virtual reality, and neuroimaging technologies, AI-generated avatars are playing an increasingly important role in social media, online entertainment, and commercial promotion. Avatar identification refers to consumers' identification with and affinity for AI-generated avatars in terms of appearance, behavior, and personality. This metric not only reflects consumers' emotional investment in the avatar but also their subjective assessment of the degree to which the avatar fits their true identity. Existing avatar evaluation methods primarily rely on questionnaires or single data sources. These methods not only have a single data dimension but are also highly susceptible to subjective cognition, making it difficult to objectively reflect consumers' true psychological reactions. Summary of the Invention

[0003] The purpose of this application is to overcome the problem of difficulty in objectively evaluating virtual avatars due to the single dimension and subjective cognitive interference in the existing technology when evaluating virtual avatars, and to provide an AI-generated virtual avatar identity evaluation method and device based on multimodal data fusion.

[0004] First, a method for evaluating the recognition of AI-generated virtual avatars based on multimodal data fusion is provided, including:

[0005] Acquiring subject data, wherein the subject data includes eye movement data, fMRI data, and behavioral data of a subject viewing a static image of a virtual avatar of the subject generated by the AI software;

[0006] Preprocessing the subject data;

[0007] Extract key features from preprocessed subject data;

[0008] Adopting late fusion strategy and cross-modal attention mechanism to achieve dynamic weighted fusion of the key features to obtain multimodal features;

[0009] Constructing an avatar identification prediction model based on the multimodal features;

[0010] The avatar identification prediction model is used to evaluate the identification of the AI-generated virtual avatar to output an evaluation result.

[0011] In some possible implementations, the evaluation results are output through an interactive interface, where the evaluation results include the predicted score of the virtual avatar's identity recognition and its confidence interval, the distribution map of the contribution weights of each modality during the fusion process, and the dynamic trends of the feature contributions of each modality and their changes over time presented using radar charts, heat maps, and line charts. The interactive interface is used to implement data export and interactive data screening.

[0012] In some possible implementations, the fMRI data is the BOLD signal collected by a functional magnetic resonance imaging device, and the behavioral data includes subjective scores and response durations.

[0013] In some possible implementations, preprocessing the eye movement data includes: data format conversion, removing noise during blinking using methods such as low-pass filtering, detecting stable fixation points, and filtering out abnormal data; preprocessing the fMRI data includes: performing head motion correction using rigid transformation, performing temporal slice correction, spatial normalization, and processing using Gaussian smoothing and frequency domain band-pass filtering; preprocessing the behavioral data includes: cleaning, handling missing values and outliers, performing data normalization, and designing specialized preprocessing processes for the different characteristics of the eye movement data, fMRI data, and behavioral data to ensure data quality.

[0014] In some possible implementations, the key features of the preprocessed eye movement data include fixation duration, saccade speed, and fixation heat map distribution; the key features of the preprocessed fMRI data include the neural score, and the calculation formula for the neural score is: Score = (∑_ voxel (BOLD_ voxel × Weight_ voxel )) / (∑_ voxel Weight_ voxel )), where BOLD_voxel is the blood oxygenation level-dependent (BOLD) signal intensity value of each voxel, and Weight_ voxel represents the importance weight corresponding to this voxel in the identity recognition prediction task, and the statistical parameter method is used to determine Weight_ voxel ; the key features of the preprocessed behavioral data include the quantitative feature vector formed by statistically analyzing the subjective scores of the subjects.

[0015] In some possible implementations, a late fusion strategy and a cross-modal attention mechanism are used to achieve dynamic weighted fusion of the key features, including:

[0016] Encoding the key features to obtain encoded vectors of a fixed dimension;

[0017] Generating a query vector Q through a fully connected layer and generating a key vector K for each modalityi Sum vector V i ;

[0018] Calculate the matching score S for each modality i : S i =f(Q, K i ), where f(·) includes scaled dot product;

[0019] Normalize the matching scores using the Softmax function to obtain weights:

[0020] ;

[0021] Among them, represents the attention weight of each modality, i represents the modality for which the attention weight is currently being calculated, and j represents the index of all modalities; S i is the matching score, exp(S i ) represents the exponential form of the matching score S i , represents the sum of the exponential forms of the matching scores for all modalities j;

[0022] Weightedly sum the encoded vectors of each modality according to the corresponding attention weights to obtain a unified multi-modal feature representation F fusion :

[0023] F fusion =∑ i α i ·V i

[0024] Among them, V i is the value vector from the i-th modality.

[0025] In some possible implementation manners, a doppelganger identity prediction model is constructed based on the multi-modal features, including:

[0026] Form a data set with the multi-modal features and the doppelganger identity labels;

[0027] Divide the data set into a training set and a test set according to a preset ratio;

[0028] Construct a regression prediction model based on a multi-layer perceptron or a convolutional neural network;

[0029] Train the regression prediction model using the training set and the test set;

[0030] Perform cross-validation and performance evaluation on the regression prediction model, and control the false positive rate through multiple comparison correction to obtain a doppelganger identity prediction model.

[0031] In a second aspect, an AI-generated virtual avatar recognition degree evaluation device based on multimodal data fusion is provided, including:

[0032] A data acquisition module for acquiring subject data, where the subject data includes eye movement data, fMRI data, and behavioral data of a subject who views a static image of a virtual avatar generated by an AI software;

[0033] A preprocessing module for preprocessing the subject data;

[0034] A feature extraction module for extracting key features from the preprocessed subject data;

[0035] A weighted fusion module for dynamically weighting and fusing the key features by adopting a late fusion strategy and a cross-modal attention mechanism to obtain multimodal features;

[0036] A model construction module for constructing an avatar recognition degree prediction model based on the multimodal features;

[0037] A prediction module for using the avatar recognition degree prediction model to evaluate the recognition degree of an AI-generated virtual avatar and output an evaluation result.

[0038] In a third aspect, a computer-readable storage medium is provided, and the computer-readable medium stores program code for a device to execute, and the program code includes steps for executing the method in any one of the implementation manners in the first aspect as described above.

[0039] In a fourth aspect, an electronic device is provided, and the electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, the method in any one of the implementation manners in the first aspect as described above is implemented.

[0040] The present application has the following beneficial effects: By adopting a late fusion and a cross-modal attention mechanism to fuse the key features in eye movement data, fMRI data, and behavioral data, it is possible to fully exploit the complementary information between modalities, which has important theoretical significance and practical value for improving the prediction accuracy and objectivity of virtual avatar recognition degree. Moreover, the avatar recognition degree prediction model constructed based on the fused features can objectively and accurately evaluate the virtual avatar recognition degree, significantly improving the prediction effect. Description of the Drawings

[0041] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application.

[0042] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0043] Figure 1 It is a flowchart of the AI-generated virtual avatar recognition degree evaluation method based on multimodal data fusion in Embodiment 1 of the present application;

[0044] Figure 2 It is a schematic diagram of the avatar recognition degree prediction model in the AI-generated virtual avatar recognition degree evaluation method based on multimodal data fusion in Embodiment 1 of the present application;

[0045] Figure 3 It is a structural block diagram of the AI-generated virtual avatar recognition degree evaluation device based on multimodal data fusion in Embodiment 2 of the present application;

[0046] Figure 4 It is an internal structural schematic diagram of the electronic device in Embodiment 4 of the present application.

[0047] Reference numerals:

[0048] 100, data acquisition module; 200, preprocessing module; 300, feature extraction module; 400, weighted fusion module; 500, model construction module; 600, prediction module. Specific embodiments

[0049] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0050] Embodiment 1

[0051] As Figure 1 shown, a method for evaluating the recognition degree of an AI-generated virtual avatar based on multimodal data fusion involved in Embodiment 1 of the present application includes:

[0052] S100. Obtain subject data, where the subject data includes eye movement data, fMRI data, and behavior data of subjects who watch static images of the test virtual avatar generated by AI software;

[0053] First, a virtual avatar for the experiment needs to be prepared. For example, ask the subject to take a photo on the day of the experiment, and use this photo as the original input for generating the virtual avatar. Then, process the taken photo with the pre-prepared AI software to generate the corresponding static image of the virtual avatar.

[0054] Then, data collection is carried out. Specifically: when the subject views the generated virtual avatar image, two types of data, namely behavioral data and fMRI data, are collected simultaneously. The collection of behavioral data includes: recording the subjective scores (such as sense of identity, degree of preference), reaction duration, and other relevant behavioral indicators of the subject during the viewing process; the collection of fMRI data includes: using a functional magnetic resonance imaging device to collect the BOLD signals in the subject's brain during the viewing process, reflecting the neural activation situation.

[0055] Finally, design an independent experimental session specifically for collecting eye movement data, and use a high-precision eye tracker to collect the eye movement data of the subject when viewing the virtual avatar image. Among them, the eye movement data includes fixation points, saccade trajectories, and pupil diameters, etc.

[0056] S200. Preprocess the data of the subject;

[0057] To ensure the data quality, specific preprocessing processes are designed for the different characteristics of eye movement data, fMRI data, and behavioral data. Specifically:

[0058] The preprocessing of eye movement data includes:

[0059] Data format conversion: Convert the original eye movement record (CSV or proprietary format) to a standard format for subsequent processing.

[0060] Noise filtering: Use low-pass filtering or other appropriate filtering algorithms to remove the interference of high-frequency noise during blinking and other periods to ensure data smoothness.

[0061] Abnormal data rejection: Use statistical methods (outlier detection) to reject significantly abnormal data points to improve data accuracy.

[0062] Feature extraction: Extract key indicators from the processed eye movement data, including fixation point position, fixation duration, saccade trajectories, and pupil diameter changes, etc., as the input features for subsequent cross-modal fusion.

[0063] The preprocessing of fMRI data includes:

[0064] Data format conversion: Convert the original fMRI data (DICOM format) to a standard format (NIfTI) for unified processing.

[0065] Head motion correction: Use rigid-body transformation to correct the subject's head motion and align images at different time points.

[0066] Slice timing correction: Perform time layer correction on each slice data to eliminate the influence of time differences during acquisition.

[0067] Spatial normalization: Each subject's image is mapped to a standard template (MNI space) through nonlinear transformation to enable cross-individual data comparison.

[0068] Spatial smoothing: Gaussian smoothing is used to process the data to reduce local noise and improve the signal-to-noise ratio.

[0069] Frequency domain bandpass filtering: Use a bandpass filter (0.01-0.1 Hz) to filter out low-frequency drift and high-frequency noise, retaining signals related to cognitive activities.

[0070] The preprocessing of behavioral data includes:

[0071] Data cleaning: Check the original behavioral data and remove erroneous records and duplicate data.

[0072] Missing value and outlier processing: Use interpolation or statistical methods to fill missing values, and identify and eliminate or correct abnormal data to ensure data integrity and accuracy.

[0073] Standardization processing: Standardize data of each dimension (such as Z-score standardization) to make data of different dimensions comparable, which facilitates subsequent multimodal data fusion and model training.

[0074] S300, extracting key features from the pre-processed subject data, specifically includes the following steps:

[0075] Extract fixation duration, saccade speed and fixation heat map distribution from the preprocessed eye movement data;

[0076] The fMRI activation images were mapped to the predefined brain area template related to virtual double identification, and the neural score of each subject was calculated using the formula: Score=(∑_voxel(BOLD_voxel×Weight_voxel)) / (∑_voxelWeight_voxel), where BOLD_voxel is the blood oxygen level dependent (BOLD) signal intensity value of each voxel, and Weight_ voxel Indicates the importance weight of the voxel in the identification prediction task, and the statistical parameter method is used to determine Weight_ voxelSpecifically, the preprocessed fMRI data is analyzed using the General Linear Model (GLM) to calculate the activation level of each voxel in tasks related to virtual avatar identification, which is usually represented by the β value. The β value reflects the response intensity of each voxel to the experimental conditions and can thus be used as the weight of the voxel in the neural score calculation. To ensure the comparability of voxel weights, these β values are normalized by normalizing the sum of the β values of all voxels to 1, so that the contribution of each voxel is reasonably reflected in the overall score. Such a method is objective and has a wide application basis, and can better reflect the true contribution of each voxel in task-related neural activities.

[0077] Statistical analysis is performed on the behavioral data to extract key indicators such as subjective scores and response durations to form a behavioral feature vector.

[0078] S400. Implement dynamic weighted fusion of the key features using a late fusion strategy and a cross-modal attention mechanism to obtain multi-modal features, specifically including:

[0079] S401. Feature encoding: Use a convolutional neural network (CNN) or a fully connected layer to encode the key features extracted from the eye movement data, fMRI data, and behavioral data respectively to obtain encoded vectors F_eye, F_fMRI, and F_behav with a fixed dimension.

[0080] Construct a cross-modal attention module:

[0081] S402. Use a fully connected layer to generate a query vector Q.

[0082] S403. Generate corresponding key vectors K i and value vectors V i (Generate the key vector K i and value vector V i ) through linear transformation; these value vectors V i represent the features to be weighted and summed in the subsequent fusion for each modality.

[0083] S404. Use a unified query vector Q (which can be generated from the encoded vectors of all modalities) and the key vector K of each modality i to calculate the matching score S i : S i =f(Q, K i ), for example, f(·) is the scaled dot product, that is , where d is the dimension of the key vector and K i is the key vector corresponding to each modality. Calculate the matching score or through MLP non-linear transformation.

[0084] S405. After these matching scores are normalized by Softmax, the attention weights for each modality are obtained , that is, the contribution ratio of the i-th modality to the final fusion result. This attention weight represents the importance of the modality features in the current task. Normalize the scores using the Softmax function: .

[0085] S406. Finally, multiply the value vector V i of each modality by the corresponding weight , and sum over all modalities to obtain a unified fused feature representation.

[0086] That is to say, the weight determines the contribution ratio of eye movement, fMRI, and behavioral data in the final fused representation, thus dynamically affecting the recognition score of the AI-generated virtual avatar. Weight the encoded vectors of each modality by the corresponding weights and sum them to form a unified fused feature F fusion :

[0087] F fusion =∑ I α i ·V i

[0088] i∈{eye, fMRI, behav}

[0089] where V i is the value vector from the i-th modality, eye represents the key features extracted from eye movement data, fMRI represents the key features extracted from fMRI data, and behav represents the key features extracted from behavioral data.

[0090] In this embodiment, adopting the late fusion and cross-modal attention mechanism can fully exploit the complementary information between modalities; it has important theoretical significance and practical value for improving the prediction accuracy and objectivity of the virtual avatar recognition degree.

[0091] S500. Construct a recognition degree prediction model for the avatar based on the multi-modal features;

[0092] As Figure 2 shown, first, the fused feature F fusionThe combination of virtual avatar identity score tags corresponding to each subject constitutes a complete dataset. This dataset is then randomly divided into a training set and a test set (e.g., 80% for training and 20% for testing) to ensure the generalization ability of the model on unknown data. Next, a regression model based on a multi-layer perceptron (MLP) is constructed to map the fused features to the predicted scores of virtual avatar identity. The model structure includes an input layer, multiple hidden layers, and an output layer. The input layer is responsible for receiving the preprocessed and normalized F fusion features; the hidden layers are composed of multiple fully connected layers, each layer using the ReLU activation function, and introducing Dropout when necessary to prevent overfitting, thus gradually capturing the non-linear relationship between the fused features and the target tags; the output layer consists of a single neuron, using a linear activation function to output the final predicted score.

[0093] During the model training process, the mean squared error (MSE) is used as the loss function, and the model parameters are iteratively updated through an optimizer (such as Adam) to minimize the error between the predicted score and the actual label. At the same time, a cross-validation method (such as K-fold cross-validation) is used to evaluate the stability and generalization performance of the model, and the root mean squared error (RMSE) and Pearson correlation coefficient are used as evaluation metrics. In addition, to control the statistical error introduced by multiple comparisons, the FDR correction method is also used to adjust the results.

[0094] In this embodiment, the prediction model constructed based on the fused features can objectively and accurately evaluate the virtual avatar identity, significantly improving the prediction effect. The algorithm processes and key formulas of each module are described in detail, enhancing the scientific nature and protection strength of the technical solution. The avatar identity prediction model of this application is applicable to multiple fields such as virtual image design, personalized generation, and marketing, and has broad commercial application value.

[0095] S600. Use the avatar identity prediction model to evaluate the identity of the AI-generated virtual avatar to output an evaluation result.

[0096] To output more and more intuitive data and to achieve data export and interactive screening, the evaluation result is output through an interactive interface. Among them, the evaluation result includes the predicted score of virtual avatar identity and its confidence interval, the distribution map of the contribution weights of each modality during the fusion process, and the dynamic trends of the contribution of each modality feature and its change over time shown by radar charts, heatmaps, and line charts, intuitively showing the contribution of each modality feature and its dynamic change over time. The interactive interface supports users to perform data export, screening, and in-depth analysis.

[0097] Embodiment 2

[0098] AsFigure 3 As shown in Figure 3 , an AI-generated virtual avatar recognition degree evaluation device based on multimodal data fusion involved in Embodiment 2 of the present application includes:

[0099] A data acquisition module 100, configured to acquire subject data, where the subject data includes eye movement data, fMRI data, and behavioral data of a subject who watches a static image of a test virtual avatar generated by an AI software;

[0100] A preprocessing module 200, configured to preprocess the subject data;

[0101] A feature extraction module 300, configured to extract key features from the preprocessed subject data;

[0102] A weighted fusion module 400, configured to implement dynamic weighted fusion of the key features by using a late fusion strategy and a cross-modal attention mechanism to obtain multimodal features;

[0103] A model construction module 500, configured to construct an avatar recognition degree prediction model based on the multimodal features;

[0104] A prediction module 600, configured to use the avatar recognition degree prediction model to evaluate the recognition degree of an AI-generated virtual avatar and output an evaluation result.

[0105] It should be noted that for other specific implementation manners of the AI-generated virtual avatar recognition degree evaluation device based on multimodal data fusion in this embodiment, reference may be made to the specific implementation manners of the AI-generated virtual avatar recognition degree evaluation method based on multimodal data fusion above. To avoid redundancy, it will not be elaborated here.

[0106] Embodiment 3

[0107] A computer-readable storage medium involved in Embodiment 3 of the present application, where the computer-readable medium stores program code for a device to execute, and the program code includes steps for executing the method in any one of the implementation manners in Embodiment 1 of the present application;

[0108] Among them, the computer-readable storage medium may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM); the computer-readable storage medium can store program code, and when the program stored in the computer-readable storage medium is executed by a processor, the processor is configured to execute the steps of the method in any one of the implementation manners in Embodiment 1 of the present application.

[0109] Embodiment 4

[0110] As Figure 4As shown, an electronic device involved in Embodiment 4 of the present application. The electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements the method in any one of the implementation manners in Embodiment 1 of the present application;

[0111] Among them, the processor may adopt a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits, and is used to execute relevant programs to implement the method in any one of the implementation manners in Embodiment 1 of the present application.

[0112] The processor may also be an integrated circuit electronic device with the ability to process signals. In the implementation process, each step of the method in any one of the implementation manners in Embodiment 1 of the present application may be completed by the integrated logic circuit in the hardware of the processor or the instruction in software form.

[0113] The above-mentioned processor may also be a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application may be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the functions required to be executed by the units included in the data processing device of the embodiments of the present application, or execute the method in any one of the implementation manners in Embodiment 1 of the present application.

[0114] The above is only a preferred specific implementation manner of the present application; however, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application, according to the technical solution of the present application and its improved concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present application.

Claims

1. An AI-generated virtual avatar recognition degree evaluation method based on multi-modal data fusion, characterized in that Including: Obtaining subject data, where the subject data includes eye movement data, fMRI data, and behavioral data of subjects who view static images of virtual avatars generated by AI software; Preprocessing the subject data; Extracting key features from the preprocessed subject data; Implementing dynamic weighted fusion of the key features using a late fusion strategy and cross-modal attention mechanism to obtain multi-modal features; Constructing an avatar identity prediction model based on the multi-modal features; Evaluating the identity of the AI-generated virtual avatar using the avatar identity prediction model to output an evaluation result.

2. The AI-generated virtual avatar recognition degree evaluation method based on multimodal data fusion according to claim 1, wherein, Outputting the evaluation result through an interactive interface, where the evaluation result includes the predicted score of the virtual avatar identity and its confidence interval, the distribution map of the contribution weights of each modality during the fusion process, and the dynamic trends of the contribution of each modality feature and its change over time presented using radar charts, heat maps, and line charts. The interactive interface is used to achieve data export and interactive data screening.

3. The AI-generated virtual avatar recognition degree evaluation method based on multi-modal data fusion according to claim 1, characterized in that The fMRI data is the BOLD signal collected by a functional magnetic resonance imaging device, and the behavioral data includes subjective scores and response durations.

4. The AI-generated virtual avatar recognition degree evaluation method based on multimodal data fusion according to claim 1, characterized in that Preprocessing the eye movement data includes: data format conversion, removing noise during blinking using methods such as low-pass filtering, detecting stable fixation points, and filtering out abnormal data; preprocessing the fMRI data includes: performing head motion correction using rigid transformation, performing temporal slice correction, spatial normalization, and using Gaussian smoothing and frequency domain band-pass filtering; preprocessing the behavioral data includes: cleaning, handling missing values and outliers, and performing data normalization.

5. The AI-generated virtual avatar identity evaluation method based on multimodal data fusion according to claim 1, characterized in that, The key features of the preprocessed eye movement data include fixation duration, saccade speed, and fixation heat map distribution; the key features of the preprocessed fMRI data include a neural score, and the calculation formula of the neural score is: Score=(∑_ voxel (BOLD_ voxel ×Weight_ voxel )) / (∑_ voxel Weight_ voxel ), where BOLD_ voxel is the blood oxygenation level-dependent signal intensity value of each voxel, and Weight_ voxel represents the importance weight corresponding to the voxel in the agreement prediction task; the key features of the preprocessed behavioral data include a quantitative feature vector formed by statistically analyzing the subjective scores of the subjects.

6. The AI-generated virtual avatar identity evaluation method based on multimodal data fusion according to claim 1, wherein Implementing dynamic weighted fusion of the key features using a late fusion strategy and cross-modal attention mechanism includes: Encoding the key features to obtain an encoded vector of a fixed dimension; Generate a query vector Q through a fully connected layer and generate a key vector K for each modality i and a value vector V i ; Calculate the matching score S for each modality i : S i = f(Q, K i ), where f(·) includes scaled dot product; Normalizing the matching scores using the Softmax function to obtain weights: ; Among them, represents the attention weight of each modality, i represents the modality for which the attention weight is currently being calculated, and j represents the index of all modalities; S i is the matching score for each modality, exp(S i ) represents the exponential form of the matching score S i ; represents the sum of the exponential forms of the matching scores for all modalities j; Weighted sum the encoded vectors of each modality according to the corresponding attention weights to obtain a unified multi-modal feature representation F fusion : F fusion =∑ i α i ·V i Among them, V i is the value vector from the i-th modality.

7. The AI-generated virtual avatar recognition degree evaluation method based on multi-modal data fusion according to claim 1, characterized in that, Constructing an avatar identity prediction model based on the multi-modal features includes: Forming a dataset with the multi-modal features and virtual avatar identity labels; Dividing the dataset into a training set and a test set according to a preset ratio; Constructing a regression prediction model based on a multi-layer perceptron or convolutional neural network; Training the regression prediction model using the training set and the test set; Performing cross-validation and performance evaluation on the regression prediction model, and controlling the false positive rate through multiple comparison correction to obtain an avatar identity prediction model.

8. An AI-generated virtual avatar recognition degree evaluation device based on multi-modal data fusion, characterized in that, Including: A data acquisition module for obtaining subject data, where the subject data includes eye movement data, fMRI data, and behavioral data of subjects who view static images of virtual avatars generated by AI software; A preprocessing module for preprocessing the subject data; A feature extraction module for extracting key features from the preprocessed subject data; A weighted fusion module for implementing dynamic weighted fusion of the key features using a late fusion strategy and cross-modal attention mechanism to obtain multi-modal features; A model construction module for constructing an avatar identity prediction model based on the multi-modal features; A prediction module, configured to use the doppelganger identity prediction model to evaluate the identity of the AI-generated virtual doppelganger and output an evaluation result.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code for execution by a device, the program code including steps for performing the method according to any one of claims 1-7.

10. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor, and when the program or instruction is executed by the processor, the method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Virtual image evaluation method, device and equipment and computer readable storage medium

    CN112116589A

  • Virtual reality video emotion recognition method and system based on time sequence characteristics

    CN114581823A

  • Multi-modal personality traits analysis method based on progressive adaptive modal enhanced attention network

    CN117520811A

  • Emotion-driven 2D supernatural digital human video generation system of multi-modal model

    CN119672601A

  • Multi-modal modeling of temporal interaction sequences

    US20140212853A1