Human-computer interaction method of naked eye 3D digital human based on AI large model
By using AI large-scale models to process multimodal input information in the human-computer interaction system of naked-eye 3D digital humans, the problems of inaccurate semantic understanding, lack of personalization and insufficient real-timeness in the prior art are solved, and a more accurate, personalized and smooth interactive experience is achieved.
Patent Information
- Application Number
- CN202510579163.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing human-computer interaction technology based on naked-eye 3D digital humans has problems such as inaccurate semantic understanding, lack of personalization of interactions and insufficient real-timeness.
A multimodal interaction method based on AI big model is adopted to optimize the interaction process through loading AI big model resources, collecting and fusion of multimodal input information, understanding user intentions, generating personalized interaction strategies and response content, and the interaction process is optimized through big data analysis.
It achieves more accurate user intention understanding, personalized interactive services and smooth response, meets the diverse needs of different users, improves user experience, and continuously improves interaction quality.
Smart Images

Figure CN120104009A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and in particular to a human-computer interaction method of a naked-eye 3D digital human based on an AI large model. Background Art
[0002] With the continuous development of science and technology, human-computer interaction technology has made remarkable progress; digital humans, as an emerging human-computer interaction carrier, can interact with users in an anthropomorphic image and have received widespread attention and application in multiple fields; naked-eye 3D technology does not require the aid of special glasses or other equipment, and users can directly view the digital human image with a three-dimensional effect, enhancing the immersion and attractiveness of the interaction.
[0003] Existing digital human interaction methods are mostly based on rule-based natural language processing interaction, which parses the natural language input by users according to pre-set rules and grammar; uses universal interaction templates to interact with all users, such as responding with the same words and processes regardless of user interests and preferences; and uses traditional computing and rendering technologies to achieve digital human interaction when generating naked-eye 3D digital human's response actions, expressions and voice replies.
[0004] However, there are still many problems in the current human-computer interaction based on naked-eye 3D digital humans: it is difficult to accurately understand the user's true intentions based on pre-set rules and grammatical analysis, which can easily lead to misunderstandings or inability to give accurate responses, resulting in inaccurate semantic understanding; using general interaction templates, it is impossible to provide customized interaction services based on the user's personalized characteristics, it is difficult to meet the diverse needs of users, and the interaction lacks personalization; when generating the response actions, expressions, and voice replies of naked-eye 3D digital humans, due to the complex calculation and rendering process involved, some systems have response delays, resulting in unsmooth interactions, affecting user experience, and insufficient real-time performance. Summary of the invention
[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a human-computer interaction method of a naked-eye 3D digital human based on an AI large model to solve the problems of inaccurate semantic understanding, lack of personalization in interaction, and insufficient real-time performance raised in the above-mentioned background technology.
[0006] To achieve the above object, the present invention provides the following technical solution: a human-computer interaction method of naked-eye 3D digital human based on AI large model, comprising: S1: AI big model resource loading: Load the resources required for naked-eye 3D digital human operation into the trained AI big model, and build an AI big model for naked-eye 3D digital human operation. The resources required for naked-eye 3D digital human operation include naked-eye 3D digital human resources, multimodal input device resources, and user historical interaction records; S2: Multimodal input information fusion: Based on the multimodal input device resources obtained in S1, the multimodal information of the user's interactive request to the naked-eye 3D digital human is collected, and then fused through the AI big model to obtain the user input representation text; S3: AI large model intention understanding: Based on the user input representation text obtained in S2, the intention is understood to obtain the user intention understanding vector; S4: Personalized strategy generation: Based on the AI big model operated by naked-eye 3D digital human, the personalized interaction strategy information text is generated through the user intention understanding vector obtained in S3; S5: Naked-eye 3D digital human response generation: Based on the AI big model of naked-eye 3D digital human operation, the corresponding naked-eye 3D digital human response content text is generated according to the personalized interaction strategy information text generated in S4, and the naked-eye 3D digital human response content text is generated through the naked-eye 3D arrangement algorithm to generate naked-eye 3D animation and display it on the screen; S6: Analyze the feedback parameters of the user after receiving the response content text through big data analysis technology, optimize according to the abnormal analysis results, and record the optimized content and transmit it to the administrator terminal.
[0007] Technical effects and advantages of the present invention: 1. The present invention uses various sensor technologies and recognition technologies to collect, identify and integrate multimodal interaction information through interaction requests initiated by users standing in front of naked-eye 3D digital human devices, forming a unified user input representation, providing more comprehensive and rich information for subsequent intention understanding, and thus more accurately determining the user's intended object; 2. The present invention realizes the fusion interaction of multimodal information by constructing an AI large model for the operation of naked-eye 3D digital humans, which can deeply understand the intentions and needs of users, and provide personalized replies and services based on the historical interaction records and real-time context of users to meet the diverse needs of different users, making the communication between naked-eye 3D digital humans and users more natural and smooth, just like the interaction between real people; 3. The present invention can continuously improve the quality of human-computer interaction by receiving user feedback and optimizing the AI large model; at the same time, it updates the personalized interaction strategy according to the user's new needs and preferences to provide better services for the next interaction, thereby continuously improving the intelligence level and interaction ability of the naked-eye 3D digital human over time. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 It is a schematic diagram of the overall process of the present invention.
[0009] Figure 2 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION
[0010] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0011] See also Figure 1 As shown, the present invention provides a naked-eye 3D digital human human-computer interaction system based on an AI large model, including an AI large model resource loading module, a multimodal input fusion module, an AI large model intention understanding module, a personalized interaction strategy generation module, a naked-eye 3D digital human response generation module, and a naked-eye 3D digital human human-computer interaction feedback and optimization module.
[0012] The AI large model resource loading module is connected to all other modules, the multimodal input fusion module is connected to the AI large model intention understanding module, the personalized interaction strategy generation module is connected to the AI large model intention understanding module and the naked-eye 3D digital human response generation module respectively, and the naked-eye 3D digital human human-computer interaction feedback and optimization module is connected to the naked-eye 3D digital human response generation module.
[0013] AI big model resource loading module: used to load the resources required for naked-eye 3D digital human operation into the trained AI big model, and build the AI big model for naked-eye 3D digital human operation. The resources required for naked-eye 3D digital human operation include naked-eye 3D digital human resources, multimodal input device resources and user historical interaction records; Multimodal input fusion module: Based on the multimodal input device resources obtained by the AI large model resource loading module, it collects and fuses the multimodal information of the user's interaction request to the naked-eye 3D digital human, and transmits the fused user request information to the AI large model intention understanding module; AI big model intention understanding module: Based on the AI big model operated by naked-eye 3D digital human, it understands the intention of the fused user request information, obtains the user intention understanding vector, and transmits it to the personalized interaction strategy generation module; Personalized interaction strategy generation module: Based on the AI big model of naked-eye 3D digital human operation, the personalized interaction strategy information text is generated through the obtained user intention understanding vector and transmitted to the naked-eye 3D digital human response generation module; Naked-eye 3D digital human response generation module: Based on the AI big model of naked-eye 3D digital human operation, the module generates the corresponding naked-eye 3D digital human response content text according to the generated personalized interaction strategy information text, and transmits it to the naked-eye 3D digital human human-computer interaction feedback and optimization module; Human-computer interaction and optimization module of naked-eye 3D digital human: Through big data analysis technology, the feedback parameters of users after receiving the response content text are analyzed, optimization is performed based on the abnormal analysis results, and the optimized content is recorded and transmitted to the administrator terminal.
[0014] See also Figure 2 As shown, a human-computer interaction method of naked-eye 3D digital human based on AI big model, including: S1: AI big model resource loading: loading the resources required for the operation of naked-eye 3D digital human into the trained AI big model, and building the AI big model for the operation of naked-eye 3D digital human, S2: multimodal input information fusion: based on the multimodal input device resources obtained in S1, collecting the multimodal information of the user initiating an interaction request to the naked-eye 3D digital human, and then fusing it through the AI big model to obtain the user input representation text, S3: AI big model intention understanding: based on the user input representation text IRT obtained in S2 text Understand the intention and obtain the user intention understanding vector. S4: Personalized strategy generation: Based on the AI big model running the naked-eye 3D digital human, generate personalized interaction strategy information text through the user intention understanding vector obtained in S3. S5: Naked-eye 3D digital human response generation: Based on the AI big model running the naked-eye 3D digital human, generate the corresponding naked-eye 3D digital human response content text according to the personalized interaction strategy information text generated by S4. S6: Through big data analysis technology, analyze the feedback parameters of the user after receiving the response content text, optimize according to the abnormal analysis results, and record the optimized content for transmission to the administrator terminal.
[0015] S1: AI big model resource loading: Load the resources required for naked-eye 3D digital human operation into the trained AI big model, and build an AI big model for naked-eye 3D digital human operation. The resources required for naked-eye 3D digital human operation include naked-eye 3D digital human resources, multimodal input device resources, and user historical interaction records; What needs to be specifically explained in this embodiment is that the resources required for the naked-eye 3D digital human include loading the 3D model of the naked-eye 3D digital human (such as the naked-eye 3D mapping algorithm), the animation library and related voice resources, and the voice resources include voice samples with different emotions and intonations; multimodal input device resources, such as sensors, microphones, cameras, etc., and calibration and testing to ensure that the user's input information can be accurately collected; user historical interaction records, such as interaction time, interaction content, user feedback, etc., provide data support for personalized interaction; loading the AI large model can ensure that the digital human can run efficiently and stably during the user interaction process, and provide smooth multimodal response capabilities.
[0016] S2: Multimodal input information fusion: Based on the multimodal input device resources obtained in S1, the multimodal information of the user's interactive request to the naked-eye 3D digital human is collected, and then fused through the AI big model to obtain the user input representation text, including the following steps: A1: First, the multimodal information data set MID in which the user initiates an interaction request to the naked-eye 3D digital human is collected through the multimodal information collection device. , M.I. i represents the i-th modal information, and n represents the number of types of multimodal information. Then, the corresponding multimodal information is identified through multimodal recognition technology to obtain the multimodal information recognition data set IRD. , IR i represents the result of the i-th modal information recognition, and n also represents the number of corresponding modal information recognition results; It should be specifically explained in this embodiment that multimodal information includes but is not limited to multimodal inputs such as voice, text, gestures, and expressions. For example, when a user consults a naked-eye 3D digital human for product information in a shopping mall, he or she can directly speak the question (voice input), or enter a text query on the interactive terminal next to him or her, or use specific gestures (such as pointing to the product area of interest) to assist in expressing needs; multimodal information collection equipment includes voice sensors, cameras, and interactive interfaces that receive text information input by users.
[0017] A2: First, the various multimodal information recognition data obtained in A1 are vectorized into a unified dimension through the sub-modal vectorization technology to obtain the multimodal recognition data vectorization data set VD1. , VD i represents the i-th modal vector, and n also represents the corresponding vectorized number of multimodal recognition data; then, through big data analysis technology, combined with the number of multimodal information combinations and the time difference between modal information i and modal information j, i<j, (i,j)∈n, the multimodal information consistency index MCI at time t is obtained. t , , C n 2 Indicates the number of multimodal information combinations obtained through the combination function, such as C 4 2 =6,∑ i<j Indicates traversing all modal combinations, (i, j)∈n, Δt ij It represents the time difference between modal information i and modal information j, and the unit is unified in ms. For example, the difference between the start of speech (t=1200ms) and the start of gesture (t=1210ms) is 10ms. 0 For comparison, if MCI ≥ MCI 0, indicating that the multimodal information consistency is good, otherwise it triggers the multimodal information collection mechanism, such as the digital human actively clarifying the inquiry, multimodal re-collection, etc.; finally, the multimodal recognition data vectorization dataset VD2 with good multimodal information consistency is obtained; What needs to be specifically explained in this embodiment is the modal vectorization technology, such as the frame-level feature extraction of Wav2Vec2.0 for the voice module, and the [CLS] tag vector of BERT-base for the text modality; multimodal information combinations such as voice-gesture, voice-text, etc.
[0018] A3: First, through the cross-modal attention mechanism, the multimodal recognition data vectorization dataset is obtained according to A2, combining each modal vector, each modal score, modality-specific projection matrix and multimodal information consistency index MCI at time t. t Generate fused feature vector E fusion , , α i represents the weight coefficient of the i-th mode vector, ,s i expresses the i-th modality score, τ represents the adjustment coefficient, , MCI t-1 represents the multimodal information consistency index at time t-1, β ij represents the cross-modal interaction coefficient, W i represents the modality-specific projection matrix, CrossAttn() represents the cross-attention mechanism function, which maps modal features of different dimensions to a unified space (such as d=512) to maintain the consistency of intent understanding processing; then, based on the AI large model running on naked-eye 3D digital humans, the fused feature vector is formed into the user input representation text IRT text ; What needs to be specifically explained in this embodiment is that the cross-modal attention mechanism generates the fusion feature E fusion is a unified semantic representation vector; multimodal recognition technology includes speech recognition technology for speech, image recognition technology for gestures, recognition technology for facial expressions, etc.; CrossAttn is a cross-attention mechanism function, which is a special attention mechanism used to process the relationship between two different sequences or modalities; different modal scores can be obtained by corresponding technologies, for example, the text modal score is obtained through the BERT model.
[0019] S3: AI big model intention understanding: user input representation text IRT based on S2 text Performing intent understanding to obtain the user intent understanding vector includes the following steps: B1: First, use the BERT model to perform IRT on the user input text text Perform hierarchical encoding to obtain the text feature vector Etext , E text =BERT base ([w 1 ,w 2 ,...w m ]) IRTtext , w n represents the mth word unit after word segmentation and part-of-speech tagging. BERT-base is a classic pre-trained language model architecture in natural language processing (NLP). Then, the memory vector M is constructed from the user's historical interaction records. t , M t =LSTM mem (E text t-1 ,E text ), E text t-1 Represents the previous conversation text feature vector, LSTM mem It is an improved long short-term memory network designed for memory enhancement; finally, the intention understanding vector IV is calculated based on the AI large model running on the naked-eye 3D digital human, IV=softmax(W×[E text ;M t ]+b), softmax() is an activation function used to convert a vector into a probability distribution, b represents a bias term, which is used to adjust the output of the model so that it can better fit the data, W represents the weight matrix, and the dimension of the weight matrix is R N×2d , N represents the number of output intent features, 2d represents the total dimension of the input text features, and the maximum, minimum and average values of the intent vector IV are combined to obtain the text intent. Figure 1 The consistency index SICI, , max(IV), min(IV) and avg(IV) represent the maximum, minimum and average values of the intention vector IV, respectively; B2: Translate the text Figure 1 Consistency Index SICI and Threshold SICI 0 For comparison, if SICI≥SICI 0 , indicating that the text intent understanding is effective, and the user intent understanding vector IV is output; otherwise, the intent understanding clarification mechanism is triggered, such as active clarification, example speech, etc., until the text intent is understood. Figure 1 Consistency index SICI ≥ SICI 0 Then stop; What needs to be specifically explained in this embodiment is that BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture, which performs well in a variety of NLP tasks, including: text classification; named entity recognition: extracting entities with specific meanings from text; question-answering system: providing relevant answers based on questions and text paragraphs, etc.; What needs to be specifically explained in this embodiment is that LSTM (Long Short-Term Memory) is used to encode text feature vectors and generate structured vectors.
[0020] S4: Personalized strategy generation: Based on the AI big model operated by naked-eye 3D digital human, the personalized interaction strategy information text is generated through the user intention understanding vector obtained in S3, including the following steps: C1: First, based on the user intention understanding vector, the MLP model is used to generate a personalized interaction strategy information vector PSF for each user. PSF=MLP([IV;IV his ]), IV his represents the historical interaction intention understanding vector under the target intention, IV represents the intention understanding vector, MLP is a multilayer perceptron function, which is a feedforward neural network; then the 3D spatial feature vector A of the intention understanding vector is calculated by combining the query matrix, key matrix and value matrix disp , , Q, K and V represent the query matrix, key matrix and value matrix respectively, and d represents the scaling factor, which is used to adjust the dimension of the vector; finally, the cosine similarity between the personalized interaction strategy information vector PSF and the historical interaction strategy information vector under the target intention is calculated through big data analysis technology, and then combined with the 3D spatial feature vector of the intention understanding vector, the personalized strategy generation effectiveness index PEI is obtained. , cos_sim represents the cosine similarity function, PSF his Represents the historical interaction strategy information vector under the target intention, ||·|| 2 represents the L2 norm of the vector, A ideal Represents the theoretical 3D space eigenvector; C2: Generate the performance index PEI and threshold PEI based on the personalized strategy 0 For comparison, if PEI ≥ PEI 0 , indicating that the generation of personalized interaction strategy is effective, and the personalized interaction strategy information text is output; otherwise, the interaction strategy adjustment mechanism is triggered, such as enabling the backup strategy template library, until the personalized strategy generation effectiveness index PEI ≥ PEI 0 Then stop; S5: naked-eye 3D digital human response generation: Based on the AI big model of naked-eye 3D digital human operation, the corresponding naked-eye 3D digital human response content text is generated according to the personalized interaction strategy information text generated in S4, and the naked-eye 3D digital human response content text is generated through the naked-eye 3D arrangement algorithm to generate naked-eye 3D animation and display it on the screen, including the following steps: D1: First, based on the AI big model of naked-eye 3D digital human operation, generate the multimodal response vector MID of naked-eye 3D digital human nD , VD nD ={m 1 ,m 2 ,...m I ...m M}=Generator(IV,PSF), Generator represents the multimodal generating function, m I represents the I-th modal response vector, M represents the number of multimodal response vectors, such as voice-action-expression trimodality, IV represents the intention understanding vector, and PSF represents the personalized interaction strategy information vector; then, through cross-modal synchronization technology, the multimodal response vector is timestamped and bound to obtain the multimodal response vector time deviation Δt sync , Δt sync =∑|Δt Ik |≤t_th, Δt Ik Indicates the time deviation between the I-th modal response and the k-th modal response, t_th indicates the time deviation threshold, for example, the deviation between the voice response and the action response is less than or equal to 40ms, I<K, (I,K)∈M; secondly, the emotion matching degree EC of the naked-eye 3D digital human is obtained, EC=cos_sim(VE,VF), cos_sim indicates the cosine similarity function, if EC≤0, it is 0, otherwise it is the EC value, VE indicates the voice emotion feature vector, VF indicates the facial emotion feature vector; finally, through big data analysis technology, combined with the time deviation Δt sync , the emotion matching degree EC and the number of visual artifacts appearing per unit time, and the multimodal response coordination index MRCI of the naked-eye 3D digital human is obtained. , Nar represents the number of visual artifacts that occur per unit time; D2: Compare the multimodal response coordination index MRCI of naked-eye 3D digital human with the threshold MRCI 0 For comparison, if MRCI ≥ MRCI 0, indicating that the multimodal response coordination of the naked-eye 3D digital human is good, the corresponding naked-eye 3D digital human response content text is output, and the naked-eye 3D digital human response content text is generated by the naked-eye 3D image arrangement algorithm to display the naked-eye 3D animation on the screen; otherwise, the multimodal response coordination mechanism is triggered, such as enabling the backup strategy template library, until the multimodal response coordination index MRCI of the naked-eye 3D digital human ≥ MRCI 0 Then stop; What needs to be specifically explained in this embodiment is that the naked-eye 3D imaging algorithm shoots multiple groups of left and right viewpoint images of different scenes and different shooting subjects, and inputs the left and right viewpoint images belonging to the same photographed subject into the constructed three-dimensional convolutional network model. After model processing, the corresponding left and right viewpoint fusion disparity map is obtained, and then the disparity value is converted into a depth distance value. Based on the disparity value, depth distance value, camera parameters and the principle of similar triangles, the three-dimensional coordinates of the photographed subject in the world coordinate system are calculated for three-dimensional reconstruction.
[0021] S6: Analyze the feedback parameters of the user after receiving the response content text through big data analysis technology, optimize according to the abnormal analysis results, and record the optimized content and transmit it to the administrator terminal, including the following steps: E1: Through big data technology, the total number of interactions between users and naked-eye 3D digital humans N_tot, the number of interactions that correctly understand user intentions N_cor, and the interaction response delay t_del are recorded to obtain the naked-eye 3D digital human interaction capability analysis index ICAI. , if (t_del-t_del 0 )≤0, it is 0, otherwise it is the calculated difference; E2: Compare the naked-eye 3D digital human interaction ability analysis index ICAI with the corresponding threshold ICAI 0 For comparison, if ICAI ≥ ICAI 0 , indicating that the naked-eye 3D digital human has good interaction capabilities. Otherwise, it means that the analysis is abnormal. Optimization is performed based on the abnormal analysis results, and the optimization content is recorded and transmitted to the administrator terminal. For example, the system conducts online learning and optimization of the AI large model and adjusts the parameters of the model to improve the accuracy of intent understanding and the quality of response generation. At the same time, according to the user's new needs and preferences, the personalized interaction strategy is updated to provide better services for the next interaction.
[0022] It should be specifically explained in this embodiment that the number of interactions refers to a complete interaction behavior, which is counted as one interaction. For example, if the user has multiple conversations with the digital human, then each conversation is counted as one interaction. After receiving the digital human's response content text, for example, if the user is not satisfied with the flight information or has other questions, such as "Is there a cheaper flight?", the user can ask the digital human again. The digital human combines the new user input with the previous interaction record to re-understand the intention and generate a response.
[0023] Secondly: In the drawings of the embodiments disclosed in the present invention, only the structures related to the embodiments disclosed in the present invention are involved, and other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of the present invention can be combined with each other; Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A human-computer interaction method of naked-eye 3D digital human based on AI large model, characterized by: include: S1: AI big model resource loading: Load the resources required for naked-eye 3D digital human operation into the trained AI big model, and build an AI big model for naked-eye 3D digital human operation. The resources required for naked-eye 3D digital human operation include naked-eye 3D digital human resources, multimodal input device resources, and user historical interaction records; S2: Multimodal input information fusion: Based on the multimodal input device resources obtained in S1, the multimodal information of the user's interactive request to the naked-eye 3D digital human is collected, and then fused through the AI big model to obtain the user input representation text; S3: AI large model intention understanding: Based on the user input representation text obtained in S2, the intention is understood to obtain the user intention understanding vector; S4: Personalized strategy generation: Based on the AI big model operated by naked-eye 3D digital human, the personalized interaction strategy information text is generated through the user intention understanding vector obtained in S3; S5: Naked-eye 3D digital human response generation: Based on the AI big model of naked-eye 3D digital human operation, the corresponding naked-eye 3D digital human response content text is generated according to the personalized interaction strategy information text generated in S4, and the naked-eye 3D digital human response content text is generated through the naked-eye 3D arrangement algorithm to generate naked-eye 3D animation and display it on the screen; S6: Analyze the feedback parameters of the user after receiving the response content text through big data analysis technology, optimize according to the abnormal analysis results, and record the optimized content and transmit it to the administrator terminal.
2. The human-computer interaction method of naked-eye 3D digital human based on AI large model according to claim 1, characterized in that: The user input representation text obtained in S2 includes: A1: firstly, a multimodal information data set MID in which the user initiates an interaction request to the naked-eye 3D digital human is collected through a multimodal information collection device, , M.I. i represents the i-th modal information, and n represents the number of types of multimodal information. Then, the corresponding multimodal information is identified through multimodal recognition technology to obtain the multimodal information recognition data set IRD. , IR i represents the result of the i-th modal information recognition, and n also represents the number of corresponding modal information recognition results.
3. The human-computer interaction method of naked-eye 3D digital human based on AI large model according to claim 2, characterized in that: The user input representation text obtained in S2 also includes: A2: firstly, various multimodal information recognition data obtained in A1 are vectorized in a unified dimension by using a sub-modal vectorization technology to obtain a multimodal recognition data vectorization data set VD1, , VD i represents the i-th modal vector, and n also represents the corresponding vectorized number of multimodal recognition data; then, through big data analysis technology, combined with the number of multimodal information combinations and the time difference between modal information i and modal information j, i<j, (i,j)∈n, the multimodal information consistency index MCI at time t is obtained. t , compare MCI with the threshold MCI0. If MCI≥MCI0, it means that the multimodal information consistency is good. Otherwise, the multimodal information collection mechanism is triggered. Finally, a multimodal recognition data vectorized dataset VD2 with good multimodal information consistency is obtained.
4. The human-computer interaction method of naked-eye 3D digital human based on AI large model according to claim 3 is characterized by: The step of obtaining the user input representation text in S2 further includes: firstly, obtaining a multimodal recognition data vectorization dataset according to A2 through a cross-modal attention mechanism, combining each modal vector, each modal score, modality-specific projection matrix and multimodal information consistency index MCI at time t t Generate fusion feature vector E fusion ; Then, based on the AI big model running on naked-eye 3D digital human, the fused feature vector is formed into the user input representation text IRT text .
5. The human-computer interaction method of naked-eye 3D digital human based on AI large model according to claim 1, characterized in that: The user intention understanding vector is obtained in S3. The following steps are involved: B1: First, use the BERT model to perform IRT on the user input text text Perform hierarchical encoding to obtain the text feature vector E text , E text =BERT base ([w1,w2,...w m ]) IRTtext , w n represents the mth lexical unit after word segmentation and part-of-speech tagging; then, the LSTM model is used to construct the memory vector M from the user's historical interaction records. t , M t =LSTM mem (E text t-1 ,E text ), E text t-1 Represents the feature vector of the previous dialogue text; finally The AI big model based on naked-eye 3D digital human operation calculates the intention understanding vector IV, IV=softmax(W×[E text ;M t ]+b), softmax() is the activation function, b represents the bias term, W represents the weight matrix, and the dimension of the weight matrix is R N×2d , N represents the number of output intent features, 2d represents the total dimension of the input text features, and the maximum, minimum and average values of the intent vector IV are combined to obtain the text intent consistency index SICI; B2: Compare the text intention consistency index SICI with the threshold SICI0. If SICI≥SICI0, it means that the text intention understanding is effective, and the user intention understanding vector IV is output; otherwise, the intention understanding clarification mechanism is triggered until the text intention consistency index SICI≥SICI0.
6. The human-computer interaction method of naked-eye 3D digital human based on AI large model according to claim 1, characterized in that: The step of generating the personalized interaction strategy information text in S4 includes the following steps: C1: First, based on the user intention understanding vector, the MLP model is used to generate a personalized interaction strategy information vector PSF for each user. PSF=MLP([IV;IV his ]), IV his represents the historical interaction intention understanding vector under the target intention, and IV represents the intention understanding vector; then the 3D spatial feature vector A of the intention understanding vector is calculated by combining the query matrix, key matrix and value matrix. disp ; Finally, through big data analysis technology, the cosine similarity between the personalized interaction strategy information vector PSF and the historical interaction strategy information vector under the target intention is calculated, and then combined with the 3D spatial feature vector of the intention understanding vector, the personalized strategy generation effectiveness index PEI is obtained; C2: Compare the personalized strategy generation effectiveness index PEI with the threshold PEI0. If PEI≥PEI0, it means that the personalized interaction strategy generation is effective, and the personalized interaction strategy information text is output; otherwise, the interaction strategy adjustment mechanism is triggered until the personalized strategy generation effectiveness index PEI≥PEI0.
7. The human-computer interaction method of naked-eye 3D digital human based on AI large model according to claim 1, characterized in that: In S5, a corresponding naked-eye 3D digital human response content text is generated. The following steps are involved: D1: First, based on the AI big model of naked-eye 3D digital human operation, generate the multimodal response vector MID of naked-eye 3D digital human nD , VD nD ={m1,m2,...m I ...m M }=Generator(IV,PSF), Generator represents the multimodal generating function, m I represents the I-th modal response vector, M represents the number of multimodal response vectors, IV represents the intention understanding vector, and PSF represents the personalized interaction strategy information vector; then, through cross-modal synchronization technology, the multimodal response vector is timestamped and bound to obtain the multimodal response vector time deviation Δt sync , Δt sync =∑|Δt Ik |≤t_th, Δt Ik represents the time deviation between the I-th modal response and the k-th modal response, t_th represents the time deviation threshold, I<K, (I,K)∈M; secondly, the emotion matching degree EC of the naked-eye 3D digital human is obtained, EC=cos_sim(VE,VF), cos_sim represents the cosine similarity function, if EC≤0, it is 0, otherwise it is the EC value, VE represents the speech emotion feature vector, VF represents the facial emotion feature vector; finally, through big data analysis technology, combined with the time deviation Δt sync , the emotion matching degree EC and the number of visual artifacts appearing per unit time, and the multimodal response coordination index MRCI of the naked-eye 3D digital human is obtained; D2: Compare the multimodal response coordination index MRCI of the naked-eye 3D digital human with the threshold MRCI0. If MRCI≥MRCI0, it means that the multimodal response coordination of the naked-eye 3D digital human is good, and the corresponding naked-eye 3D digital human response content text is output, and the naked-eye 3D digital human response content text is generated through the naked-eye 3D image arrangement algorithm to display the naked-eye 3D animation on the screen; otherwise, the multimodal response coordination mechanism is triggered until the multimodal response coordination index MRCI≥MRCI0 of the naked-eye 3D digital human.
8. The human-computer interaction method of naked-eye 3D digital human based on AI large model according to claim 1, characterized in that: The S6 comprises the following steps: E1: Through big data technology, the total number of interactions between users and naked-eye 3D digital humans N_tot, the number of interactions that correctly understand user intentions N_cor, and the interaction response delay t_del are recorded to obtain the naked-eye 3D digital human interaction capability analysis index ICAI; E2: Compare the naked-eye 3D digital human interaction ability analysis index ICAI with the corresponding threshold ICAI0. If ICAI ≥ ICAI0, it means that the naked-eye 3D digital human interaction ability is good. Otherwise, it means that the analysis is abnormal. Optimize according to the abnormal analysis results, and record the optimization content and transmit it to the administrator terminal.
Citation Information
Patent Citations
Digital human control method and device based on multiple modes and electronic equipment
CN119441403A
Interactive digital human generation method and system based on artificial intelligence
CN119600159A
Multi-modal interaction method and apparatus, controller, system, automobile, and storage medium
WO2025066925A1
Cited By
Interaction optimization implementation method and device applied to digital human
CN120631190A
Interaction method and interaction terminal fusing multi-mode perception and naked eye 3D
CN122363500A