A Human-Computer Interaction Method for a Naked-Eye 3D Digital Human Based on an AI Large Model
Through multimodal information fusion and personalized strategy generation based on AI big model, the problems of inaccurate semantic understanding and insufficient real-timeness in human-computer interaction of naked-eye 3D digital humans are solved, and a personalized and efficient user interaction experience is achieved.
Patent Information
- Application Number
- CN202510579163.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The existing human-computer interaction methods for naked-eye 3D digital humans have problems such as inaccurate semantic understanding, lack of personalization of interactions and insufficient real-time performance, resulting in poor user experience.
Using an AI big model-based method, user information is collected through multimodal input devices, AI big model is used for intention understanding and personalized strategy generation, combined with user historical interaction records, personalized interaction response is generated, and interaction quality is optimized through big data analysis.
It achieves more accurate user intention understanding, provides personalized interactive services, improves the naturalness and fluency of interaction, and continuously improves interaction capabilities through feedback optimization.
Smart Images

Figure CN120104009B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and specifically to a human-computer interaction method for a naked-eye 3D digital human based on an AI large model. Background Art
[0002] With the continuous development of technology, significant progress has been made in human-computer interaction technology; as an emerging human-computer interaction carrier, digital humans can interact with users in an anthropomorphic image and have received extensive attention and applications in multiple fields; the naked-eye 3D technology enables users to directly view the digital human image with a three-dimensional effect without the need for special devices such as glasses, enhancing the immersion and attraction of the interaction.
[0003] Existing digital human interaction methods are mostly rule-based natural language processing interactions, which parse the natural language input by users according to pre-set rules and grammars; use a general interaction template to interact with all users, for example, regardless of the user's interests and preferences, respond with the same words and processes; when generating the response actions, expressions, and voice replies of the naked-eye 3D digital human, traditional computing and rendering technologies are used to achieve the interaction of the digital human.
[0004] However, there are still many problems in the current human-computer interaction based on naked-eye 3D digital humans: it is difficult to accurately understand the true intentions of users by parsing according to pre-set rules and grammars, which is prone to misunderstandings or inability to give accurate responses, resulting in inaccurate semantic understanding; using a general interaction template, it is impossible to provide customized interaction services according to the personalized characteristics of users, and it is difficult to meet the diverse needs of users, and the interaction lacks personalization; when generating the response actions, expressions, and voice replies of the naked-eye 3D digital human, due to the complex computing and rendering processes involved, some systems have problems with response delays, resulting in unsmooth interactions and affecting the user experience, and the real-time performance is insufficient. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a human-computer interaction method for a naked-eye 3D digital human based on an AI large model to solve the problems of inaccurate semantic understanding, lack of personalization in interaction, and insufficient real-time performance mentioned in the above background art.
[0006] To achieve the above object, the present invention provides the following technical solution: A human-computer interaction method for a naked-eye 3D digital human based on an AI large model, including:
[0007] S1: Loading of AI large model resources: Load the resources required for the operation of the naked-eye 3D digital human into the trained AI large model to construct an AI large model for the operation of the naked-eye 3D digital human. The resources required for the operation of the naked-eye 3D digital human include naked-eye 3D digital human resources, multi-modal input device resources, and user historical interaction records;
[0008] S2: Multi-modal Input Information Fusion: Based on the multi-modal input device resources obtained in S1, collect the multi-modal information of the user's interaction request initiated to the naked-eye 3D digital human, and then fuse it through the AI large model to obtain the user input representation text;
[0009] S3: AI Large Model Intent Understanding: Conduct intent understanding based on the user input representation text obtained in S2 to obtain the user intent understanding vector;
[0010] S4: Personalized Strategy Generation: Based on the AI large model running on the naked-eye 3D digital human, generate the personalized interaction strategy information text through the user intent understanding vector obtained in S3;
[0011] S5: Naked-eye 3D Digital Human Response Generation: Based on the AI large model running on the naked-eye 3D digital human, generate the corresponding naked-eye 3D digital human response content text according to the personalized interaction strategy information text generated in S4, and generate the naked-eye 3D animation through the naked-eye 3D layout algorithm and display it on the screen;
[0012] S6: Through big data analysis technology, analyze the feedback parameters after the user receives the response content text, optimize according to the abnormal analysis results, and record the optimization content and transmit it to the administrator terminal.
[0013] Technical Effects and Advantages of the Present Invention:
[0014] 1. Through various sensor technologies and recognition technologies, the present invention collects multi-modal interaction information through the interaction request initiated by the user standing in front of the naked-eye 3D digital human device, identifies and integrates it to form a unified user input representation, providing more comprehensive and rich information for subsequent intent understanding, and thus more accurately determining the user's intent object;
[0015] 2. By constructing the AI large model running on the naked-eye 3D digital human, the present invention realizes the fusion interaction of multi-modal information, can deeply understand the user's intent and needs, combines the user's historical interaction records and real-time situations, provides personalized responses and services, meets the diverse needs of different users, and makes the communication between the naked-eye 3D digital human and the user more natural and fluent, just like the interaction between real people;
[0016] 3. By receiving the user's feedback and optimizing the AI large model, the present invention can continuously improve the quality of human-computer interaction; at the same time, according to the user's new needs and preferences, update the personalized interaction strategy to provide better services for the next interaction, and thus continuously improve the intelligent level and interaction ability of the naked-eye 3D digital human over time. Brief Description of the Drawings
[0017] Figure 1This is the overall process schematic diagram of the present invention.
[0018] Figure 2 This is the method process schematic diagram of the present invention. Detailed implementation manners
[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0020] Please refer to Figure 1 As shown, the present invention provides a human-computer interaction system for a naked-eye 3D digital human based on an AI large model, including an AI large model resource loading module, a multi-modal input fusion module, an AI large model intention understanding module, a personalized interaction strategy generation module, a naked-eye 3D digital human response generation module, and a human-computer interaction feedback and optimization module for the naked-eye 3D digital human.
[0021] The AI large model resource loading module is connected to all the other modules. The multi-modal input fusion module is connected to the AI large model intention understanding module. The personalized interaction strategy generation module is respectively connected to the AI large model intention understanding module and the naked-eye 3D digital human response generation module. The human-computer interaction feedback and optimization module for the naked-eye 3D digital human is connected to the naked-eye 3D digital human response generation module.
[0022] AI large model resource loading module: It is used to load the resources required for the operation of the naked-eye 3D digital human for the trained AI large model, and construct the AI large model for the operation of the naked-eye 3D digital human. The resources required for the operation of the naked-eye 3D digital human include naked-eye 3D digital human resources, multi-modal input device resources, and user historical interaction records.
[0023] Multi-modal input fusion module: Based on the multi-modal input device resources obtained by the AI large model resource loading module, it collects and fuses the multi-modal information of the user's interaction request to the naked-eye 3D digital human, and transmits the fused user request information to the AI large model intention understanding module.
[0024] AI large model intention understanding module: Based on the AI large model for the operation of the naked-eye 3D digital human, it conducts intention understanding on the fused user request information to obtain the user intention understanding vector, and transmits it to the personalized interaction strategy generation module.
[0025] Personalized interaction strategy generation module: Based on the AI large model for the operation of the naked-eye 3D digital human, it generates the personalized interaction strategy information text through the obtained user intention understanding vector, and transmits it to the naked-eye 3D digital human response generation module.
[0026] Bare-eye 3D digital human response generation module: Based on the AI large model for the operation of the bare-eye 3D digital human, generate the corresponding bare-eye 3D digital human response content text according to the generated personalized interaction strategy information text, and transmit it to the human-computer interaction feedback and optimization module of the bare-eye 3D digital human;
[0027] Human-computer interaction and optimization module of the bare-eye 3D digital human: Analyze the feedback parameters after the user receives the response content text through big data analysis technology, optimize according to the abnormal analysis results, and record the optimization content and transmit it to the administrator terminal.
[0028] Please refer to Figure 2 As shown, a human-computer interaction method for a bare-eye 3D digital human based on an AI large model includes: S1: AI large model resource loading: Load the resources required for the operation of the bare-eye 3D digital human for the trained AI large model, and construct the AI large model for the operation of the bare-eye 3D digital human. S2: Multi-modal input information fusion: Based on the multi-modal input device resources obtained in S1, collect the multi-modal information of the user's interaction request to the bare-eye 3D digital human, and then fuse it through the AI large model to obtain the user input representation text. S3: AI large model intention understanding: Based on the user input representation text IRT text obtained in S2, perform intention understanding to obtain the user intention understanding vector. S4: Personalized strategy generation: Based on the AI large model for the operation of the bare-eye 3D digital human, generate the personalized interaction strategy information text through the user intention understanding vector obtained in S3. S5: Bare-eye 3D digital human response generation: Based on the AI large model for the operation of the bare-eye 3D digital human, generate the corresponding bare-eye 3D digital human response content text according to the personalized interaction strategy information text generated in S4. S6: Through big data analysis technology, analyze the feedback parameters after the user receives the response content text, optimize according to the abnormal analysis results, and record the optimization content and transmit it to the administrator terminal.
[0029] S1: AI large model resource loading: Load the resources required for the operation of the bare-eye 3D digital human for the trained AI large model, and construct the AI large model for the operation of the bare-eye 3D digital human. The resources required for the operation of the bare-eye 3D digital human include bare-eye 3D digital human resources, multi-modal input device resources, and user historical interaction records;
[0030] In this embodiment, it should be specifically noted that the resources required for the naked-eye 3D digital human include loading the 3D model of the naked-eye 3D digital human (such as the naked-eye 3D layout algorithm), the animation library, and the relevant voice resources. The voice resources include voice samples with different emotions and intonations; multi-modal input device resources, such as sensors, microphones, cameras, etc., and calibration and testing are carried out to ensure that the input information of the user can be accurately collected; the user's historical interaction records, such as interaction time, interaction content, user feedback, etc., provide data support for personalized interaction; loading the AI large model can ensure that the digital human can operate efficiently and stably during user interaction and provide smooth multi-modal response capabilities.
[0031] S2: Multi-modal input information fusion: Based on the multi-modal input device resources obtained in S1, collect the multi-modal information of the user's interaction request to the naked-eye 3D digital human, and then fuse it through the AI large model to obtain the user input representation text, including the following steps:
[0032] A1: First, through the multi-modal information collection device, collect the multi-modal information dataset MID of the user's interaction request to the naked-eye 3D digital human, , MI i represents the i-th type of modal information, and n represents the number of types of multi-modal information; then through multi-modal recognition technology, identify the corresponding multi-modal information to obtain the multi-modal information recognition dataset IRD, , IR i represents the result of the i-th type of modal information recognition, and n also represents the number of results of the corresponding modal information recognition;
[0033] In this embodiment, it should be specifically noted that the multi-modal information includes but is not limited to multi-modal inputs such as voice, text, gestures, expressions, etc. For example, when a user consults product information from a naked-eye 3D digital human in a mall, the user can directly speak out the question (voice input), or enter text for query on the adjacent interaction terminal, or use specific gesture actions (such as pointing to the area of the product of interest) to assist in expressing the needs; the multi-modal information collection device includes voice sensors, cameras, and the interaction interface for receiving the text information input by the user, etc.
[0034] A2: First, through the per-modal vectorization technology, vectorize the various multi-modal information recognition data obtained in A1 to a unified dimension to obtain the multi-modal recognition data vectorization dataset VD1, , VD i represents the i-th type of modal vector, and n also represents the corresponding vectorization quantity of the multi-modal recognition data; then through big data analysis technology, combined with the number of multi-modal information combinations and the time difference between modal information i and modal information j, i < j, (i, j) ∈ n, to obtain the multi-modal information consistency index MCI at time t t , , C n 2 represents the number of multimodal information combinations obtained through the combination function, such as C4 2 = 6, ∑ i<j represents traversing all modal combinations, (i, j) ∈ n, Δt ij represents the time difference between modal information i and modal information j, with the unit unified as ms. For example, the difference between the start of speech (t = 1200ms) and the start of gesture (t = 1210ms) is 10ms. Compare MCI with the threshold MCI0. If MCI ≥ MCI0, it indicates good multimodal information consistency; otherwise, trigger the multimodal information acquisition mechanism, such as the digital human actively clarifying and asking questions, multimodal re-acquisition, etc. Finally, obtain the vectorized dataset VD2 of multimodal recognition data with good multimodal information consistency;
[0035] It should be specifically noted in this embodiment that the cross-modal vectorization technology, for example, for the speech module, the frame-level feature extraction of Wav2Vec2.0 is adopted, and for the text modality, the [CLS] token vector of BERT-base is adopted, etc.; multimodal information combinations such as speech-gesture, speech-text, etc.
[0036] A3: First, through the cross-modal attention mechanism, obtain the vectorized dataset of multimodal recognition data according to A2, and combine each modal vector, each modal score, the modal-specific projection matrix, and the multimodal information consistency index MCI at time t t Generate the fused feature vector E fusion , , α i represents the weight coefficient of the i-th modal vector, , s i expresses the score of the i-th modality, τ represents the adjustment coefficient, , MCI t-1 represents the multimodal information consistency index at time t - 1, β ij represents the cross-modal interaction coefficient, W i represents the modal-specific projection matrix, CrossAttn() represents the cross-attention mechanism function, which maps modal features of different dimensions to a unified space (such as d = 512) to maintain the consistency of intent understanding processing; then, based on the AI large model running on the naked-eye 3D digital human, form the user input representation text IRT from the fused feature vector text ;
[0037] It should be specifically noted in this embodiment that the cross-modal attention mechanism generates the fused feature E fusionis a unified semantic representation vector; multi-modal recognition technologies include speech recognition technology for speech, image recognition technology for gestures, recognition technology for expressions, etc.; CrossAttn is a cross-attention mechanism function, which is a special attention mechanism used to process the relationship between two different sequences or modalities; different modality scores can be obtained by corresponding technologies. For example, the text modality score is obtained through the BERT model.
[0038] S3: AI large model intention understanding: Based on the user input representation text IRT obtained in S2 text perform intention understanding to obtain the user intention understanding vector, including the following steps:
[0039] B1: First, use the BERT model to perform hierarchical encoding on the user input text IRT text to obtain the text feature vector E text , E text =BERT base ([w1, w2,... w m ) IRTtext , w n represents the m-th lexical unit after word segmentation and part-of-speech tagging. BERT-base is a classic pre-trained language model architecture in natural language processing (NLP); then construct the memory vector M from the user's historical interaction records t , M t =LSTM mem (E text t-1 , E text ), E text t-1 represents the previous dialogue text feature vector. LSTM mem is an improved long short-term memory network designed specifically for memory enhancement; finally, calculate the intention understanding vector IV based on the AI large model running on the naked-eye 3D digital human. IV = softmax(W × [E text ; M t + b), softmax() is an activation function used to convert a vector into a probability distribution, b represents the bias term used to adjust the output of the model so that it can better fit the data, and W represents the weight matrix. The dimension of the weight matrix is R N×2d , N represents the number of output intention features, 2d represents the total dimension of the input text features. Combine the maximum value, minimum value, and average value of the intention vector IV to obtain the text Figure 1 consistency index SICI, , max(IV), min(IV), and avg(IV) represent the maximum value, minimum value, and average value of the intention vector IV respectively;
[0040] B2: Interpret the text meaning Figure 1 Compare the semantic consistency index SICI with the threshold SICI0. If SICI ≥ SICI0, it indicates that the text intention understanding is effective, and the user intention understanding vector IV is output; otherwise, trigger the intention understanding clarification mechanism, such as active clarification, example words, etc., until the text meaning Figure 1 The semantic consistency index SICI ≥ SICI0, then stop;
[0041] It should be specifically noted in this embodiment that BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture, which performs well in a variety of NLP tasks, including: text classification; named entity recognition: extracting entities with specific meanings from text; question answering systems: providing relevant answers according to questions and text paragraphs, etc.; It should be specifically noted in this embodiment that LSTM (Long Short-Term Memory) is used to encode the text feature vector to generate a structured vector.
[0042] S4: Personalized strategy generation: Based on the AI large model for the operation of the naked-eye 3D digital human, generate personalized interaction strategy information text through the user intention understanding vector obtained in S3, including the following steps:
[0043] C1: First, based on the user intention understanding vector, through the MLP model, generate a personalized interaction strategy information vector PSF for each user, PSF = MLP([IV;IV his )), IV his represents the historical interaction intention understanding vector under the target intention, IV represents the intention understanding vector, and MLP is the Multilayer Perceptron function, which is a feedforward neural network; then combine the query matrix, key matrix, and value matrix to calculate the 3D spatial feature vector A of the intention understanding vector disp , , Q, K, and V represent the query matrix, key matrix, and value matrix respectively, and d represents the scaling factor used to adjust the dimension of the vector; finally, through big data analysis technology, calculate the cosine similarity between the personalized interaction strategy information vector PSF and the historical interaction strategy information vector under the target intention, and then combine the 3D spatial feature vector of the intention understanding vector to obtain the personalized strategy generation efficiency index PEI, , cos_sim represents the cosine similarity function, PSF his represents the historical interaction strategy information vector under the target intention, ||·||2 represents the L2 norm of the vector, and A ideal represents the theoretical 3D spatial feature vector;
[0044] C2: Compare the personalized policy generation efficiency index PEI with the threshold PEI0. If PEI≥PEI0, it indicates that the generation of the personalized interaction policy is effective, and output the personalized interaction policy information text; otherwise, trigger the interaction policy adjustment mechanism, such as enabling the backup policy template library, and stop until the personalized policy generation efficiency index PEI≥PEI0;
[0045] S5: Generation of the response of the naked-eye 3D digital human: Based on the AI large model for the operation of the naked-eye 3D digital human, generate the corresponding naked-eye 3D digital human response content text according to the personalized interaction policy information text generated in S4, and generate a naked-eye 3D animation through the naked-eye 3D layout algorithm and display it on the screen, including the following steps:
[0046] D1: First, based on the AI large model for the operation of the naked-eye 3D digital human, generate the multi-modal response vector MID nD , VD nD ={m1,m2,...m I ...m M}=Generator(IV,PSF), where Generator represents the multi-modal generation function, m I represents the response vector of the I-th modality, M represents the number of multi-modal response vectors, such as the three modalities of voice-action-expression, IV represents the intention understanding vector, and PSF represents the personalized interaction policy information vector; then, through the cross-modal synchronization technology, perform timestamp binding on the multi-modal response vector to obtain the time deviation Δt sync , Δt sync =∑|Δt Ik |≤t_th, where Δt Ik represents the time deviation between the response of the I-th modality and the response of the k-th modality, and t_th represents the time deviation threshold. For example, the deviation between the voice response and the action response is less than or equal to 40ms, I<K, (I,K)∈M; Secondly, obtain the emotional matching degree EC of the naked-eye 3D digital human, EC=cos_sim(VE,VF), where cos_sim represents the cosine similarity function. If EC≤0, it is 0, otherwise it is the EC value, VE represents the voice emotional feature vector, and VF represents the facial emotional feature vector; Finally, through the big data analysis technology, combine the time deviation Δt sync , the emotional matching degree EC, and the number of visual artifacts appearing per unit time to obtain the multi-modal response coordination index MRCI of the naked-eye 3D digital human, , Nar represents the number of visual artifacts appearing per unit time;
[0047] D2: Compare the multi-modal response coordination index MRCI of the naked-eye 3D digital human with the threshold MRCI0. If MRCI ≥ MRCI0, it indicates that the multi-modal response of the naked-eye 3D digital human is well coordinated. Output the corresponding text of the naked-eye 3D digital human response content, and generate a naked-eye 3D animation through the naked-eye 3D layout algorithm and display it on the screen; otherwise, trigger the multi-modal response coordination mechanism, such as enabling the alternative strategy template library, until the multi-modal response coordination index MRCI of the naked-eye 3D digital human ≥ MRCI0 and then stop;
[0048] Specifically in this embodiment, the naked-eye 3D layout algorithm captures multiple groups of left and right viewpoint images of different scenes and different shooting subjects, inputs the left and right viewpoint images belonging to the same shooting subject into the constructed three-dimensional convolutional network model, and obtains the corresponding left and right viewpoint fusion disparity maps after model processing. Then, the disparity values are converted into depth distance values, and the three-dimensional coordinates of the shooting subject in the world coordinate system are calculated based on the disparity values, depth distance values, camera parameters, and the principle of similar triangles for three-dimensional reconstruction.
[0049] S6: Through big data analysis technology, analyze the feedback parameters after the user receives the response content text, optimize according to the anomaly analysis results, and record the optimization content and transmit it to the administrator terminal, including the following steps:
[0050] E1: Through big data technology, record the total number of interactions N_tot between the user and the naked-eye 3D digital human and the number of interactions N_cor that correctly understand the user's intention, and the interaction response delay t_del, to obtain the interaction ability analysis index ICAI of the naked-eye 3D digital human. , if (t_del - t_del0) ≤ 0, then it is 0, otherwise it is the calculated difference;
[0051] E2: Compare the interaction ability analysis index ICAI of the naked-eye 3D digital human with the corresponding threshold ICAI0. If ICAI ≥ ICAI0, it indicates that the interaction ability of the naked-eye 3D digital human is good; otherwise, it indicates that the analysis is abnormal. Optimize according to the anomaly analysis results, and record the optimization content and transmit it to the administrator terminal. For example, the system performs online learning and optimization on the AI large model, adjusts the parameters of the model to improve the accuracy of intention understanding and the quality of response generation; at the same time, update the personalized interaction strategy according to the new needs and preferences of the user to provide better services for the next interaction.
[0052] It should be specifically noted in this embodiment that the number of interactions refers to a complete interaction behavior, which is counted as one interaction. For example, if a user has multiple conversations with a digital human, each conversation is counted as one interaction; after the user receives the response text from the digital human, for example, if they are not satisfied with the flight information or have other questions, such as "Are there any cheaper flights?", they can ask the digital human again. The digital human combines the new user input with the previous interaction records and re-performs intent understanding and response generation.
[0053] Secondly: In the accompanying drawings of the disclosed embodiments of the present invention, only the structures related to the disclosed embodiments are involved. Other structures can refer to the general design. Without conflict, the same embodiment and different embodiments of the present invention can be combined with each other;
[0054] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A human-computer interaction method for a naked-eye 3D digital human based on an AI large model, characterized in that: Including: S1: Loading AI large model resources: Load the resources required for the operation of the naked-eye 3D digital human by the trained AI large model, and construct the AI large model for the operation of the naked-eye 3D digital human. The resources required for the operation of the naked-eye 3D digital human include naked-eye 3D digital human resources, multi-modal input device resources, and user historical interaction records; S2: Multi-modal input information fusion: Based on the multi-modal input device resources obtained in S1, collect the multi-modal information of the user's interaction request to the naked-eye 3D digital human, and then fuse it through the AI large model to obtain the user input representation text; S3: AI large model intention understanding: Conduct intention understanding based on the user input representation text obtained in S2 to obtain the user intention understanding vector; S4: Personalized strategy generation: Based on the AI large model for the operation of the naked-eye 3D digital human, generate the personalized interaction strategy information text through the user intention understanding vector obtained in S3; S5: Naked-eye 3D digital human response generation: Based on the AI large model for the operation of the naked-eye 3D digital human, generate the corresponding naked-eye 3D digital human response content text according to the personalized interaction strategy information text generated in S4, and generate a naked-eye 3D animation through the naked-eye 3D layout algorithm and display it on the screen; Generating the corresponding naked-eye 3D digital human response content text in S5, including the following steps: D1: First, based on the AI large model for the operation of the naked-eye 3D digital human, generate the multi-modal response vector MID of the naked-eye 3D digital human nD , MID nD = {m1, m2,... m I ... m M} = Generator(IV, PSF), where Generator represents the multi-modal generation function, m I represents the I-th modal response vector, M represents the number of multi-modal response vectors, IV represents the intention understanding vector, and PSF represents the personalized interaction strategy information vector; then, through the cross-modal synchronization technology, timestamp binding is performed on the multi-modal response vector to obtain the time deviation Δt of the multi-modal response vector sync , Δt sync = ∑|Δt Ik | ≤ t_th, where Δt Ik represents the time deviation between the I-th modal response and the k-th modal response, t_th represents the time deviation threshold, I < K, I ∈ M, K ∈ M; secondly, obtain the emotional matching degree EC of the naked-eye 3D digital human, EC = cos_sim(VE, VF), where cos_sim represents the cosine similarity function. If EC ≤ 0, then it is 0, otherwise it is the EC value. VE represents the voice emotion feature vector, and VF represents the facial emotion feature vector; finally, through the big data analysis technology, combined with the time deviation Δt sync , the emotional matching degree EC, and the number of visual artifacts appearing per unit time, obtain the multi-modal response coordination index MRCI of the naked-eye 3D digital human; D2: Compare the multi-modal response coordination index MRCI of the naked-eye 3D digital human with the threshold MRCI0. If MRCI≥MRCI0, it means that the multi-modal response of the naked-eye 3D digital human is well coordinated. Output the corresponding naked-eye 3D digital human response content text, and generate a naked-eye 3D animation through the naked-eye 3D layout algorithm and display it on the screen; otherwise, trigger the multi-modal response coordination mechanism until the multi-modal response coordination index MRCI of the naked-eye 3D digital human≥MRCI0 and then stop; S6: Through big data analysis technology, analyze the feedback parameters after the user receives the response content text, optimize according to the abnormal analysis results, and record the optimized content and transmit it to the administrator terminal.
2. The human-computer interaction method of a naked-eye 3D digital human based on an AI large model according to claim 1, wherein: The user input representation text obtained in S2 includes: A1: First, through a multi-modal information collection device, a multi-modal information data set MID of the user's interaction request with the naked-eye 3D digital human is collected. , MI i represents the i-th type of modal information, and n represents the number of types of multi-modal information; then, through multi-modal recognition technology, the corresponding multi-modal information is recognized to obtain a multi-modal information recognition data set IRD. , IR i represents the result of the recognition of the i-th type of modal information, and n also represents the number of results of the recognition of the corresponding modal information.
3. The human-computer interaction method of a naked-eye 3D digital human based on an AI large model according to claim 2, wherein: The user input obtained in S2 further includes: A2: First, through the cross-modal vectorization technology, the various multi-modal information recognition data obtained in A1 are vectorized to a unified dimension to obtain a cross-modal recognition data vectorization dataset VD1. , VD i represents the i-th modal vector, and n also represents the corresponding vectorization quantity of the multi-modal recognition data; then, through the big data analysis technology, combined with the cross-modal information combination quantity and the time difference between modal information i and modal information j, i < j, i ∈ n, j ∈ n, the cross-modal information consistency index MCI at time t is obtained. t , MCI t is compared with the threshold MCI0. If MCI t ≥ MCI0, it indicates that the cross-modal information consistency is good; otherwise, the cross-modal information acquisition mechanism is triggered; finally, a cross-modal recognition data vectorization dataset VD2 with good cross-modal information consistency is obtained.
4. The human-computer interaction method of a naked-eye 3D digital human based on an AI large model according to claim 3, wherein: The user input representation text obtained in S2 further includes: First, through the cross-modal attention mechanism, a multi-modal recognition data vectorized dataset is obtained according to A2, and combined with each modal vector, each modal score, the modal-specific projection matrix, and the multi-modal information consistency index MCI at time t t to generate a fused feature vector E fusion ; Then, based on the AI large model running on the naked-eye 3D digital human, the fused feature vector is formed into the user input representation text IRT text .
5. A human-computer interaction method for a naked-eye 3D digital human based on an AI large model according to claim 1, characterized in that: Obtaining the user intention understanding vector in S3, including the following steps: B1: First, the BERT model is used to perform hierarchical encoding on the user input representation text IRT text to obtain the text feature vector E text , E text = BERT base ([w1, w2,... w m ) IRTtext , w m represents the m-th lexical unit after word segmentation and part-of-speech tagging; then, through the LSTM model, the memory vector M is constructed from the user's historical interaction records t , M t = LSTM mem (E text t-1 , E text ), E text t-1 represents the previous dialogue text feature vector; finally The AI large model running based on the naked-eye 3D digital human calculates the intention understanding vector IV, IV = softmax(W × [E text ;M t +b), where softmax() is the activation function, b represents the bias term, W represents the weight matrix, and the dimension of the weight matrix is R N×2d , N represents the number of output intention features, 2d represents the total dimension of the input text features, and by combining the maximum value, minimum value, and average value of the intention vector IV, the text intention consistency index SICI is obtained; B2: Compare the text intention consistency index SICI with the threshold SICI0. If SICI≥SICI0, it means that the text intention understanding is effective, and output the user intention understanding vector IV; otherwise, trigger the intention understanding clarification mechanism until the text intention consistency index SICI≥SICI0 and then stop.
6. A human-computer interaction method for a naked-eye 3D digital human based on an AI large model according to claim 1, characterized in that: Generating the personalized interaction strategy information text in S4 includes the following steps: C1: First, based on the user intention understanding vector, through the MLP model, generate a personalized interaction strategy information vector PSF for each user, PSF = MLP([IV;IV his ), where IV his represents the historical interaction intention understanding vector under the target intention, and IV represents the intention understanding vector; then, combine the query matrix, key matrix, and value matrix to calculate the 3D spatial feature vector A disp of the intention understanding vector; finally, through big data analysis technology, calculate the cosine similarity between the personalized interaction strategy information vector PSF and the historical interaction strategy information vector under the target intention, and then combine it with the 3D spatial feature vector of the intention understanding vector to obtain the personalized strategy generation efficiency index PEI; C2: Compare the personalized strategy generation efficiency index PEI with the threshold PEI0. If PEI≥PEI0, it means that the personalized interaction strategy generation is effective, and output the personalized interaction strategy information text; otherwise, trigger the interaction strategy adjustment mechanism until the personalized strategy generation efficiency index PEI≥PEI0 and then stop.
7. A human-computer interaction method for a naked-eye 3D digital human based on an AI large model according to claim 1, characterized in that: S6 includes the following steps: E1: Through big data technology, record the total number of interactions N_tot between the user and the naked-eye 3D digital human and the number of interactions N_cor that correctly understand the user's intention, as well as the interaction response delay t_del, to obtain the interaction ability analysis index ICAI of the naked-eye 3D digital human; E2: Compare the interaction ability analysis index ICAI of the naked-eye 3D digital human with the corresponding threshold ICAI0. If ICAI ≥ ICAI0, it indicates that the interaction ability of the naked-eye 3D digital human is good; otherwise, it indicates an abnormal analysis. Optimize according to the abnormal analysis results and record the optimization content for transmission to the administrator terminal.
Citation Information
Patent Citations
Digital human control method and device based on multiple modes and electronic equipment
CN119441403A
Multi-modal interaction method and apparatus, controller, system, automobile, and storage medium
WO2025066925A1