Digital human interaction method and system based on naked-eye 3D visualization

Through multimodal input processing and hierarchical retrieval strategies, combined with preset content databases and enterprise knowledge bases, a digital human-driven engine is used to generate 3D visual answers, solving the response efficiency and security problems of the existing system under complex enterprise businesses, and achieving efficient and accurate user interaction.

CN119782466BActive Publication Date: 2025-08-15SHENZHEN YIGU CULTURE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411847986.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-08-15
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

The existing intelligent digital human interaction system is difficult to respond to user problems efficiently and accurately when facing complex enterprise business and diversified user queries, and the large number of calls to the common language model leads to depletion of computing resource consumption and performance, and there is a risk of sensitive information leakage.

Method used

The multimodal input processing module is used for semantic understanding, combined with the preset content database and the enterprise knowledge base for matching search, if not satisfied, the pre-trained language model is called, the answer is generated through the digital human-driven engine, and the naked-eye 3D visualization is achieved using a transparent LCD screen.

Benefits of technology

It realizes efficient and accurate user response, reduces computing resource consumption, protects sensitive information, and provides an intelligent and natural immersive human-computer interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782466B_ABST
    Figure CN119782466B_ABST
Patent Text Reader

Abstract

The present application provides a digital human interaction method and system based on naked-eye 3D visualization, comprising: obtaining query content provided by a user in a multimodal manner, using natural language processing technology to perform semantic understanding and intent analysis on the query content to obtain user intent data; performing a matching search in a preset content database based on the user intent data, and if a highly relevant answer template or video clip exists, calling the preset content to provide feedback to the user; adopting corresponding digital human video visualization solutions based on different answer sources, wherein a digital human driving engine is used to match lip movements, expressions, and movements to synthesize a corresponding digital human video; and displaying the digital human video through a transparent LCD screen.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a digital human interaction method and system based on naked-eye 3D visualization. Background Art

[0002] When building intelligent digital human interactive visualization systems, ensuring efficient and accurate responses to user queries is a core challenge. When matching answers to different types of questions based on a pre-set template database, practical applications often encounter the following specific issues: Despite preparing a wealth of answer templates, due to the complexity of enterprise business operations and the diversity of user questions, some questions often lack fully matched answers within the pre-set template database. In this case, relying directly on pre-trained general language models to search massive amounts of information online to generate responses not only consumes significant computing resources but can also result in the generated answers being inconsistent with internal enterprise data, potentially leaking sensitive information. Even if a general language model captures a certain number of relevant answers, these answers may not meet the enterprise's specific needs, industry standards, or the latest internal regulations. For example, some answers may reference outdated product information, terms of service, or legal interpretations that are not applicable to the enterprise. This not only affects the accuracy of responses but also poses potential risks to the enterprise and its users. Furthermore, whenever an unmatched question is encountered, the system must invoke the pre-trained language model for large-scale information retrieval and text generation. This process consumes significant computing resources, especially in the case of high concurrent requests. This significantly increases server load, potentially degrading system performance and impacting user experience. At the same time, after finding the answer content, it also needs to be converted into an adaptive three-dimensional digital human video and presented in a way that imitates the expressions and movements of real humans, while providing users with an efficient and natural communication experience. Summary of the Invention

[0003] The present invention provides a digital human interaction method based on naked-eye 3D visualization, which mainly includes:

[0004] Obtaining query content provided by users in a multimodal manner, using natural language processing technology to perform semantic understanding and intent analysis on the query content to obtain user intent data;

[0005] Based on the user intent data, a matching search is performed in the preset content database. If there is a highly relevant answer template or video clip, the preset content is called to feedback to the user;

[0006] If the preset content database fails to meet the matching search requirements, the query content is passed to the enterprise knowledge base, and key information in the enterprise knowledge base is extracted using information retrieval technology. If the relevance of the information extracted from the enterprise knowledge base is lower than a preset threshold, the query content is submitted to the pre-trained language model;

[0007] After receiving the query content, the general large model uses a pre-trained large-scale language model and the massive amount of information resources on the Internet to input the question and perform reasoning based on the context information, and dynamically constructs the answer through text generation technology.

[0008] Based on the different sources of answers, corresponding digital human video visualization solutions are adopted, using a digital human driving engine to match lip movements, expressions, and movements to synthesize appropriate digital human videos;

[0009] The digital human video is displayed through a transparent LCD screen.

[0010] The present invention provides a digital human interaction system based on naked-eye 3D visualization, which mainly includes:

[0011] A multimodal input processing module is used to obtain the query content provided by the user in a multimodal manner, and use natural language processing technology to perform semantic understanding and intent analysis on the query content to obtain user intent data;

[0012] A preset content matching module is used to perform a matching search in a preset content database based on the user intent data, and if there is a highly relevant answer template or video clip, the preset content is called to feedback to the user;

[0013] an enterprise knowledge base retrieval module, configured to, if the preset content database fails to meet the matching search requirements, transmit the query content to the enterprise knowledge base, extract key information from the enterprise knowledge base using information retrieval technology, and submit the query content to a pre-trained language model if the relevance of the information extracted from the enterprise knowledge base is lower than a preset threshold;

[0014] The general large-scale model underlying module is used to receive the query content. It uses a pre-trained large-scale language model and massive information resources on the Internet to input the question and perform reasoning based on context information, and dynamically construct the answer through text generation technology.

[0015] The digital human video synthesis module is used to adopt corresponding digital human video visualization solutions based on different answer sources. It uses the digital human driving engine to match lip movements, expressions, and movements to synthesize appropriate digital human videos;

[0016] The naked-eye 3D display module is used to display the digital human video through a transparent LCD screen.

[0017] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:

[0018] The present invention discloses a digital human interaction method and system based on naked-eye 3D visualization. The method understands the user's intention through multimodal input and adopts a hierarchical retrieval strategy to obtain the answer content: first, match in the preset content library, and if it fails, search the enterprise knowledge base, and finally generate it by a large language model. For answers from different sources, the present invention adopts corresponding visualization solutions: directly call the preset video, or use text-to-speech and digital human driving engine to synthesize new videos. In the process of digital human behavior generation, real-person video data is modeled and imitated through machine learning algorithms, and reinforcement learning technology is used to optimize the interaction strategy. Finally, the present invention displays the digital human video through a special transparent LCD screen, uses the principle of optical refraction to create a naked-eye 3D effect, and dynamically adjusts the display parameters according to the environment. This method realizes an intelligent, natural, and immersive human-computer interaction experience, which can be widely used in customer service, guided tours and other scenarios to improve service efficiency and user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Flowchart of the digital human interaction method based on naked-eye 3D visualization of the present invention.

[0020] Figure 2 Schematic diagram of the digital human interaction system based on naked-eye 3D visualization of the present invention. DETAILED DESCRIPTION

[0021] The following will describe the technical solutions in the embodiments of the present invention in detail with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention.

[0022] like Figure 1 The digital human interaction method based on naked-eye 3D visualization in this embodiment may specifically include:

[0023] Step S101: obtain query content provided by the user in a multimodal manner, use natural language processing technology to perform semantic understanding and intent analysis on the query content, and obtain user intent data.

[0024] A voice signal and a facial image sequence are acquired through a voice acquisition device and a face acquisition device, and a phoneme sequence, a pitch curve, facial key point coordinate data, and expression feature parameters are extracted from the voice signal and the facial image sequence; the voice content is segmented and annotated according to the phoneme sequence and the pitch curve, and a feature fusion algorithm is used to perform multimodal feature fusion of the pre-trained word vectors corresponding to the segmentation results and the facial key point coordinate data to obtain a fused feature input matrix; semantic templates are matched in a preset dialogue knowledge base for the fused feature input matrix to obtain semantic understanding results and intention labels.

[0025] Specifically, a voice acquisition device acquires voice signals and extracts phoneme sequences, pitch curves, and speech rate parameters. A facial acquisition device simultaneously acquires facial image sequences and extracts facial landmark coordinate data and expression feature parameters. Based on the voice signal's phoneme sequence and pitch curve, a conditional random field algorithm is used to segment and annotate the speech content. Semantic feature vectors are generated based on the pre-trained word vectors corresponding to the segmentation results. A feature fusion algorithm is used to perform multimodal feature fusion of the semantic feature vectors with the facial landmark coordinate data, generating a fused feature input matrix. The fused feature input matrix is then matched against semantic templates stored in a pre-set conversation knowledge base to generate semantic understanding results and corresponding intent labels. Based on the intent labels, the corresponding digital human expression parameters and action parameter sequences are retrieved from a pre-set action library. The expression parameters and action parameter sequences are used to drive the 3D digital human model, rendering the digital human's expressions and body movements on a naked-eye 3D display. The speech synthesis module converts the semantic understanding results into speech waveform data, which is then combined with digital human body parameters for synchronized audio and video output. In multimodal human-computer interaction scenarios, activation is supported via voice wake-up or facial recognition. For voice wake-up, the system monitors the audio signal in the environment. Once the preset wake-up word is detected, the voice signal processing process begins. For face recognition wake-up, the camera continuously captures an image sequence. When the registered user's face is recognized, facial image processing and other related modules are initialized. The voice signal is first preprocessed to extract the phoneme sequence. A 25-millisecond analysis window is used for each acoustic frame, with a frame shift of 10 milliseconds. 13-dimensional fundamental frequency features and first-order difference features are extracted using Mel-frequency cepstral coefficients. The fundamental frequency contour features are also extracted from the pitch curve. Syllable duration normalization is used to extract speech rate parameters. A 68-point annotation scheme is used during facial image sequence acquisition to extract facial landmark coordinates. Facial deformation parameters are obtained by calculating the Euclidean distance between adjacent landmarks. Principal component analysis is used to reduce the dimensionality of facial expression changes and obtain expression feature parameters. The voice signal and facial image are time-aligned based on the synchronized clocks of the audio and video acquisition hardware. After word segmentation, speech content is processed to generate a word sequence. Each word corresponds to a 300-dimensional pre-trained word vector. A bidirectional long short-term memory network is used to extract features from the word vector sequence. A convolutional neural network is used to extract spatial features from facial key point coordinate data, resulting in a 128-dimensional expression feature vector. The speech feature vector and expression feature vector are concatenated and then fused through a fully connected layer to output a 256-dimensional fused feature vector. A pre-set conversation knowledge base stores structured semantic templates, each of which contains an intent category, a keyword set, and a corresponding response strategy. The fused feature vector is then matched against the semantic template for similarity using cosine similarity calculation with a similarity threshold of 0.85. The template with the highest similarity is selected as the semantic understanding result.The digital human's expression parameters include weight values for 17 basic action units, ranging from 0 to 1. The action parameter sequence uses a frequency of 25 frames per second to describe joint position information such as head rotation angle and torso posture. Based on naked-eye 3D display technology, parallax rendering is used to generate parallax images for the left and right eyes, with the parallax angle range controlled between 2 and 8 degrees to ensure viewing comfort. Speech synthesis uses a vocoder model to convert text into acoustic parameters and then generate a speech waveform with a 16kHz sampling rate. The synchronization error between the lip parameters and the speech is calculated using an auditory and visual perception model, and the error is controlled to within 40 milliseconds. The digital human's expression parameters are adjusted according to the fundamental frequency and loudness characteristics of the speech to achieve expressive multimodal interaction effects.

[0026] Step S102: performing a matching search in a preset content database based on the user intention data. If there is a highly relevant answer template or video clip, the preset content is called and fed back to the user.

[0027] Keywords and context features are extracted based on user intention data to generate a feature vector of the intention. A semantic search is performed on a preset content database through a vector space model to obtain a similarity score. The similarity scores are sorted in descending order. If the similarity is greater than a similarity threshold, the corresponding action data is extracted from the preset content database. A three-dimensional space transformation algorithm is used to align the coordinate system of the action data. The key frames of the action data are smoothly interpolated through Bezier curves to obtain a continuous action trajectory. The motion parameters between adjacent frames are calculated based on the continuous action trajectory. If the change in the motion parameters exceeds the motion parameter change threshold, transition action data is obtained from the action library, and a left-eye stereoscopic image is generated for the action data through a parallax rendering method.

[0028] Specifically, a deep learning algorithm extracts keywords and contextual features from user intent data to generate a 512-dimensional intent feature vector. A vector space model is then used to perform semantic retrieval of answer templates and video materials from a pre-set content database, generating a similarity score between the intent and the content. The search results are sorted in descending order based on the similarity score. A similarity threshold of 0.8 is set, and answer templates and video materials with similarities exceeding the threshold are selected as candidate content. The corresponding motion data is then extracted. A three-dimensional spatial transformation algorithm is used to align the extracted motion data, and keyframes are smoothly interpolated using Bezier curves to generate continuous motion trajectories. Based on the motion trajectory data, motion parameters between adjacent frames are calculated. If the change in motion parameters exceeds a pre-set threshold, transition motion data is retrieved from the motion library for supplementation. The supplemented motion sequence is then 3D reprojected and combined with parallax rendering to generate left-eye and right-eye stereo images. The digital human animation is then output on a naked-eye 3D display device. Based on the lip-sync keyframes in the motion sequence, the corresponding phoneme sequence is extracted, and a speech synthesizer generates a matching speech waveform for synchronized audio and video output. To understand user intent, deep learning features are extracted using a multi-layer perceptron architecture. The input layer receives a sequence of word vectors representing the user's query text, while the middle layer consists of 512 neurons. A nonlinear transformation is performed using the ReLU activation function, ultimately outputting a 512-dimensional intent feature vector. For the query "check today's weather," the extracted feature vector contains semantic information such as "check," "weather," and "time," with each dimension ranging from -1 to 1. Similarity is calculated using the cosine similarity method, calculating the cosine of the angle between the intent feature vector and the pre-set content feature vector. When the query intent achieves a similarity score of 0.92 with a "weather query" answer template from the pre-set content library, exceeding the set threshold of 0.8, the corresponding digital human video clip is selected. The video clip contains motion data such as the digital human's facial expressions and gestures as it explains the weather. This motion data is represented using keyframes, each containing the 3D coordinates of 67 human joints. Interpolation between adjacent keyframes is performed using a cubic Bezier curve, with the interpolation function P(t) = (1-t). 3 P0+3t(1-t) 2 P1+3t 2 (1-t)P2+t 3P3, where t ranges from 0 to 1, P0 and P3 are keyframe positions, i.e., the starting and ending points of the curve, and P1 and P2 are control points. When the displacement of joint points between adjacent keyframes exceeds a preset threshold of 30 cm, a transitional action is retrieved from the action library. This transitional action retrieval calculates similarity based on the Euclidean distance of the joint positions, selecting the transitional action sequence with the closest starting pose to supplement the initial pose and achieve a smooth transition. During 3D reprojection, binocular parallax is used to generate left and right eye images. The parallax value is related to the viewing distance and pupil distance. At a viewing distance of 2 meters, a baseline parallax value of 6.5 cm is set. Glasses-free 3D display uses a grating structure to project the left and right eye images separately to the viewer's eyes, achieving a stereoscopic visual effect. Speech synthesis uses phonemes as the basic unit, extracting phoneme labels corresponding to lip-sync keyframes from the action sequence. Each phoneme lasts approximately 50 to 200 milliseconds. By adjusting the phoneme duration parameter, speech output is synchronized with lip-sync changes, keeping the time deviation between phonemes and lip-sync within 40 milliseconds. Based on the duration information of the phonemes, the playback speed of the action sequence is adjusted accordingly to ensure the coordination and consistency of sound and action.

[0029] In step S103, if the pre-set content database fails to meet the matching search requirements, the query content is transferred to the enterprise knowledge base, and key information in the enterprise knowledge base is extracted using information retrieval technology. If the relevance of the information extracted from the enterprise knowledge base is lower than a pre-set threshold, the query content is submitted to the pre-trained language model.

[0030] A text classifier is used to divide the query content into topics, and a query semantic vector is obtained from the query content through a keyword extraction method; relevant documents are retrieved from the enterprise knowledge base based on the query semantic vector, and entity information is extracted from the relevant documents using a named entity recognition method, and the entity information is labeled and classified using a conditional random field algorithm; for the labeled and classified entity information, the cosine similarity between the entity information and the query semantic vector is calculated, and if the cosine similarity is higher than the cosine similarity threshold, a candidate knowledge matrix is generated; the entity information in the candidate knowledge matrix is sorted in descending order of similarity, and the entity information combination with the highest similarity is selected to generate an answer semantic vector, the answer text is constructed through a template filling method, and the digital human motion data is rendered using a three-dimensional reprojection method.

[0031] Specifically, a text classifier is used to thematically classify user queries. Keyword extraction methods are used to extract core entities and attributes from the query, constructing a query semantic vector. Matches are retrieved from the preset content layer. If the match similarity falls below a threshold of 0.7, a query at the knowledge base layer is triggered. Relevant documents are retrieved from the enterprise knowledge base based on the query semantic vector. Named entity recognition methods are used to extract entity information such as people, organizations, time, location, and numerical values contained in the documents. This entity information is annotated and classified using a conditional random field algorithm. For the annotated and classified entity information, cosine similarity is calculated between the annotated and classified entity information and the query semantic vector. Entities are then screened using a similarity threshold of 0.6 to generate a candidate knowledge matrix. Entities in the candidate knowledge matrix are sorted in descending order of similarity, and the top three entities with the highest similarity are selected and combined. Based on the selected entities, a response semantic vector is generated, and a template filling method is used to construct the response text. Based on the response text, the corresponding digital human expression and motion data sequence is retrieved from the action library. The digital human motion data is rendered using a 3D reprojection method and displayed on a naked-eye 3D display device. If the response similarity falls below a threshold of 0.5, a switch is made to a general large language model. When a user queries a company's product return and exchange policy, the text classifier first identifies the topic of the query content and extracts keywords through the word frequency-inverse document frequency method, including core words such as "return", "exchange", and "policy", to construct a 300-dimensional query semantic vector. The matching calculation at the preset content layer obtains a similarity of 0.65, which is lower than the 0.7 threshold, triggering a query at the knowledge base layer. The knowledge base retrieval locates the company's after-sales service documents, and named entity recognition extracts key entity information in the document, including the return period of "7 days", the exchange condition of "complete packaging", and the refund method of "original return". The conditional random field algorithm labels these entities, with the return period labeled as time type, the exchange condition as status type, and the refund method as method type. The cosine similarity between the labeled entity information and the query semantic vector is calculated using the following formula:

[0032] cos(θ) = (A·B) / (||A||·||B||), where A is the query vector and B is the entity information vector. The calculated similarity of return-related entity information is 0.85, and the similarity of exchange-related entity information is 0.78, both exceeding the 0.6 threshold. This generates a candidate knowledge matrix. The knowledge matrix is sorted in descending order of similarity, and the most relevant entity information sets (return time limit, exchange conditions, and refund method) are selected. Based on the selected entity information, the pre-set answer template "product within X days from receipt date, under Y conditions, and handled in Z manner" is used to populate the answer text, generating a standardized answer text. Based on the semantic category of the answer text, the corresponding digital human expression sequence and gesture action sequence are retrieved from the action library. The expression sequence includes basic action units such as eyebrow raising, gaze, and smiling, each containing the coordinate data of 17 facial feature points. The gesture action sequence includes directional gestures and explanatory gestures, each containing the 3D coordinate data of 21 hand joints. The digital human action data is 3D reprojected and displayed on a naked-eye 3D display device. The reprojection process uses the parallax rendering principle. At a 2-meter viewing distance, the parallax value for the left and right eyes is set to 6.5 cm to generate a binocular parallax image. If the similarity between the generated answer text and the query is 0.45, which is below the 0.5 threshold, the system automatically switches to a general large language model to generate a more general answer. The answer text is converted into a speech waveform using a speech synthesizer with a sampling rate of 16 kHz. A phoneme sequence is extracted to drive the digital population model. The action sequence is resampled based on the phoneme duration to achieve precise synchronization between sound and action, with a time error of less than 40 milliseconds.

[0033] In step S104, after receiving the query content, the general large model bottom layer uses a pre-trained large-scale language model, combined with massive information resources on the Internet, inputs the question and performs reasoning based on context information, and dynamically constructs an answer through text generation technology.

[0034] Part-of-speech tagging data is obtained based on the user query content, and the part-of-speech tagging data is obtained by the question understander performing word segmentation and part-of-speech tagging on the user query content; the part-of-speech tagging data is used to generate a question vector and a context vector, and a retrieval vector is formed by splicing the question vector and the context vector, and the retrieval vector is used to obtain information fragments; the subject-predicate-object grammatical structure is extracted from the information fragment, and the subject-predicate-object grammatical structure is used to construct knowledge triples based on syntactic dependencies. After constructing the knowledge triples, an initial answer knowledge graph is obtained; the initial answer knowledge graph is subjected to knowledge expansion to obtain an extended answer knowledge graph, and the extended answer knowledge graph is used to extract key facts, and the key facts are filled in through sentence templates to obtain the answer text.

[0035] Specifically, the question understander performs word segmentation and part-of-speech tagging on the user's query, extracting question keywords and contextual features to generate a question vector. A context vector is then constructed by extracting the last five conversation turns from the conversation transcript. A retrieval vector is generated by concatenating the question vector and context vector. Relevant information fragments are retrieved from the internet knowledge base and sorted by relevance calculation. The subject-verb-object grammatical structure is extracted from the filtered information fragments, and knowledge triples are constructed based on syntactic dependencies to generate an initial answer knowledge graph. A knowledge graph reasoning method is used to expand the initial answer knowledge graph, supplementing it with the complete knowledge nodes required for the answer. Key facts are extracted from the expanded answer knowledge graph, and a standardized answer text is generated using a sentence template filling method. Emotional tendencies and tone features are identified from the answer text and mapped to corresponding facial expression parameters in the digital human expression library. Based on the facial expression parameters, matching body movement data is retrieved from the action library to generate a complete digital human action sequence. The digital human action sequence is reprojected in 3D using parallax rendering and presented in real time on a naked-eye 3D display device. When a user asks "Applications of artificial intelligence in the medical field," the question understander first performs word segmentation, extracting keywords such as "artificial intelligence," "medical care," and "application." It also annotates "medical care" as a domain term and "application" as a subject term, constructing a 512-dimensional question vector. Contextual information, including basic AI concepts previously inquired about by the user, is extracted from the most recent five conversations, and these are concatenated to form a 768-dimensional search vector. Based on the search vector, the system searches for relevant content in the internet knowledge base, calculating relevance using cosine similarity with a threshold of 0.75. The top 10 relevant information segments are then extracted. Each information segment undergoes syntactic analysis, identifying grammatical components such as the subject "artificial intelligence," the predicate "application," and the object "assisted diagnosis." Knowledge triples such as "artificial intelligence - application to - medical imaging" and "deep learning - diagnosis - lung disease" are extracted. Knowledge graph reasoning is used to expand the initial answer, supplemented by association rules with specific data points such as "95% CT image recognition accuracy" and "60% reduction in diagnosis time." Using pre-set professional answer templates, medical scenarios are divided into three dimensions: diagnosis, treatment, and management. Standardized answer texts are generated according to the "domain-method-effect" structure. Sentiment analysis identifies positive and professional tone characteristics in the answers, which are then mapped to the "confident" and "focused" expression parameter groups in the digital human expression library. Facial expression parameters include weights for 17 action units, such as an eyebrow raise of 0.7 and a mouth corner raise of 0.5. Based on these expression parameters, matching gestures are retrieved from the action library to generate an action sequence containing both explanatory and indicative gestures. The action sequence is then 3D reprojected and displayed on a naked-eye 3D display. This reprojection utilizes binocular parallax, with a parallax of 6.5 cm between the left and right eyes at a viewing distance of 2 meters. While explaining medical applications, the digital human maintains a professional smile with gestures such as raising his hand and pointing.A speech synthesizer generates speech output at a 16kHz sampling rate, precisely controlling lip shape changes based on the phoneme sequence, with a time error of less than 40 milliseconds. Based on sound loudness and speech rate parameters, the digital human's movement amplitude and rhythm are dynamically adjusted to achieve natural expression with synchronized sound and image. Pre-set common sentence templates are used to construct a hierarchical answer structure. When elaborating on medical application cases, the system gradually expands on basic diagnosis, specialized treatment, and smart hospitals, establishing logical connections between each part through semantically cohesive vocabulary. Specific statistical data is inserted appropriately during the answer process to enhance persuasiveness and credibility, while maintaining the professionalism and accuracy of the language expression.

[0036] To meet the specific needs of different enterprises, lightweight private models are trained to replace or supplement general large models, generating response text while protecting data privacy.

[0037] A text segmenter is used to segment professional documents, and the conceptual relationship in the segmentation result is obtained through a dependency syntax analyzer, and a knowledge vector is generated according to the conceptual relationship; semantic understanding parameters are extracted from a general language model according to the knowledge vector, and the semantic understanding parameters are subjected to matrix pruning and compression to obtain a lightweight semantic understanding model; cosine similarity is calculated for the knowledge vector, and if the similarity is higher than a similarity threshold, the knowledge vectors are clustered to the same knowledge node, and an encrypted knowledge base is constructed through the knowledge nodes; a lightweight semantic understanding model is used to perform semantic analysis on the query content to obtain a query intent vector, and the knowledge content with the highest matching degree is retrieved from the encrypted knowledge base according to the query intent vector.

[0038] Specifically, based on the professional documents and historical Q&A records provided by the enterprise, a text segmenter is used to extract professional terminology. Dependency parsing is used to identify conceptual relationships and generate an initial knowledge vector. A knowledge distillation method is used to transfer semantic understanding capabilities from a general language model, pruning and compressing the parameter matrix to generate a lightweight semantic understander. A domain knowledge base is constructed based on the initial knowledge vector. Knowledge is clustered using cosine similarity, and relationships between knowledge nodes are established. A hashing algorithm is used to desensitize sensitive data in the knowledge base, generating an encrypted knowledge base. A lightweight semantic understander is used to understand user queries and extract query intent and key attributes. Based on query intent, relevant content is retrieved from the encrypted knowledge base and a list of candidate answers is generated by sorting by similarity. Based on the tone and intonation characteristics of the candidate answers, matching digital human expression parameters are retrieved from the expression library. Based on the expression parameters, a digital human action sequence is generated and displayed on a naked-eye 3D display using parallax rendering. In the financial enterprise scenario, when extracting professional terminology from internal documents, a tokenizer identifies core financial concepts such as "financial product," "yield rate," and "risk level." Dependency analysis then identifies conceptual relationships such as "financial product - inclusion - yield rate" and "risk level - impact - yield rate." Each concept is represented by a 300-dimensional vector, forming an initial set of knowledge vectors. During the knowledge distillation process, the original language model is compressed using structural pruning, reducing the 12-layer Transformer architecture to 4 layers. The number of attention heads is reduced from 12 to 4, the hidden layer dimension is reduced from 768 to 256, and the number of model parameters from 150 million to 15 million, while maintaining core semantic understanding capabilities. A financial domain knowledge base is constructed based on the knowledge vectors. K-means clustering is used to group knowledge points, with a K value of 20. Cosine similarity is used to calculate the strength of association between knowledge points. The similarity calculation formula is cos(θ) = (A·B) / (||A||·||B||), where A and B are knowledge vectors. A relationship is established when the similarity is greater than 0.8. Sensitive data such as customer information and transaction amounts are hashed and encrypted using the SHA-256 algorithm to generate fixed-length ciphertext. The original data "Zhang San - Purchase - Financial Product A - Amount 1 million" is converted into a hash value and stored in the encrypted knowledge base. When a user asks for "high-yield financial product recommendations," the lightweight semantic understander extracts the query intent "product recommendation" and the key attribute "high yield," searches the encrypted knowledge base for relevant content, and sorts it in descending order of similarity to generate a list of candidate answers. Based on the tone characteristics of the answer content, matching digital human expression parameters are retrieved from the expression library. In the professional explanation scenario, facial parameters such as eyebrow lift of 0.6, gaze focus of 0.8, and smile intensity of 0.4 are set to create a professional and credible expression effect. Combined with speech synthesis, speech output is generated at a 16kHz sampling rate, and lip movements are controlled through phoneme sequences.The digital human's motion sequences include explanatory and indicative gestures, each described by the 3D coordinate data of 21 hand joints. The motion data is then reprojected in 3D and presented on a glasses-free 3D display. A parallax value of 6.5 cm is set at a viewing distance of 2 meters to generate stereoscopic images for both eyes. Gestures are synchronized with the speech rhythm, with an error limit of 40 milliseconds, achieving a natural and smooth expression. Enterprise knowledge is integrated into the chain from document extraction, knowledge base construction, to answer generation. Knowledge distillation and parameter compression ensure lightweight and efficient private deployment. Encryption algorithms are also used to ensure data security, enabling professional and accurate knowledge transfer through the digital human.

[0039] Step S105: adopt corresponding digital human video visualization solutions according to different answer sources, wherein a digital human driving engine is used to match lip movements, expressions, and movements to synthesize a corresponding digital human video.

[0040] A content recognition method is used to classify and identify the answer content. If the recognition result comes from a preset content database or an enterprise knowledge base, the corresponding video material identifier is obtained from the preset content database or the enterprise knowledge base; video material data is extracted according to the video material identifier, and a digital human expression sequence and audio data are obtained from the video material data through a video framing method; a speech synthesis method is used to generate a standard audio waveform for the audio data, and a phoneme label sequence and lip shape parameters are extracted from the standard audio waveform through a phoneme recognizer; a standard lip shape library is retrieved based on the phoneme label sequence to obtain a lip shape parameter sequence, and digital human animation data is generated according to the lip shape parameter sequence and the expression sequence, and the digital human animation data is three-dimensionally reprojected through a parallax rendering method.

[0041] Specifically, a content recognition method is used to classify the source of the answer. The source type is determined by calculating text feature vectors. If the answer originates from a preset content library or enterprise knowledge base, the corresponding material identifier is retrieved from the video material library. Video material is extracted based on the material identifier, and video frame segmentation is used to obtain the digital human's expression sequence, action sequence, and corresponding audio data. For the text generated by the language model, a speech synthesis method is used to generate a standard audio waveform. A phoneme recognizer is used to extract phoneme labels and duration information. A standard lip shape library is retrieved based on the phoneme label sequence to generate a lip shape parameter sequence. Parameter smoothing is used to eliminate sudden changes between adjacent lip shapes. A speech emotion recognizer is used to extract intonation features from the audio waveform. Based on the intonation features, an expression library is retrieved to generate facial expression parameters. Based on the expression parameter sequence, the corresponding body movement data is retrieved from the action library. Motion trajectory smoothing is used to interpolate the movement data. A motion fusion method is used to integrate the lip shape parameters, expression parameters, and body movement data to generate a complete digital human animation sequence. A parallax rendering method is used to reproject the digital human animation in 3D, and a stereoscopic visual effect is output in real time on a naked-eye 3D display device. In an enterprise service scenario, when a user inquires about product return procedures, the content identifier first classifies the source of the answer using text feature vectors, extracting word frequency and syntactic features to construct a 512-dimensional feature vector. Cosine similarity is then used to determine whether the answer originates from a pre-set content library, with a similarity threshold of 0.85. A video clip with the identifier "return_policy_001" is retrieved from the pre-set content library and parsed at a frame rate of 25 frames per second. The coordinates of 17 facial landmarks in the digital human's expression sequence are extracted, while the action sequence contains the 3D coordinates of 21 human joints. The audio data is sampled at a 16kHz rate. If the answer originates from a language model, the speech synthesizer converts the text into an audio waveform. A phoneme segmentation algorithm is used to segment the speech into basic phoneme units based on acoustic characteristics. Each phoneme is annotated with its start time and duration, generating a phoneme label sequence such as " / h / / e / / l / / o / ." Based on the phoneme sequence, corresponding lip shape parameters are retrieved from a standard lip shape library. Each phoneme is mapped to specific lip shape parameters, including values such as mouth opening and lip angle extension. Cubic spline interpolation can be used to smooth adjacent lip shape parameters. Speech emotion recognition extracts acoustic features such as fundamental frequency and energy from the audio, calculates emotional dimension scores such as positivity and activation, and maps them to a facial expression parameter group. For example, a positive tone corresponds to parameter values such as 0.7 for eyebrow uplift and 0.5 for mouth corner uplift. The action library stores standardized gesture action sequences, each of which includes parameters such as arm joint angle and finger curvature. The action data is interpolated through motion trajectory smoothing, and the joint positions of the intermediate frames are calculated using Bezier curves to ensure natural and smooth transitions. During the action fusion process, parameters such as lip shape, expression, and gesture are unified into a three-dimensional coordinate space, and the final vertex position is calculated based on the weighted superposition method.Lip-shape weighting dynamically adjusts with speech loudness, expression weighting positively correlates with emotional intensity, and gesture weighting peaks at keyframes. Parallax rendering uses a 6.5cm baseline parallax at a 2-meter viewing distance, calculating the horizontal offset between the left and right eye images based on depth information. Using a grating structure, the left and right eye images are projected separately to the viewer's eyes, achieving a naked-eye 3D effect. Rendering resolution is 3840x2160, with a refresh rate of 90Hz, meeting the requirements of real-time interaction.

[0042] During the synthesis process of digital human videos, massive amounts of real-life video data are modeled and imitated for learning, while the digital human's response strategy and interaction method are optimized based on user feedback dialogue.

[0043] The human body key point coordinate sequence and facial feature point sequence in real-life video data are obtained to obtain a human body motion feature matrix; the human body motion feature matrix is converted using a joint angle mapping method to obtain a standardized motion feature vector; an action generator is constructed using a recurrent neural network for the standardized motion feature vector, and the action frame sequence is interpolated and completed; the completed action frame sequence is smoothed using a Bezier curve to obtain a smooth action sequence; a state transition matrix is constructed based on user satisfaction scores, the state transition probability is updated using a reinforcement learning method, and the optimal action response sequence is selected.

[0044] Specifically, real-life video data is processed using a deep neural network to extract 67 human key point coordinate sequences and 68 facial feature point sequences, generating a human motion feature matrix. Joint angle mapping is used to convert the human motion feature matrix to the digital human's skeletal structure, generating a standardized motion feature vector. Based on the standardized motion feature vector, a recurrent neural network is used to construct a motion generator, which interpolates and completes consecutive motion frames. For the completed motion frame sequence, Bezier curves are used to smooth the motion trajectory, generating a smooth motion sequence. An emotion recognizer analyzes the voice intonation and text sentiment in user feedback to generate a user satisfaction score. A reinforcement learning method is used to construct a motion state transition matrix, and state transition probabilities are updated based on the user satisfaction score. Based on the state transition matrix, the optimal motion response sequence is selected to generate the digital human's expressions and motion control parameters. Parallax rendering is used to reproject the digital human in 3D, and the interactive screen is presented in real time on a naked-eye 3D display device. In a customer service scenario, a deep neural network extracts features from recorded live receptionist videos, including the coordinates of 67 human body keypoints and 68 facial landmarks. Each keypoint is represented by three-dimensional (x, y, z) coordinates. Spatial features are extracted through convolutional layers, and dimensionality reduction is performed through pooling layers to output a 512-dimensional motion feature matrix. These feature points describe standard gestures, such as the upward corner of the mouth when the receptionist smiles or the rotation of the head when nodding. Joint angle mapping uses Euler angles to convert the spatial positions of human joints into a sequence of joint angles. For example, for an arm raise, the rotation angles of the shoulder, elbow, and wrist joints are calculated to generate a standardized angle vector. The mapped motion feature vector contains the Euler angle values for each joint, normalized to a range of -1 to 1. A recurrent neural network employs a long short-term memory (LSTM) network architecture. The input layer receives the standardized motion feature vector, and the hidden layer contains 256 neurons. Time series prediction is used to generate motion parameters for intermediate frames. For a 25-frame-per-second motion sequence, the network predicts transitions between adjacent keyframes. A cubic Bezier curve is used to smooth the motion trajectory, with the curve equation being P(t) = (1-t). 3 P0+3t(1-t) 2 P1+3t 2 (1-t)P2+t 3P3, where t ranges from 0 to 1, P0 and P3 are keyframe positions, and P1 and P2 are control points. For example, in an arm-waving motion, acceleration and deceleration effects can be achieved by adjusting the control point positions. When analyzing user feedback, the emotion recognizer extracts acoustic features such as the voice pitch curve and energy distribution, and simultaneously matches the text against a sentiment dictionary to calculate a positivity score. The satisfaction score is calculated on a 5-point scale, weighted by a weight of 0.6 for voice emotion features and 0.4 for text emotion features. The state transition matrix constructed using reinforcement learning has dimensions of 1000x1000, with each state corresponding to a set of standard action poses. Transition probabilities are updated based on user satisfaction scores, with higher-scoring state transitions receiving greater rewards. The state transition strategy is optimized using value iteration. During the generation of digital human expressions and motion parameters, the action state transition matrix and real-time feedback scores are combined to select the optimal action response sequence. For example, if user satisfaction feedback is high, the weight of smiling expressions is increased, and gesture amplitude is increased to enhance the positive interaction effect. Parallax rendering uses the principle of binocular parallax, setting a baseline parallax of 6.5 cm at a 2-meter viewing distance. It then combines motion depth information to calculate the horizontal offset of the left and right eye images. The rendering resolution is 3840x2160, and the left and right eye images are projected separately to the viewer's eyes via a grating structure, achieving a real-time stereoscopic display effect.

[0045] Step S106: displaying the digital human video through a transparent LCD screen.

[0046] The method acquires illumination data of the display environment through a light sensor array, and generates an environmental parameter matrix containing ambient light intensity values, light incident angle values, and color temperature distribution values based on the illumination data. The method processes the environmental parameter matrix using a light parameter mapping method to obtain a display parameter vector containing display brightness values, contrast values, and color saturation values. The original digital human image is subjected to brightness gain processing based on the display parameter vector to obtain a brightness-compensated image, and the brightness-compensated image is color-corrected based on the color temperature distribution values in the environmental parameter matrix to obtain a color-corrected image. The left and right eye disparity values of the color-corrected image are calculated, and a stereoscopic disparity image pair is generated using a light refraction compensation method. The stereoscopic disparity image pair is then subjected to viewing angle compensation based on the observation distance value to obtain the final display image.

[0047] Specifically, a light sensor array collects illumination data from the display environment, including ambient light intensity, incident angle, and color temperature distribution parameters, to generate an environmental parameter matrix. A light parameter mapping method is used to convert the environmental parameter matrix into a display parameter vector, containing display brightness, contrast, and color saturation values. Based on the display parameter vector, brightness gain processing is performed on the original digital human image to generate a brightness-compensated image. Color temperature correction is then performed on the brightness-compensated image, adjusting the image's color balance based on the ambient color temperature to generate a color-corrected image. In the digital human holographic cabin, this color correction is crucial to ensuring the realism and immersion of the virtual character, as holographic projection technology requires precise color reproduction to maintain 3D image quality. Based on the color-corrected image, left-eye and right-eye disparity values are calculated, and a light refraction compensation method is used to generate stereoscopic disparity image pairs. These stereoscopic disparity image pairs are used to create a realistic 3D effect, giving participants the feeling of being immersed in the digital human environment. The disparity images are separated using the grating structure of the transparent LCD screen to generate separate images for the left and right eyes. This technology is not only applicable to traditional 3D displays but is also particularly well-suited for transparent LCD screens within holographic cabins. It ensures clear stereoscopic images within a 45-degree viewing angle, allowing users to experience the stereoscopic effect without the need for special glasses. The parallax angle is calculated based on the viewing distance, and visual compensation is applied to the independent left and right eye images to generate the final display image. Real-time monitoring is used to continuously collect environmental parameter changes, dynamically updating the display parameter vector based on the magnitude of the change. In shopping mall displays, the light sensor array consists of 16 light intensity sensors and 4 color temperature sensors, evenly distributed around the transparent LCD screen or the digital human holographic cabin display. The light intensity sensor has a measurement range of 0-100,000 lux, and the color temperature sensor has a measurement range of 2,000-10,000 Kelvin, with a sampling frequency of 10 Hz. The sensor data is filtered to form a 16x3 environmental parameter matrix containing light intensity, incident angle, and color temperature values. The mapping of ambient parameters to display parameters uses a piecewise linear function. When the ambient light intensity ranges from 0 to 1000 lux, the display brightness is proportionally increased by 0.5, and when it ranges from 1000 to 10,000 lux, it is proportionally increased by 0.3. This helps optimize the visibility and contrast of digital humans in transparent LCD screens or holographic cabins. Contrast increases with increasing ambient light intensity. The mapping function is contrast = 1 + log(ambient_light / 1000), where contrast is the image contrast value, ambient_light is the ambient light intensity, and log is the natural logarithm function, which converts linear changes in ambient light intensity to changes on a logarithmic scale. Color saturation is inversely proportional to the ambient color temperature. In holographic cabins, to make the digital human appear more natural, color saturation requires fine-tuning based on the ambient color temperature. Brightness gain processing uses pixel-level multiplication, multiplying the RGB values of the original image by the brightness gain coefficient.In an indoor environment of 5000 lux, the brightness gain factor is 1.5, corresponding to a display brightness of 450 nits. Color temperature correction is achieved by adjusting the relative intensities of the three RGB channels, mapping the standard white point of 6500 Kelvin to the actual ambient color temperature. Parallax image generation utilizes the principle of binocular parallax, with a baseline parallax value of 6.5 cm for the left and right eyes at a standard viewing distance of 2 meters. The transparent LCD screen's grating structure consists of a strip lens array with a lens pitch of 0.5 mm, a focal length of 2 mm, and a refractive index of 1.5. Light is refracted through the lens array to form nine viewing angles, each differing by 8 degrees. As the viewing distance changes, the parallax angle adjusts inversely. When the viewing distance increases from 2 meters to 3 meters, the parallax angle decreases from 8 degrees to 5.3 degrees. Display parameters are updated every 1000 lux change in ambient light intensity. The display brightness can be adjusted from 150-600 nits, the contrast ratio from 1000:1 to 3000:1, and the color temperature from 5000-7500 Kelvin. Through real-time parameter adjustments, the digital human image maintains optimal display quality in various scenarios, including natural morning light, strong midday sunlight, and dim evening light. The grating structure ensures clear stereoscopic images within a 45-degree viewing angle, allowing the viewer to freely move around to experience a realistic 3D effect. The real-time monitoring system's response latency is kept within 100 milliseconds, making the adjustment of display parameters imperceptible to the viewer.

[0048] The present invention provides a digital human interaction system based on naked-eye 3D visualization, which mainly includes:

[0049] A multimodal input processing module is used to obtain the query content provided by the user in a multimodal manner, and use natural language processing technology to perform semantic understanding and intent analysis on the query content to obtain user intent data;

[0050] A preset content matching module is used to perform a matching search in a preset content database based on the user intent data, and if there is a highly relevant answer template or video clip, the preset content is called to feedback to the user;

[0051] an enterprise knowledge base retrieval module, configured to, if the preset content database fails to meet the matching search requirements, transmit the query content to the enterprise knowledge base, extract key information from the enterprise knowledge base using information retrieval technology, and submit the query content to a pre-trained language model if the relevance of the information extracted from the enterprise knowledge base is lower than a preset threshold;

[0052] The general large-scale model underlying module is used to receive the query content. It uses a pre-trained large-scale language model and massive information resources on the Internet to input the question and perform reasoning based on context information, and dynamically construct the answer through text generation technology.

[0053] The digital human video synthesis module is used to adopt corresponding digital human video visualization solutions based on different answer sources. It uses the digital human driving engine to match lip movements, expressions, and movements to synthesize appropriate digital human videos;

[0054] The naked-eye 3D display module is used to display the digital human video through a transparent LCD screen.

[0055] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the present application is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. A digital human interaction method based on naked-eye 3D visualization, characterized in that: The method comprises: Obtaining query content provided by users in a multimodal manner, using natural language processing technology to perform semantic understanding and intent analysis on the query content to obtain user intent data; Based on the user intent data, a matching search is performed in the preset content database. If there is a highly relevant answer template or video clip, the preset content is called to feedback to the user; If the preset content database fails to meet the matching search requirements, the query content is passed to the enterprise knowledge base, and key information in the enterprise knowledge base is extracted using information retrieval technology. If the relevance of the information extracted from the enterprise knowledge base is lower than a preset threshold, the query content is submitted to the pre-trained language model; After receiving the query content, the general large model uses a pre-trained large-scale language model and the massive amount of information resources on the Internet to input the question and perform reasoning based on the context information, and dynamically constructs the answer through text generation technology. Based on the different sources of answers, corresponding digital human video visualization solutions are adopted, using a digital human driving engine to match lip movements, expressions, and movements to synthesize appropriate digital human videos; Displaying the digital human video through a transparent LCD screen; The method of performing a matching search in a preset content database based on the user intention data and calling the preset content for feedback to the user if a highly relevant answer template or video clip exists includes: Extract keywords and context features based on user intent data, generate a feature vector of intent, and perform semantic retrieval on the preset content database using a vector space model to obtain a similarity score. sorting the similarity scores in descending order, and if the similarity is greater than a similarity threshold, extracting corresponding action data from the preset content database; A three-dimensional space transformation algorithm is used to align the coordinate system of the motion data, and a Bezier curve is used to smoothly interpolate the key frames of the motion data to obtain a continuous motion trajectory; Calculating motion parameters between adjacent frames based on the continuous motion trajectory, and if a change in the motion parameters exceeds a motion parameter change threshold, obtaining transition motion data from a motion library, and generating left-eye and right-eye stereo images for the motion data using a parallax rendering method; According to different sources of answers, corresponding digital human video visualization solutions are adopted, including: Using a content recognition method to classify and recognize the answer content, if the recognition result comes from a preset content database or an enterprise knowledge base, obtaining a corresponding video material identifier from the preset content database or the enterprise knowledge base; Extracting video material data according to the video material identifier, and obtaining a digital human expression sequence and audio data from the video material data by a video frame segmentation method; Generating a standard audio waveform using a speech synthesis method for the audio data, and extracting a phoneme label sequence and lip shape parameters from the standard audio waveform using a phoneme recognizer; A standard lip shape library is retrieved based on the phoneme tag sequence to obtain a lip shape parameter sequence, digital human animation data is generated according to the lip shape parameter sequence and expression sequence, and the digital human animation data is three-dimensionally reprojected using a parallax rendering method.

2. The method according to claim 1, characterized in that The method of obtaining the query content provided by the user in a multimodal manner and performing semantic understanding and intent analysis on the query content using natural language processing technology to obtain user intent data includes: Acquire a voice signal and a facial image sequence through a voice acquisition device and a face acquisition device, and extract a phoneme sequence, a pitch curve, facial key point coordinate data, and expression feature parameters from the voice signal and the facial image sequence; Segmenting and labeling the speech content according to the phoneme sequence and the pitch curve, and using a feature fusion algorithm to perform multimodal feature fusion on the pre-trained word vectors corresponding to the segmentation results and the facial key point coordinate data to obtain a fused feature input matrix; The fusion feature input matrix is matched with semantic templates in a preset dialogue knowledge base to obtain semantic understanding results and intent labels.

3. The method according to claim 1, characterized in that If the preset content database fails to meet the matching search requirements, the query content is transferred to the enterprise knowledge base, and key information in the enterprise knowledge base is extracted using information retrieval technology; If the relevance of the information extracted from the enterprise knowledge base is lower than a preset threshold, the query content is submitted to the pre-trained language model, including: Using a text classifier to classify the query content into topics, and obtaining a query semantic vector from the query content using a keyword extraction method; Retrieving relevant documents in the enterprise knowledge base according to the query semantic vector, extracting entity information from the relevant documents using a named entity recognition method, and labeling and classifying the entity information using a conditional random field algorithm; For the labeled and classified entity information, the cosine similarity between the entity information and the query semantic vector is calculated. If the cosine similarity is higher than the cosine similarity threshold, a candidate knowledge matrix is generated. The entity information in the candidate knowledge matrix is sorted in descending order of similarity, and the entity information combination with the highest similarity is selected to generate an answer semantic vector. The answer text is constructed through a template filling method, and the digital human motion data is rendered using a three-dimensional reprojection method.

4. The method according to claim 1, wherein After receiving the query content, the general large model uses a pre-trained large-scale language model, combined with massive information resources on the Internet, inputs the question and performs reasoning based on context information, and dynamically constructs an answer through text generation technology, including: Obtaining part-of-speech tagging data based on the user query content, wherein the part-of-speech tagging data is obtained by performing word segmentation and part-of-speech tagging on the user query content by the question understander; Generate a question vector and a context vector using the part-of-speech tagging data, and form a search vector by concatenating the question vector and the context vector, wherein the search vector is used to obtain information fragments; Extracting a subject-predicate-object grammatical structure from the information fragment, wherein the subject-predicate-object grammatical structure is used to construct a knowledge triple based on syntactic dependency, and after constructing the knowledge triple, obtaining an initial answer knowledge graph; Performing knowledge expansion on the initial answer knowledge graph to obtain an extended answer knowledge graph, wherein the extended answer knowledge graph is used to extract key facts, and the key facts are filled in through a sentence template to obtain an answer text; It also includes: training lightweight private models to replace or supplement general large models for the specific needs of different enterprises, and generating response texts while protecting data privacy.

5. The method according to claim 4, characterized in that The aforementioned lightweight private models are trained to meet the specific needs of different enterprises, replacing or supplementing the general large models to generate response text while protecting data privacy, including: Use a text word segmenter to segment professional documents, use a dependency parser to obtain the conceptual relationships in the word segmentation results, and generate knowledge vectors based on the conceptual relationships; Extracting semantic understanding parameters from a general language model according to the knowledge vector, and performing matrix pruning and compression on the semantic understanding parameters to obtain a lightweight semantic understanding model; Performing cosine similarity calculation on the knowledge vectors, and if the similarity is higher than a similarity threshold, clustering the knowledge vectors into the same knowledge node, and constructing an encrypted knowledge base through the knowledge nodes; A lightweight semantic understanding model is used to perform semantic analysis on the query content to obtain a query intention vector, and the knowledge content with the highest matching degree is retrieved in the encrypted knowledge base based on the query intention vector.

6. The method according to claim 1, characterized in that The method adopts corresponding digital human video visualization solutions according to different answer sources, wherein a digital human driving engine is used to match lip movements, expressions, and movements to synthesize a corresponding digital human video, and further includes: During the synthesis process of digital human videos, massive amounts of real-life video data are modeled and imitated for learning, while the digital human's response strategy and interaction method are optimized based on user feedback dialogue.

7. The method according to claim 6, characterized in that In the process of synthesizing digital human videos, modeling and imitation learning are performed on massive amounts of real-life video data, and the digital human's response strategy and interaction method are optimized based on user feedback dialogue, including: Obtain the human body key point coordinate sequence and facial feature point sequence in real-person video data to obtain the human body action feature matrix; According to the human body motion feature matrix, a joint angle mapping method is used to convert the matrix to obtain a standardized motion feature vector; Constructing an action generator based on the standardized action feature vector through a recurrent neural network, and performing interpolation and completion processing on the action frame sequence; The completed action frame sequence is smoothed using Bezier curve to obtain a smooth action sequence; A state transition matrix is constructed based on user satisfaction scores, and the state transition probability is updated through reinforcement learning method to select the optimal action response sequence.

8. The method according to claim 1, characterized in that The step of displaying the digital human video through a transparent LCD screen comprises: Acquire illumination data of the display environment through a light sensor array, and generate an environmental parameter matrix including ambient light intensity values, light incident angle values, and color temperature distribution values according to the illumination data; Processing the environmental parameter matrix using a lighting parameter mapping method to obtain a display parameter vector including display brightness values, contrast values, and color saturation values; Performing brightness gain processing on the original digital human image according to the display parameter vector to obtain a brightness compensated image, and performing color correction on the brightness compensated image according to the color temperature distribution value in the environmental parameter matrix to obtain a color corrected image; The left-eye and right-eye disparity values are calculated for the color-corrected image, a light refraction compensation method is used to generate a stereoscopic disparity image pair, and viewing angle compensation is performed on the stereoscopic disparity image pair according to the observation distance value to obtain a final display image.

9. A digital human interaction system based on naked-eye 3D visualization, characterized in that: The system comprises: A multimodal input processing module is used to obtain the query content provided by the user in a multimodal manner, and use natural language processing technology to perform semantic understanding and intent analysis on the query content to obtain user intent data; A preset content matching module is used to perform a matching search in a preset content database based on the user intent data, and if there is a highly relevant answer template or video clip, the preset content is called to feedback to the user; an enterprise knowledge base retrieval module, configured to, if the preset content database fails to meet the matching search requirements, transmit the query content to the enterprise knowledge base, extract key information from the enterprise knowledge base using information retrieval technology, and submit the query content to a pre-trained language model if the relevance of the information extracted from the enterprise knowledge base is lower than a preset threshold; The general large-scale model underlying module is used to receive the query content. It uses a pre-trained large-scale language model and massive information resources on the Internet to input the question and perform reasoning based on context information, and dynamically construct the answer through text generation technology. The digital human video synthesis module is used to adopt corresponding digital human video visualization solutions based on different answer sources. It uses the digital human driving engine to match lip movements, expressions, and movements to synthesize appropriate digital human videos; A naked-eye 3D display module, used to display the digital human video through a transparent LCD screen; The preset content matching module performs a matching search in the preset content database based on the user intention data. If there is a highly relevant answer template or video clip, the preset content is called to provide feedback to the user, including: Extract keywords and context features based on user intent data, generate a feature vector of intent, and perform semantic retrieval on the preset content database using a vector space model to obtain a similarity score. sorting the similarity scores in descending order, and if the similarity is greater than a similarity threshold, extracting corresponding action data from the preset content database; A three-dimensional space transformation algorithm is used to align the coordinate system of the motion data, and a Bezier curve is used to smoothly interpolate the key frames of the motion data to obtain a continuous motion trajectory; Calculating motion parameters between adjacent frames based on the continuous motion trajectory, and if a change in the motion parameters exceeds a motion parameter change threshold, obtaining transition motion data from a motion library, and generating left-eye and right-eye stereo images for the motion data using a parallax rendering method; The digital human video synthesis module adopts corresponding digital human video visualization solutions according to different answer sources, including: Using a content recognition method to classify and recognize the answer content, if the recognition result comes from a preset content database or an enterprise knowledge base, obtaining a corresponding video material identifier from the preset content database or the enterprise knowledge base; Extracting video material data according to the video material identifier, and obtaining a digital human expression sequence and audio data from the video material data by a video frame segmentation method; Generating a standard audio waveform using a speech synthesis method for the audio data, and extracting a phoneme label sequence and lip shape parameters from the standard audio waveform using a phoneme recognizer; A standard lip shape library is retrieved based on the phoneme tag sequence to obtain a lip shape parameter sequence, digital human animation data is generated according to the lip shape parameter sequence and expression sequence, and the digital human animation data is three-dimensionally reprojected using a parallax rendering method.

Citation Information

Patent Citations

  • Digital human construction method and system based on multi-modal large model

    CN118627519A

  • Semi-autonomous digital human posturing

    US20130197887A1