User operation guiding method based on digital human and intelligent terminal

By using digital people-based user operation guidance methods in real estate business management, the high labor cost problem caused by the complex process of real estate business processing is solved, the business process is simplified and intelligent, and user satisfaction is improved.

CN120010664APending Publication Date: 2025-05-16WUHAN DIGITAL MAP INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510090183.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-29
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The process of real estate business processing is complicated, which leads to customers who need to frequently consult customer service, which has a long consultation time and high labor cost for business processing.

Method used

Using a user operation guidance method based on digital people, by creating a guided digital person with multiple preset images, using intelligent service terminals to perform voice recognition, analysis and keyword extraction for users, match response texts and generate corresponding audio and lip-shaped timing data, to realize dynamic display of digital person's image and behavior.

Benefits of technology

Through digital human guidance, the process of users conducting business consultation and processing is simplified, quickly understand user needs and provide corresponding services, reduce the consultation time and cost of manual customer service, and improve user satisfaction and the intelligence of business processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010664A_ABST
    Figure CN120010664A_ABST
Patent Text Reader

Abstract

The invention discloses a user operation guiding method based on a digital person and an intelligent terminal, and the method comprises the steps: creating a navigation digital person with a plurality of preset images, and setting a first dynamic shape motion, a first dynamic facial expression and a navigation voice according to the images of the navigation digital person; obtaining voice information of a user, performing voice recognition and analysis on the voice information to obtain an analysis text, and extracting a service keyword according to the analysis text; matching a response text based on the business keyword, and obtaining response audio and lip shape time sequence data according to the response text; and setting a second dynamic facial expression and second dynamic shape data of the guide digital person based on the lip shape time sequence data, displaying the guide digital person through the intelligent service terminal, and playing response audio. According to the invention, the user is guided through the guide digital person, and the voice questions input by the user are answered by videos and audios, so that the intelligent degree of real estate business handling and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent handling of real estate business, and in particular to a user operation guidance method and intelligent terminal based on digital human. Background Art

[0002] Digital virtual people refer to virtual characters with digital appearance that exist in the non-physical world. With the development of digital technology, the metaverse, as an important manifestation of the digital economy, has ushered in new development opportunities. Digital virtual people, digital virtual items, etc. are the core concepts in the metaverse, important elements that link the real and virtual worlds, and the basis for realizing human-computer interaction. A complex with multiple human characteristics created by computer means such as computer graphics, graphics rendering, motion capture, deep learning, and speech synthesis. Digital virtual people must have three main characteristics: human appearance, human behavior, and human thoughts, that is, in terms of appearance, they must have specific appearance, gender, personality and other character characteristics, in terms of behavioral expression, they must have the ability to express themselves with language, facial expressions and body movements, and in terms of thought interaction, they must have the ability to recognize the external environment and communicate and interact with people.

[0003] With the promotion and pilot application of three-dimensional cadastral management and technical methods, people can use digital data of real estate registration, integrate cryptography, advanced data encryption, big data, artificial intelligence, blockchain, Web3.0 (decentralized Internet running on blockchain technology) and other technologies as a prerequisite to build a "metaverse + real estate registration" world and conduct multi-sensory information interaction. Guiding users to perform real estate registration operations through digital people has also become a future development trend.

[0004] Therefore, it is necessary to propose a user operation guidance method and intelligent terminal based on digital humans, which can guide users to perform real estate registration operations through digital humans, improve the intelligence level of business handling, and reduce the labor cost of business handling. Summary of the invention

[0005] In view of this, the present invention provides an online method and system for handling real estate registration business, which is used to solve the technical problem that due to the complicated process of real estate business handling, customers need to frequently consult customer service, resulting in long consultation time and high labor costs for business handling.

[0006] In order to achieve the above technical objectives, the present invention adopts the following technical solutions:

[0007] The present invention provides a user operation guidance method based on digital human, comprising:

[0008] Creating a tour guide digital human with multiple preset images, setting a first dynamic body movement, a first dynamic facial expression and a tour guide voice according to the image of the tour guide digital human, and displaying the tour guide digital human and playing the tour guide voice through the intelligent service terminal;

[0009] Obtain the user's voice information, perform voice recognition and analysis on the voice information, obtain parsed text, and extract business keywords based on the parsed text;

[0010] Match the response text based on the business keywords, and obtain the response audio and lip shape timing data according to the response text;

[0011] The second dynamic facial expression and the second dynamic body data of the guide digital human are set based on the lip shape time sequence data, the guide digital human is displayed through the intelligent service terminal, and the response audio is played.

[0012] Further, obtaining the response audio according to the response text includes:

[0013] A speech synthesis model is established based on the Transformer architecture, and the speech synthesis model is trained for the first time using the first training set, so that the model learns the correlation features between adjacent audio frames to output continuous audio corresponding to the response text;

[0014] The speech synthesis model is further trained for a second time using the second training set, so that the model learns to generate Chinese speech corresponding to the Chinese text, so as to output an audio intonation corresponding to the response text;

[0015] The speech synthesis model is trained for the third time using the third training set containing Chinese sentences and corresponding speech to ensure that the speech generated by the model is highly consistent with the input text;

[0016] After three trainings, a fully trained speech synthesis model is obtained, and the response text is input into the speech synthesis model to obtain the corresponding initial response audio;

[0017] A speech waveform generation model is established based on a multi-layer deep convolutional neural network to synthesize the speech timing waveform corresponding to the initial response audio to obtain the final output response audio.

[0018] Furthermore, the loss function of the first training is the spectral distortion between the output audio and the verification audio, and the model optimizer uses the AdaGrad optimizer;

[0019] After the first training, a text embedding layer is added to the speech synthesis model, the input text is represented as a vector embedding, and a text-speech joint attention mechanism is added so that the generated speech features can reflect the input text information; a convolutional layer is added to the output layer for smoothing, and a PostNet module is added to optimize the speech quality;

[0020] After the second training, a multi-head attention mechanism and position encoding are added to the speech synthesis model to optimize the intonation connection of the local context.

[0021] Further, obtaining lip shape timing data according to the response text includes:

[0022] Analyze the phoneme sequence corresponding to the response text, mark the emphasized syllables and stressed syllables in the phoneme sequence, and divide the phoneme sequence into common phonemes and emphasized phonemes;

[0023] Determine the lip shape data corresponding to each phoneme in the phoneme sequence; wherein the lip shape data includes the lip opening and closing degree and the lip shape state parameter, and the lip opening and closing degree corresponding to the key phoneme is greater than the lip shape data corresponding to the common phoneme;

[0024] The lip shape timing data is obtained according to the lip shape data of all the phonemes in the phoneme sequence.

[0025] Furthermore, setting the second dynamic facial expression and the second dynamic body data of the guide digital human based on the lip shape time sequence data includes:

[0026] Importing the lip shape timing data into the facial model of the tour guide digital human, associating the lip shape movement at each time point with the facial model of the tour guide digital human, and obtaining a facial expression timing diagram;

[0027] The facial expressions at adjacent time points are smoothed by an interpolation algorithm to generate lip animation to simulate the actual movement of the lips during pronunciation;

[0028] In the integrated animation editing environment, a facial expression timing diagram and body dynamic data corresponding to the lip animation are generated to obtain a second dynamic facial expression and a second dynamic body data.

[0029] Furthermore, the creation of a digital human guide with multiple preset images includes:

[0030] Obtain a video of the target person's image, analyze the video frame by frame, identify and locate multiple joint points of the target person's image, and obtain the two-dimensional coordinates of each joint point in the image pixel coordinate system;

[0031] The two-dimensional coordinates of each joint point are converted into three-dimensional coordinates in a preset coordinate system in combination with the structural characteristics of the human body, and a data file describing the bone hierarchy and joint angle and position information is obtained based on the three-dimensional coordinates of all joint points;

[0032] In the preset 3D modeling software, the data file is combined with the character model corresponding to the target character image to obtain a guide digital human;

[0033] The camera records the action and generates an action file, and the guide digital human corresponding to the target character image is driven according to the action file.

[0034] Furthermore, the step of creating a digital human guide having a plurality of preset images further includes:

[0035] Standardize the two-dimensional coordinates, map the standardized two-dimensional coordinates to three-dimensional space, and generate corresponding three-dimensional joint point coordinates;

[0036] Import the generated 3D joint point data into Blender to build a skeleton system, bind the skeleton system to the 3D model of the corresponding character, set keyframes for the skeleton system at different time points, and define character animation;

[0037] The video production interface of the 3D rendering software is called to export the animation video, and the preset audio is combined with the animation video for output.

[0038] The present invention also provides a user operation guidance intelligent terminal based on digital human, comprising:

[0039] The human-computer interaction unit is used to obtain the voice information input by the user, display the image and actions of the tour guide digital human, and output the tour guide audio and response audio;

[0040] A data analysis unit, configured to perform speech recognition and analysis on the voice information input by the user to obtain a parsed text, extract business keywords based on the parsed text, match the answer text based on the business keywords, and obtain the answer audio and lip shape timing data based on the answer text;

[0041] The digital human control unit is used to create a tour guide digital human with a variety of preset images, set a first dynamic body movement, a first dynamic facial expression and a tour guide voice according to the image of the tour guide digital human; and is also used to set a second dynamic facial expression and a second dynamic body data of the tour guide digital human based on the lip shape timing data.

[0042] Furthermore, the data analysis unit adopts a serverless cloud function cluster architecture;

[0043] The cloud function cluster architecture includes an object storage module, a parameter server module, a function cluster module and a configuration module.

[0044] The object storage module is used to save the input basic model and the model checkpoint sent by the parameter server;

[0045] The parameter service module is used to set various parameters of the model during model training;

[0046] The configuration module is used to obtain resource configuration requirements required for model operation;

[0047] The function cluster is used to call the model checkpoint saved on the parameter server to perform model training or use the model to predict results.

[0048] Furthermore, the function cluster adopts a Lambda function cluster.

[0049] Compared with the prior art, the user operation guidance method and system based on digital human provided by the present invention realizes the function of displaying different images and behaviors according to different scenes and user needs by creating a guide digital human with multiple preset images, making the intelligent guidance service of digital human more personalized and improving user satisfaction; by recognizing, analyzing and extracting keywords of the user's voice information, simplifying the business process of users' business consultation and handling, being able to quickly understand user needs and call the response text matching the consultation matter; matching the corresponding response text by business keywords, obtaining the response audio and lip shape timing data according to the response text, and generating the corresponding digital human body movement and expression movement, thereby combining voice, vision, hearing and other methods to interact with the user, and providing the user with the operation guidance service of real estate business. The method of this embodiment ensures a unified service standard through the preset digital human image, action and language, and explains the operation process of complex real estate business. Through this guidance method, it is convenient for users to solve common basic problems by themselves, avoid frequent help from manual customer service, save user time, and also reduce the workload of customer service personnel. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 A schematic diagram of a flow chart of an embodiment of a user operation guidance method based on a digital human provided by the present invention;

[0051] Figure 2 A schematic diagram of adjusting the skeleton model of a tour guide digital human provided by the present invention;

[0052] Figure 3 A schematic diagram of a flow chart of an embodiment of the guide digital human provided by the present invention responding to a user input voice;

[0053] Figure 4 A structural schematic diagram of an embodiment of a user operation guidance intelligent terminal based on a digital human provided by the present invention;

[0054] Figure 5 A schematic diagram of an embodiment of a cloud function cluster architecture provided by the present invention. DETAILED DESCRIPTION

[0055] The preferred embodiments of the present invention are described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not used to limit the scope of the present invention.

[0056] The present invention provides a user operation guidance method based on digital human and an intelligent terminal, which are respectively described below.

[0057] A specific embodiment of the present invention discloses a user operation guidance method based on digital human. Figure 1 is a flow chart of the user operation guidance method based on digital human, such as Figure 1 As shown, the method includes:

[0058] Step S101: creating a tour guide digital human with multiple preset images, setting a first dynamic body movement, a first dynamic facial expression and a tour guide voice according to the image of the tour guide digital human, and displaying the tour guide digital human and playing the tour guide voice through the intelligent service terminal;

[0059] Step S102: Acquire the user's voice information, perform voice recognition and analysis on the voice information to obtain a parsed text, and extract business keywords based on the parsed text;

[0060] Step S103: matching the response text based on the business keywords, and obtaining the response audio and lip timing data according to the response text;

[0061] Step S104: setting the second dynamic facial expression and the second dynamic body data of the tour guide digital human based on the lip shape timing data, displaying the tour guide digital human through the intelligent service terminal, and playing the response audio.

[0062] The method of this embodiment, by creating a digital human guide with multiple preset images, realizes the function of displaying different images and behaviors according to different scenarios and user needs, making the intelligent guidance service of the digital human more personalized and improving user satisfaction; by recognizing, analyzing and extracting keywords from the user's voice information, it simplifies the business process of the user's business consultation and handling, and can quickly understand the user's needs and call the response text matching the consultation matter; by matching the corresponding response text with business keywords, the response audio and lip shape timing data are obtained according to the response text, and the corresponding digital human body movements and facial expressions are generated, so as to interact with the user in a variety of ways such as voice, vision, and hearing, and provide the user with the operation guidance service of real estate business. The method of this embodiment ensures a unified service standard through the preset digital human image, action and language, and explains the operation process of complex real estate business. Through this guidance method, it is convenient for users to solve common basic problems by themselves, avoid frequent help from manual customer service, save users' time, and also reduce the workload of customer service personnel.

[0063] As a preferred embodiment, the creation of a digital human guide with multiple preset images includes:

[0064] Obtain a video of the target person's image, analyze the video frame by frame, identify and locate multiple joint points of the target person's image, and obtain the two-dimensional coordinates of each joint point in the image pixel coordinate system;

[0065] The two-dimensional coordinates of each joint point are converted into three-dimensional coordinates in a preset coordinate system in combination with the structural characteristics of the human body, and a data file describing the bone hierarchy and joint angle and position information is obtained based on the three-dimensional coordinates of all joint points;

[0066] In the preset 3D modeling software, the data file is combined with the character model corresponding to the target character image to obtain a guide digital human;

[0067] The camera records the action and generates an action file, and the guide digital human corresponding to the target character image is driven according to the action file.

[0068] In practical applications, we use blender 3D modeling software to manually model the virtual digital human image. This modeling method can make the virtual digital human model more plastic, more accurate in details, more flexible in adjusting the model, and also make the model more unique in the designer's style.

[0069] Character joint detection is the core of virtual digital human technology. Its purpose is to accurately locate and identify key parts of the human body, such as the head, hands and feet, from images or videos. In this regard, the YOLO series of technologies have attracted much attention. We selected YOLO3 and used it to identify character joints in the video, thereby extracting the two-dimensional joint coordinates. It adopts real-time target detection and recognition methods, and stands out for its high efficiency and speed.

[0070] We used the VideoTo3dPoseAndBvh algorithm to fuse the YOLO3 model to convert the two-dimensional coordinates of the joint points identified frame by frame by YOLO3 into three-dimensional coordinates. Due to the close connection between human bones, meridians and muscles, the three-dimensional coordinates in the specified coordinate system can be derived based on the coordinates of the joint points in the two-dimensional video. After obtaining the three-dimensional coordinates, coordinate conversion is performed to move the coordinates of the joint points to the center of the coordinate system, and then the three-dimensional coordinates of each frame are written into the Bvh skeletal animation file in the specified format.

[0071] Finally, the skeletal animation is combined with the character model in 3D modeling software such as Blender, the movements are recorded by the camera, and the action files are generated to achieve the effect of driving the virtual digital human.

[0072] As a preferred embodiment, the creation of a digital human guide having a plurality of preset images further comprises:

[0073] Standardize the two-dimensional coordinates, map the standardized two-dimensional coordinates to three-dimensional space, and generate corresponding three-dimensional joint point coordinates;

[0074] Import the generated 3D joint point data into Blender to build a skeleton system, bind the skeleton system to the 3D model of the corresponding character, set keyframes for the skeleton system at different time points, and define character animation;

[0075] The video production interface of the 3D rendering software is called to export the animation video, and the preset audio is combined with the animation video for output.

[0076] Specifically, in the process of skeletal animation generation, we use a preset generation algorithm, which is abbreviated as VideoTo3dPoseAndBvh algorithm. Its core steps are as follows:

[0077] Step 1: 2D to 3D conversion: To ensure data consistency and facilitate subsequent calculations, the extracted 2D coordinates are standardized so that they are evenly distributed in the range of -1 to 1. This step ensures the consistency of data from different video sources. Based on the standardized 2D coordinates, the VideoTo3dPoseAndBvh algorithm maps them to 3D space and generates the corresponding 3D joint point coordinates. These 3D coordinates include the position information of the x, y, and z axes, providing key data for the subsequent skeletal animation file generation.

[0078] Step 2: Coordinate axis conversion: In the process of compositing and rendering the video, the relevant interfaces of the blender SDK are called. Since different 3D rendering software such as blender may have different reference systems, optional coordinate axis conversion is added to the video synthesis script.

[0079] Step 3: Model skeleton binding: When driving the virtual digital human model, the model needs to be bound to the skeleton. However, the model size and the skeleton size may not match each other. Therefore, before binding, the skeleton size needs to be adjusted according to the model size ratio so that the model size and the skeleton size match each other. Figure 2 As shown, Figure 2 A schematic diagram showing adjustments to the model skeleton.

[0080] Step 4: Video synthesis: After the model skeleton is bound, the video synthesis script will load the prefabricated video environment, then call the video production interface of the 3D rendering software, export the animation video, and finally combine the audio and video, and export the video clip file to the specified area. After the last part of the video clip is completed, call the stream processing script to integrate multiple small video clip files into a complete video storage.

[0081] As a specific embodiment, the first dynamic body movement is waving the upper limbs to say hello or bow, etc. by default, the first dynamic facial expression is smiling and blinking at preset intervals by default, and the guide voice is a preset welcome word.

[0082] In some embodiments, after acquiring the user's voice information, voice recognition and analysis are performed on the voice information, which are achieved through ASR technology and NLP technology respectively.

[0083] The user's voice input is converted into text through ASR technology, which can adapt to different voice input environments and languages. The text output by ASR is processed through NLP technology, including word segmentation, entity recognition, keyword extraction, intent understanding, etc. These steps help understand the user's intent and extract key information. The parsed text is obtained, and business keywords are extracted based on the parsed text. Based on the results of NLP output, an intelligent question-and-answer system can be designed to match the corresponding answer text. This can be achieved through rule-based matching and machine learning models (such as classifiers or pre-trained language models). The above processes can be implemented by corresponding existing technologies and will not be explained in detail here.

[0084] As a preferred embodiment, the step of obtaining the response audio according to the response text includes:

[0085] A speech synthesis model is established based on the Transformer architecture, and the speech synthesis model is trained for the first time using the first training set, so that the model learns the correlation features between adjacent audio frames to output continuous audio corresponding to the response text;

[0086] The speech synthesis model is further trained for a second time using the second training set, so that the model learns to generate Chinese speech corresponding to the Chinese text, so as to output an audio intonation corresponding to the response text;

[0087] The speech synthesis model is trained for the third time using the third training set containing Chinese sentences and corresponding speech to ensure that the speech generated by the model is highly consistent with the input text;

[0088] After three trainings, a fully trained speech synthesis model is obtained, and the response text is input into the speech synthesis model to obtain the corresponding initial response audio;

[0089] A speech waveform generation model is established based on a multi-layer deep convolutional neural network to synthesize the speech timing waveform corresponding to the initial response audio to obtain the final output response audio.

[0090] In some embodiments, in the virtual digital human technology, the key to giving it the ability to speak is speech synthesis. It mainly includes the following processes:

[0091] (1) Text-to-speech conversion: In the pre-training stage, a large amount of speech data is used to learn the text-to-speech conversion rules, so as to predict the next audio frame of the input text. This fill-in-the-blank method is used to grasp the relationship between text and speech.

[0092] (2) Model retraining: In this stage, more Chinese training data is added to the model to make the model more robust to Chinese speech synthesis. Through a large amount of audio learning, the model learns how to map from Chinese text to the corresponding Chinese speech.

[0093] (3) Model fine-tuning: In this stage, annotated text data, such as sentences and their corresponding speech, are used to fine-tune the pre-trained model parameters to ensure that the output speech closely matches the input text.

[0094] (4) WaveNet vocoder output: In order to produce high-quality speech output, the method of combining WaveNet vocoder is chosen. WaveNet is an advanced vocoder that can produce natural speech waveforms.

[0095] Although there are many related algorithms, here we improve the Transformer architecture as the basic model. As a preferred embodiment, the loss function of the first training is the spectral distortion between the output audio and the verification audio, and the model optimizer uses the AdaGrad optimizer;

[0096] After the first training, a text embedding layer is added to the speech synthesis model, the input text is represented as a vector embedding, and a text-speech joint attention mechanism is added so that the generated speech features can reflect the input text information; a convolutional layer is added to the output layer for smoothing, and a PostNet module is added to optimize the speech quality;

[0097] After the second training, a multi-head attention mechanism and position encoding are added to the speech synthesis model to optimize the intonation connection of the local context.

[0098] As a specific embodiment, in actual application, the digital human can be further optimized to support multiple languages ​​or dialects, improve cultural adaptability, and enable the digital human to better serve user groups in different countries and regions.

[0099] In some embodiments, text can also be converted to speech by integrating a TTL engine.

[0100] As a preferred embodiment, the step of obtaining lip shape timing data according to the response text includes:

[0101] Analyze the phoneme sequence corresponding to the response text, mark the emphasized syllables and stressed syllables in the phoneme sequence, and divide the phoneme sequence into common phonemes and emphasized phonemes;

[0102] Determine the lip shape data corresponding to each phoneme in the phoneme sequence; wherein the lip shape data includes the lip opening and closing degree and the lip shape state parameter, and the lip opening and closing degree corresponding to the key phoneme is greater than the lip shape data corresponding to the common phoneme;

[0103] The lip shape timing data is obtained according to the lip shape data of all the phonemes in the phoneme sequence.

[0104] As a specific embodiment, for the emphasized phonemes and the common phonemes, the changes in their lip shape data over time are calculated respectively. Specific measurement indicators of lip shape can be used, such as the degree of opening and closing of the lips (opening), the front and back position of the lips (extended or retracted), the rounded or unrounded state of the lips, etc. The timing diagram of the expression is dynamically adjusted according to the real-time generated sound waveform or the stress and syllable strength of the speech text. When emphasizing syllables and stresses, the speed and amplitude of the expression changes can be increased, while subtle adjustments are made to weaker or unimportant parts to maintain overall fluidity and expressiveness.

[0105] In addition to displaying the digital human's body language and facial expressions, the response text must also display the text information synchronously to facilitate user understanding.

[0106] As a preferred embodiment, the second dynamic facial expression and the second dynamic body data of the guide digital human are set based on the lip shape time sequence data, including:

[0107] Importing the lip shape timing data into the facial model of the tour guide digital human, associating the lip shape movement at each time point with the facial model of the tour guide digital human, and obtaining a facial expression timing diagram;

[0108] The facial expressions at adjacent time points are smoothed by an interpolation algorithm to generate lip animation to simulate the actual movement of the lips during pronunciation;

[0109] In the integrated animation editing environment, a facial expression timing diagram and body dynamic data corresponding to the lip animation are generated to obtain a second dynamic facial expression and a second dynamic body data.

[0110] The second dynamic body data determines the duration of the body data according to the duration, and is combined with the lip shape animation to compile body movements that match the facial expressions. Figure 3 As shown, Figure 3 A flow chart showing the digital tour guide responding to user input voice.

[0111] like Figure 4 As shown, the embodiment of the present invention further provides a user operation guidance intelligent terminal 400 based on a digital human, comprising:

[0112] The human-computer interaction unit 401 is used to obtain the voice information input by the user, display the image and actions of the tour guide digital human, and output the tour guide audio and the response audio;

[0113] The data analysis unit 402 is used to perform speech recognition and analysis on the voice information input by the user to obtain a parsed text, extract business keywords based on the parsed text, match the response text based on the business keywords, and obtain the response audio and lip timing data based on the response text;

[0114] The digital human control unit 403 is used to create a tour guide digital human with a variety of preset images, set the first dynamic body movement, the first dynamic facial expression and the tour guide voice according to the image of the tour guide digital human; and set the second dynamic facial expression and the second dynamic body data of the tour guide digital human based on the lip shape timing data.

[0115] As a preferred embodiment, the data analysis unit adopts a serverless cloud function cluster architecture;

[0116] The cloud function cluster architecture includes an object storage module, a parameter server module, a function cluster module and a configuration module.

[0117] The object storage module is used to save the input basic model and the model checkpoint sent by the parameter server;

[0118] The parameter service module is used to set various parameters of the model during model training;

[0119] The configuration module is used to obtain resource configuration requirements required for model operation;

[0120] The function cluster is used to call the model checkpoint saved on the parameter server to perform model training or use the model to predict results.

[0121] As a preferred embodiment, the function cluster adopts a Lambda function cluster.

[0122] like Figure 5 As shown, Figure 5 The schematic diagram of the cloud function cluster architecture is shown. Since video data can be processed at the frame level concurrently, in order to meet the requirements of multi-task concurrency of different videos and frame-level concurrency within the video as much as possible, the system deployment uses container images to create cloud functions. This method can ensure the isolation of computing resources between different tasks and within the same task, as well as the rapid startup of cloud functions, so as to quickly increase computing power and complete multi-level concurrent processing.

[0123] During the deployment phase, operators first use the cloud resource monitoring service to perform load testing. By performing single-frame or small time segment load testing on the model in a single cloud function, the number of concurrent cloud functions required for a single model can be calculated based on real-time requirements. The purpose of this process is to obtain the processing power of a single cloud function for the algorithm through the cloud monitoring service, calculate the number of cloud functions that need to be pulled within the corresponding time based on the customer's expected processing speed, and finally configure the cloud function parameters in the cloud function warehouse. Based on the test data, the appropriate cloud function concurrency is obtained, and the system can configure the number of cloud function triggers to meet the operator's requirements for task processing speed.

[0124] During the use phase, operators only need to upload data streams to the cloud and trigger cloud functions through the stream processing service provided by the cloud vendor. The cloud function will process the data stream and return the processing results to the parameter server, and finally save the integrated data stream to the object storage service. This process realizes the automation of data stream processing and storage, greatly improving the efficiency and flexibility of data processing.

[0125] By properly configuring the concurrency of cloud functions, you can flexibly handle tasks of different scales and real-time requirements, dynamically adjust the number and concurrency of cloud functions according to actual needs, and avoid wasting resources.

[0126] In some embodiments, the human-computer interaction module also includes a user satisfaction input module, and the data analysis unit can also analyze service evaluations based on historical data, perform trend analysis and prediction on user questions, etc., to further improve the level of intelligence in business processing.

[0127] It should be noted that, in addition to interacting on the smart terminal, the method of the present invention can also interact through the user's mobile terminal through web pages, applets, etc.

[0128] The intelligent terminal provided in this embodiment solves the problems of traditional self-service equipment lacking human-computer interaction, poor functional operation friendliness, and lack of real-time business guidance. It provides a new method for connecting service personnel and customers, realizes the standardization of the service system, and truly supports the development and construction of unmanned business halls for real estate processing.

[0129] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by any technician familiar with the technical field within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A user operation guidance method based on digital human, characterized in that: include: Creating a tour guide digital human with multiple preset images, setting a first dynamic body movement, a first dynamic facial expression and a tour guide voice according to the image of the tour guide digital human, and displaying the tour guide digital human and playing the tour guide voice through the intelligent service terminal; Obtain the user's voice information, perform voice recognition and analysis on the voice information, obtain parsed text, and extract business keywords based on the parsed text; Match the response text based on the business keywords, and obtain the response audio and lip shape timing data according to the response text; The second dynamic facial expression and the second dynamic body data of the guide digital human are set based on the lip shape time sequence data, the guide digital human is displayed through the intelligent service terminal, and the response audio is played.

2. The user operation guidance method based on digital human according to claim 1, characterized in that: The step of obtaining the response audio according to the response text includes: A speech synthesis model is established based on the Transformer architecture, and the speech synthesis model is trained for the first time using the first training set, so that the model learns the correlation features between adjacent audio frames to output continuous audio corresponding to the response text; The speech synthesis model is further trained for a second time using the second training set, so that the model learns to generate Chinese speech corresponding to the Chinese text, so as to output an audio intonation corresponding to the response text; The speech synthesis model is trained for the third time using the third training set containing Chinese sentences and corresponding speech to ensure that the speech generated by the model is highly consistent with the input text; After three trainings, a fully trained speech synthesis model is obtained, and the response text is input into the speech synthesis model to obtain the corresponding initial response audio; A speech waveform generation model is established based on a multi-layer deep convolutional neural network to synthesize the speech timing waveform corresponding to the initial response audio to obtain the final output response audio.

3. The user operation guidance method based on digital human according to claim 2, characterized in that: The loss function of the first training is the spectral distortion between the output audio and the verification audio, and the model optimizer uses the AdaGrad optimizer; After the first training, a text embedding layer is added to the speech synthesis model, the input text is represented as a vector embedding, and a text-speech joint attention mechanism is added so that the generated speech features can reflect the input text information; a convolutional layer is added to the output layer for smoothing, and a PostNet module is added to optimize the speech quality; After the second training, a multi-head attention mechanism and position encoding are added to the speech synthesis model to optimize the intonation connection of the local context.

4. The user operation guidance method based on digital human according to claim 2, characterized in that: The step of obtaining lip shape timing data according to the response text includes: Analyze the phoneme sequence corresponding to the response text, mark the emphasized syllables and stressed syllables in the phoneme sequence, and divide the phoneme sequence into common phonemes and emphasized phonemes; Determine the lip shape data corresponding to each phoneme in the phoneme sequence; wherein the lip shape data includes the lip opening and closing degree and the lip shape state parameter, and the lip opening and closing degree corresponding to the key phoneme is greater than the lip shape data corresponding to the common phoneme; The lip shape timing data is obtained according to the lip shape data of all the phonemes in the phoneme sequence.

5. The user operation guidance method based on digital human according to claim 3, characterized in that: The second dynamic facial expression and the second dynamic body data of the guide digital human are set based on the lip shape time sequence data, including: Importing the lip shape timing data into the facial model of the tour guide digital human, associating the lip shape movement at each time point with the facial model of the tour guide digital human, and obtaining a facial expression timing diagram; The facial expressions at adjacent time points are smoothed by an interpolation algorithm to generate lip animation to simulate the actual movement of the lips during pronunciation; In the integrated animation editing environment, a facial expression timing diagram and body dynamic data corresponding to the lip animation are generated to obtain a second dynamic facial expression and a second dynamic body data.

6. The user operation guidance method based on digital human according to claim 1, characterized in that: The method of creating a tour guide digital person with multiple preset images includes: Obtain a video of the target person's image, analyze the video frame by frame, identify and locate multiple joint points of the target person's image, and obtain the two-dimensional coordinates of each joint point in the image pixel coordinate system; The two-dimensional coordinates of each joint point are converted into three-dimensional coordinates in a preset coordinate system in combination with the structural characteristics of the human body, and a data file describing the bone hierarchy and joint angle and position information is obtained based on the three-dimensional coordinates of all joint points; In the preset 3D modeling software, the data file is combined with the character model corresponding to the target character image to obtain a guide digital human; The camera records the action and generates an action file, and the guide digital human corresponding to the target character image is driven according to the action file.

7. The user operation guidance method based on digital human according to claim 6, characterized in that: The step of creating a digital tour guide having a plurality of preset images further comprises: Standardize the two-dimensional coordinates, map the standardized two-dimensional coordinates to three-dimensional space, and generate corresponding three-dimensional joint point coordinates; Import the generated 3D joint point data into Blender to build a skeleton system, bind the skeleton system to the 3D model of the corresponding character, set keyframes for the skeleton system at different time points, and define character animation; The video production interface of the 3D rendering software is called to export the animation video, and the preset audio is combined with the animation video for output.

8. A user operation guidance intelligent terminal based on digital human, characterized in that: include: The human-computer interaction unit is used to obtain the voice information input by the user, display the image and actions of the tour guide digital human, and output the tour guide audio and response audio; A data analysis unit, configured to perform speech recognition and analysis on the speech information input by the user to obtain a parsed text, extract business keywords based on the parsed text, match a response text based on the business keywords, and obtain a response audio and lip shape timing data based on the response text; The digital human control unit is used to create a tour guide digital human with a variety of preset images, set a first dynamic body movement, a first dynamic facial expression and a tour guide voice according to the image of the tour guide digital human; and is also used to set a second dynamic facial expression and a second dynamic body data of the tour guide digital human based on the lip shape timing data.

9. The user operation guidance intelligent terminal based on digital human according to claim 8, characterized in that: The data analysis unit adopts a serverless cloud function cluster architecture; The cloud function cluster architecture includes an object storage module, a parameter server module, a function cluster module and a configuration module; The object storage module is used to save the input basic model and the model checkpoint sent by the parameter server; The parameter service module is used to set various parameters of the model during model training; The configuration module is used to obtain resource configuration requirements required for model operation; The function cluster is used to call the model checkpoint saved on the parameter server to perform model training or use the model to predict results.

10. The user operation guidance intelligent terminal based on digital human according to claim 9, characterized in that: The function cluster adopts a Lambda function cluster.

Citation Information

Cited By

  • Method for quickly switching and multiplexing digital human images

    CN120298558A

  • Interaction method based on lightweight modular three-dimensional digital human

    CN121861175A