Digital human automatic explanation method and system based on large model, and storage medium
By using a large model for semantic recognition and style analysis in the digital human automatic explanation system, the explanation content and style are automatically determined, and the problem of single explanation content in the existing technology is solved, improving user experience and content diversity.
Patent Information
- Application Number
- CN202510126940.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-30
AI Technical Summary
The existing digital humans have a single explanation content during automatic explanation, resulting in a decrease in user experience.
By obtaining the user's speech, using pre-trained large models for semantic recognition, automatically determining the target explanation content, and conducting style analysis based on the user and explanation content, and controlling digital people to provide content explanation.
It improves the accuracy and content diversity of digital human automatic explanations, and improves the user experience.
Smart Images

Figure CN120066262A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital humans, and in particular, to a digital human automatic explanation method, system and storage medium based on a large model. Background Art
[0002] A digital human, also known as a virtual human or virtual digital human, is a virtual character that simulates human images, behaviors, and intelligence through computer technology. Digital humans have a highly realistic human appearance, including facial features, body shapes, clothing accessories, etc. During the use of digital humans, they can be used as explainers to explain content.
[0003] In the existing digital human automatic explanation process, generally, a fixed explanation content is played for explanation, resulting in a single explanation content and reducing the user experience. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a digital human automatic explanation method, system and storage medium based on a large model to solve the problem of single explanation content in the prior art.
[0005] The embodiments of the present invention are implemented as follows. A digital human automatic explanation method based on a large model, the method includes:
[0006] Obtain the explanation requirement voice sent by the user, and input the explanation requirement voice into the pre-trained large model for semantic recognition to obtain the explanation requirement semantics;
[0007] Determine the target explanation content according to the explanation requirement semantics, and perform style analysis on the user and the target explanation content to obtain the target explanation style;
[0008] Perform display analysis on the user and the explanation requirement semantics to obtain the target digital human, and control the target digital human to explain the content according to the target explanation content and the target explanation style.
[0009] Preferably, before inputting the explanation requirement voice into the pre-trained large model for semantic recognition, it further includes:
[0010] Obtain a voice sample, and input the voice sample into the large model for endpoint detection;
[0011] Segment the voice sample according to the endpoint detection result to obtain a segmented sample, and frame the segmented sample to obtain a framed sample;
[0012] Perform text prediction on the framed sample to obtain a text prediction sequence, and extract features from the text prediction sequence to obtain text sample features;
[0013] Perform semantic prediction on the text sample features to obtain sample predicted semantics, and determine the model loss according to the sample predicted semantics, the text prediction sequence, and the text sample features;
[0014] Update the parameters of the large model according to the model loss until the large model converges, and obtain the pre-trained large model.
[0015] Preferably, perform text prediction on the framed sample to obtain a text prediction sequence, including:
[0016] Extract features from the framed sample to obtain sample features, and perform feature mapping on the sample features to obtain an acoustic sample sequence;
[0017] Perform feature encoding on the sample features to obtain an acoustic sample vector, and perform word prediction on the acoustic sample vector to obtain a word sample sequence;
[0018] Determine the text prediction sequence according to the acoustic sample sequence and the word sample sequence.
[0019] Preferably, perform feature extraction on the text prediction sequence to obtain text sample features, including:
[0020] Perform word embedding processing and position encoding on the text prediction sequence to obtain a word embedding vector and a position encoding vector, and combine the word embedding vector and the position encoding vector to obtain a text prediction vector;
[0021] Perform multi-head attention mechanism calculation on the text prediction vector to obtain an attention vector, and splice the attention vectors to obtain a spliced vector;
[0022] Perform linear transformation on the spliced vector to obtain the text sample features.
[0023] Preferably, perform style analysis on the user and the target explanation content to obtain a target explanation style, including:
[0024] Perform semantic recognition on the target explanation content to obtain an explanation semantics, and determine an explanation emotion according to the explanation semantics;
[0025] Obtain the age, gender, and occupation of the user, and construct an explanation style matrix according to the age, gender, occupation, and explanation emotion;
[0026] Perform vector transformation on the explanation style matrix to obtain an explanation style vector, and calculate the similarity between the explanation style vector and a preset style vector to obtain a style similarity;
[0027] Determine the preset style vector corresponding to the maximum of the style similarities as the target style vector, and determine the explanation style corresponding to the target style vector as the target explanation style.
[0028] Preferably, controlling the target digital human to conduct content explanation according to the target explanation content and the target explanation style includes:
[0029] Obtain the speech pronunciation dictionary corresponding to the target explanation style, and perform phoneme conversion on the target explanation content to obtain a phoneme string;
[0030] Divide the phoneme string according to the content words in the target explanation content to obtain phoneme pairs, and match the phoneme pairs with the speech pronunciation dictionary to obtain pronunciation information;
[0031] Perform audio conversion on the target explanation content according to the pronunciation information to obtain the target explanation audio, and control the target digital human to conduct content explanation according to the target explanation audio.
[0032] Preferably, performing display analysis on the semantics of the user and the explanation requirement to obtain a target digital human includes:
[0033] Judge whether there is a digital human identifier in the semantics of the explanation requirement;
[0034] If there is the digital human identifier in the semantics of the explanation requirement, determine the target digital human according to the digital human identifier;
[0035] If there is no such digital human identifier in the semantics of the explanation requirement, obtain the historical explanation data of the user, and obtain the explanation duration of the historical digital human in the historical explanation data;
[0036] Determine the historical digital human corresponding to the maximum of the explanation durations as the target digital human.
[0037] Another object of the embodiments of the present invention is to provide a digital human automatic explanation system based on a large model, and the system includes:
[0038] A semantic recognition module, configured to obtain the explanation requirement speech sent by the user, and input the explanation requirement speech into a pre-trained large model for semantic recognition to obtain the semantics of the explanation requirement;
[0039] A style analysis module, configured to determine the target explanation content according to the semantics of the explanation requirement, and perform style analysis on the user and the target explanation content to obtain the target explanation style;
[0040] A content explanation module, configured to perform display analysis on the user and the semantics of the explanation requirement to obtain a target digital human, and control the target digital human to perform content explanation according to the target explanation content and the target explanation style.
[0041] Preferably, the semantic recognition module is further configured to:
[0042] Obtain a voice sample, and input the voice sample into the large model for endpoint detection;
[0043] Segment the voice sample according to the endpoint detection result to obtain a segmented sample, and frame the segmented sample to obtain a framed sample;
[0044] Perform text prediction on the framed sample to obtain a text prediction sequence, and extract features from the text prediction sequence to obtain text sample features;
[0045] Perform semantic prediction on the text sample features to obtain a sample prediction semantics, and determine a model loss according to the sample prediction semantics, the text prediction sequence, and the text sample features;
[0046] Update the parameters of the large model according to the model loss until the large model converges to obtain the pre-trained large model.
[0047] In an embodiment of the present invention, by inputting the voice of the explanation requirement into the pre-trained large model for semantic recognition, the semantic information of the voice of the explanation requirement can be effectively extracted. Based on the semantics of the explanation requirement, the target explanation content can be automatically determined. By analyzing the styles of the user and the target explanation content, the target explanation style suitable for the target explanation content can be effectively determined, improving the accuracy of the automatic explanation of the digital human. By performing display analysis on the user and the semantics of the explanation requirement, the target digital human suitable for the user's needs can be automatically confirmed. By controlling the target digital human to perform content explanation according to the target explanation content and the target explanation style, the content explanation operation can be effectively executed. In an embodiment of the present invention, by analyzing the user's explanation requirement, the target explanation content can be accurately determined, improving the diversity of content explanation and the user's experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a flowchart of a method for automatic digital human explanation based on a large model provided in the first embodiment of the present invention;
[0049] Figure 2 is a schematic structural diagram of a system for automatic digital human explanation based on a large model provided in the second embodiment of the present invention;
[0050] Figure 3 is a schematic diagram of automatic digital human explanation provided in the second embodiment of the present invention;
[0051] Figure 4 It is a schematic structural diagram of the terminal device provided by the third embodiment of the present invention. Detailed implementation manners
[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] In order to illustrate the technical solutions described in the present invention, the following will be described through specific embodiments.
[0054] Embodiment 1
[0055] Please refer to Figure 1 , which is a flowchart of the digital human automatic explanation method based on a large model provided by the first embodiment of the present invention. The digital human automatic explanation method based on a large model can be applied to any device or system. The digital human automatic explanation method based on a large model includes the following steps:
[0056] Step S10: Obtain the speech of the explanation requirement sent by the user, and input the speech of the explanation requirement into the pre-trained large model for semantic recognition to obtain the semantics of the explanation requirement;
[0057] Among them, by inputting the speech of the explanation requirement into the pre-trained large model for semantic recognition, the semantic information of the speech of the explanation requirement can be effectively extracted.
[0058] Optionally, before inputting the speech of the explanation requirement into the pre-trained large model for semantic recognition, it further includes:
[0059] Obtain a speech sample, and input the speech sample into the large model for endpoint detection; among them, by performing endpoint detection on the speech sample, the start and end positions of the speech signal in the speech sample can be determined.
[0060] Segment the speech sample according to the endpoint detection result to obtain a segmented sample, and frame the segmented sample to obtain a framed sample; among them, based on the endpoint detection result, the continuous audio stream is segmented into independent speech segments to obtain a segmented sample, which facilitates subsequent processing of each effective speech segment, reduces the processing amount of invalid data, and improves the recognition efficiency;
[0061] Perform text prediction on the framed sample to obtain a text prediction sequence, and extract features from the text prediction sequence to obtain text sample features; among them, by performing text prediction on the framed sample, the text prediction sequence corresponding to the framed sample can be effectively predicted;
[0062] Perform semantic prediction on the text sample features to obtain sample predicted semantics, and determine the model loss based on the sample predicted semantics, the text prediction sequence, and the text sample features;
[0063] Update the parameters of the large model according to the model loss until the large model converges to obtain the pre-trained large model; wherein, calculate the similarities between the sample predicted semantics and the standard semantics, between the text prediction sequence and the text standard sequence, and between the text sample features and the text standard features to obtain semantic similarity, sequence similarity, and feature similarity, and perform a weighted operation on the semantic similarity, sequence similarity, and feature similarity to obtain the model loss. During the weighted operation, the weighting coefficients of the semantic similarity, sequence similarity, and feature similarity can be set according to requirements.
[0064] Further, perform text prediction on the framed samples to obtain a text prediction sequence, including:
[0065] Extract features from the framed samples to obtain sample features, and perform feature mapping on the sample features to obtain an acoustic sample sequence; wherein, by performing feature mapping on the sample features, each acoustic probability distribution is output to obtain the acoustic sample sequence;
[0066] Perform feature encoding on the sample features to obtain an acoustic sample vector, perform word prediction on the acoustic sample vector to obtain a word sample sequence, and determine the text prediction sequence based on the acoustic sample sequence and the word sample sequence; wherein, by decoding the acoustic sample sequence and the word sample sequence, through a decoding algorithm such as the Viterbi algorithm, search for the most likely word sequence, that is, calculate the text sequence with the highest probability as the recognition result based on the acoustic information and the language probability information to obtain the text prediction sequence.
[0067] Even further, perform feature extraction on the text prediction sequence to obtain text sample features, including:
[0068] Perform word embedding processing and position encoding on the text prediction sequence to obtain a word embedding vector and a position encoding vector, and combine the word embedding vector and the position encoding vector to obtain a text prediction vector; wherein, map the words or sub-words in the text prediction sequence to a low-dimensional vector representation through a word embedding matrix to obtain the word embedding vector, each word corresponds to a vector with a fixed dimension, and the word embedding vector contains the initial representation of the semantic information of the word. By performing position encoding on the text prediction sequence, a unique vector representation can be effectively assigned to each position in the text prediction sequence;
[0069] Perform multi-head attention mechanism calculation on the text prediction vector to obtain an attention vector, splice the attention vectors to obtain a spliced vector, and perform a linear transformation on the spliced vector to obtain the text sample features; among them, pass the spliced vector through a linear transformation layer to obtain the final feature representation. The linear transformation layer can map the vector to a new dimensional space, further fuse and adjust the features, and obtain a feature vector that can better capture the text semantic information and context relationship.
[0070] Step S20, determine the target explanation content according to the semantics of the explanation requirement, and perform a style analysis on the user and the target explanation content to obtain the target explanation style;
[0071] Among them, perform entity recognition on the semantics of the explanation requirement, determine the explanation object and explanation purpose in the semantics of the explanation requirement based on the entity recognition result, and match the explanation object and explanation purpose with the explanation content database to obtain the target explanation content.
[0072] Optionally, performing a style analysis on the user and the target explanation content to obtain the target explanation style includes:
[0073] Perform semantic recognition on the target explanation content to obtain the explanation semantics, and determine the explanation emotion according to the explanation semantics; among them, perform vector conversion on the explanation semantics to obtain a semantic vector, calculate the emotion similarity between the semantic vector and the emotion vector corresponding to the preset emotion, and determine the preset emotion corresponding to the maximum emotion similarity as the explanation emotion;
[0074] Obtain the age, gender and occupation of the user, and construct an explanation style matrix according to the age, the gender, the occupation and the explanation emotion; among them, according to the preset mapping relationship, perform element mapping on the age, gender, occupation and explanation emotion respectively to obtain matrix element values, and construct a matrix according to the matrix element values corresponding to the age, gender, occupation and explanation emotion to obtain the explanation style matrix;
[0075] Perform vector conversion on the explanation style matrix to obtain an explanation style vector, and calculate the similarity between the explanation style vector and a preset style vector to obtain a style similarity;
[0076] Determine the preset style vector corresponding to the maximum style similarity as the target style vector, and determine the explanation style corresponding to the target style vector as the target explanation style; among them, the preset style vector can be set according to requirements, and the target explanation style includes style such as emotion style, pause style and tone color.
[0077] Step S30: Perform display analysis on the user and the semantics of the explanation requirement to obtain a target digital human, and control the target digital human to conduct content explanation according to the target explanation content and the target explanation style;
[0078] Among them, through performing display analysis on the user and the semantics of the explanation requirement, the target digital human suitable for the user's needs can be automatically confirmed. By controlling the target digital human to conduct content explanation through the target explanation content and the target explanation style, the content explanation operation can be effectively executed.
[0079] Optionally, controlling the target digital human to conduct content explanation according to the target explanation content and the target explanation style includes:
[0080] Obtain the speech pronunciation dictionary corresponding to the target explanation style, and perform phoneme conversion on the target explanation content to obtain a phoneme string; among them, match the target explanation style with the dictionary data table to obtain the speech pronunciation dictionary, and different target explanation styles and the corresponding speech pronunciation dictionaries are stored in the dictionary data table.
[0081] Divide the phoneme string according to the content words in the target explanation content to obtain phoneme pairs, and match the phoneme pairs with the speech pronunciation dictionary to obtain pronunciation information; among them, dividing the phoneme string according to the content words in the target explanation content can effectively divide the phoneme string based on the phrase relationship to obtain phoneme pairs, and by matching the phoneme pairs with the speech pronunciation dictionary, the pronunciation information corresponding to the phoneme pairs can be obtained.
[0082] Perform audio conversion on the target explanation content according to the pronunciation information to obtain the target explanation audio, and control the target digital human to conduct content explanation according to the target explanation audio; among them, combine the pronunciation information according to the order of the content words in the target explanation content to obtain the target explanation audio.
[0083] Furthermore, performing display analysis on the user and the semantics of the explanation requirement to obtain a target digital human includes:
[0084] Judge whether there is a digital human identifier in the semantics of the explanation requirement; among them, the digital human identifier is used to represent the number of the digital human.
[0085] If there is the digital human identifier in the semantics of the explanation requirement, determine the target digital human according to the digital human identifier; among them, if there is a digital human identifier in the semantics of the explanation requirement, it is determined that the user needs to use the specified digital human for explanation.
[0086] If the digital human identifier does not exist in the semantics of the explanation requirement, obtain the historical explanation data of the user, obtain the explanation duration of the historical digital human in the historical explanation data, and determine the historical digital human corresponding to the maximum explanation duration as the target digital human; among them, by obtaining the explanation duration of each historical digital human, the preference of the user for the intelligent agent can be effectively determined based on the explanation duration. Therefore, determining the historical digital human corresponding to the maximum explanation duration as the target digital human improves the user experience.
[0087] In this embodiment, by inputting the explanation requirement voice into the pre-trained large model for semantic recognition, the semantic information of the explanation requirement voice can be effectively extracted. Based on the semantics of the explanation requirement, the target explanation content can be automatically determined. By analyzing the style of the user and the target explanation content, the target explanation style suitable for the target explanation content can be effectively determined, improving the accuracy of the digital human's automatic explanation. In this embodiment, by analyzing the user's explanation requirement, the target explanation content can be accurately determined, improving the diversity of content explanation and the user experience.
[0088] Embodiment 2
[0089] Please refer to Figure 2 , which is a schematic structural diagram of the digital human automatic explanation system 100 based on a large model provided by the second embodiment of the present invention, including:
[0090] A semantic recognition module 10, configured to obtain the explanation requirement voice sent by the user, and input the explanation requirement voice into the pre-trained large model for semantic recognition to obtain the semantics of the explanation requirement.
[0091] Optionally, the semantic recognition module 10 is further configured to: obtain a voice sample, and input the voice sample into the large model for endpoint detection;
[0092] Segment the voice sample according to the endpoint detection result to obtain a segmented sample, and frame the segmented sample to obtain a framed sample;
[0093] Perform text prediction on the framed sample to obtain a text prediction sequence, and perform feature extraction on the text prediction sequence to obtain text sample features;
[0094] Perform semantic prediction on the text sample features to obtain sample prediction semantics, and determine the model loss according to the sample prediction semantics, the text prediction sequence, and the text sample features;
[0095] Update the parameters of the large model according to the model loss until the large model converges to obtain the pre-trained large model.
[0096] Further, the semantic recognition module 10 is further configured to: extract features from the framed samples to obtain sample features, and perform feature mapping on the sample features to obtain an acoustic sample sequence;
[0097] perform feature encoding on the sample features to obtain an acoustic sample vector, and perform word prediction on the acoustic sample vector to obtain a word sample sequence;
[0098] determine the text prediction sequence according to the acoustic sample sequence and the word sample sequence.
[0099] Even further, the semantic recognition module 10 is further configured to: perform word embedding processing and positional encoding on the text prediction sequence to obtain a word embedding vector and a positional encoding vector, and combine the word embedding vector and the positional encoding vector to obtain a text prediction vector;
[0100] perform multi-head attention mechanism calculation on the text prediction vector to obtain an attention vector, and splice the attention vectors to obtain a spliced vector;
[0101] perform a linear transformation on the spliced vector to obtain the text sample features.
[0102] The style analysis module 11 is configured to determine target explanation content according to the semantic of the explanation requirement, and perform style analysis on the user and the target explanation content to obtain a target explanation style.
[0103] Optionally, the style analysis module 11 is further configured to: perform semantic recognition on the target explanation content to obtain an explanation semantics, and determine an explanation emotion according to the explanation semantics;
[0104] obtain the age, gender and occupation of the user, and construct an explanation style matrix according to the age, the gender, the occupation and the explanation emotion;
[0105] perform vector transformation on the explanation style matrix to obtain an explanation style vector, and calculate the similarity between the explanation style vector and a preset style vector to obtain a style similarity;
[0106] determine the preset style vector corresponding to the maximum style similarity as the target style vector, and determine the explanation style corresponding to the target style vector as the target explanation style.
[0107] The content explanation module 12 is configured to perform display analysis on the user and the semantic of the explanation requirement to obtain a target digital human, and control the target digital human to perform content explanation according to the target explanation content and the target explanation style.
[0108] Optionally, the content explanation module 12 is further configured to: obtain the speech pronunciation dictionary corresponding to the target explanation style, and perform phoneme conversion on the target explanation content to obtain a phoneme string;
[0109] Divide the phoneme string according to the content words in the target explanation content to obtain phoneme pairs, and match the phoneme pairs with the speech pronunciation dictionary to obtain pronunciation information;
[0110] Perform audio conversion on the target explanation content according to the pronunciation information to obtain a target explanation audio, and control the target digital human to perform content explanation according to the target explanation audio.
[0111] Furthermore, the content explanation module 12 is further configured to: determine whether there is a digital human identifier in the semantics of the explanation requirement;
[0112] If there is the digital human identifier in the semantics of the explanation requirement, determine the target digital human according to the digital human identifier;
[0113] If there is no such digital human identifier in the semantics of the explanation requirement, obtain the historical explanation data of the user, and obtain the explanation duration of the historical digital human in the historical explanation data;
[0114] Determine the historical digital human corresponding to the maximum explanation duration as the target digital human.
[0115] Please refer to Figure 3 , there is a universal window on the large screen of the exhibition hall. The digital human automatic explanation system based on the large model can automatically realize functions such as voice scheduling, automatic explanation, company / product introduction video, ppt (ppt can be flipped through by voice scheduling), table opening, video opening, data summary, BI report viewing, etc.
[0116] In this embodiment, the speech input of the user is received, the speech signal is converted into text based on ASR (Automatic Speech Recognition), the semantics of the explanation requirement is determined based on the text, and the target explanation content is determined according to the semantics of the explanation requirement. The target explanation content is converted into a speech signal based on TTS (Text-to-Speech Synthesis) to obtain a target explanation audio. The data is stored and the knowledge base is managed based on KMS (Knowledge Management System) to support the decision-making and response of the AI large model. The business logic is analyzed and processed based on BI (Business Intelligence). Based on the AI large model, various modalities of information such as text and speech can be processed simultaneously to inject a soul into the digital human, so as to interact with the user more naturally and richly. Based on AGENT (Intelligent Agent) as the core of the AI large model, the work of each module is coordinated.
[0117] In this embodiment, through the deep learning ability of the large model, the digital human's understanding and processing ability of natural language are significantly improved, enabling the digital human to more accurately understand the user's intentions and emotions, and solving the problem of limited NLU ability in traditional technologies. The introduction of the large model enables the digital human to better process context information and integrate long-term and short-term memory, thus maintaining coherence in the conversation and solving the problem that digital humans in the prior art are difficult to maintain the conversation context. The digital human is allowed to make personalized adjustments based on the user's historical interaction data and preferences, and provide customized explanation content, solving the problems of single digital human service and lack of personalization in the prior art.
[0118] In this embodiment, by inputting the speech of the explanation requirement into the pre-trained large model for semantic recognition, the semantic information of the speech of the explanation requirement can be effectively extracted. Based on the semantics of the explanation requirement, the target explanation content can be automatically determined. By analyzing the styles of the user and the target explanation content, the target explanation style suitable for the target explanation content can be effectively determined, improving the accuracy of the digital human's automatic explanation. By performing a display analysis on the user and the semantics of the explanation requirement, the target digital human suitable for the user's needs can be automatically confirmed. By controlling the target digital human to perform content explanation through the target explanation content and the target explanation style, the content explanation operation can be effectively executed. In this embodiment, by analyzing the user's explanation requirement, the target explanation content can be accurately determined, improving the diversity of content explanation and the user's experience.
[0119] Embodiment III
[0120] Figure 4 is a structural block diagram of a terminal device 2 provided in the third embodiment of the present application. As Figure 4 shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for the digital human automatic explanation method based on the large model. When the processor 20 executes the computer program 22, the steps in each of the above embodiments of the digital human automatic explanation method based on the large model are implemented.
[0121] Exemplarily, the computer program 22 can be divided into one or more modules. The one or more modules are stored in the memory 21 and executed by the processor 20 to complete the present application. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.
[0122] The so-called processor 20 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0123] The memory 21 may be an internal storage unit of the terminal device 2, such as the hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the terminal device 2. Further, the memory 21 may also include both the internal storage unit and the external storage device of the terminal device 2. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or is to be output.
[0124] In addition, in each embodiment of the present application, the various functional modules may be integrated in one processing unit, may also exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0125] When an integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0126] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A digital human automatic explanation method based on a large model, characterized in that: The method comprises: Obtaining the explanation request voice sent by the user, and inputting the explanation request voice into the pre-trained large model for semantic recognition to obtain the explanation request semantics; Determine the target explanation content according to the explanation requirement semantics, and perform style analysis on the user and the target explanation content to obtain the target explanation style; The user and the explanation requirement semantics are displayed and analyzed to obtain a target digital person, and the target digital person is controlled to explain the content according to the target explanation content and the target explanation style.
2. The method for automatic interpretation of digital human based on a large model as claimed in claim 1, characterized in that: Before inputting the explanation demand voice into the pre-trained large model for semantic recognition, the method further includes: Acquire a speech sample, and input the speech sample into the large model for endpoint detection; Segmenting the speech sample according to the endpoint detection result to obtain segmented samples, and framing the segmented samples to obtain framed samples; Performing text prediction on the frame samples to obtain a text prediction sequence, and performing feature extraction on the text prediction sequence to obtain text sample features; Performing semantic prediction on the text sample features to obtain sample prediction semantics, and determining a model loss according to the sample prediction semantics, the text prediction sequence and the text sample features; The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
3. The large-model-based digital human automatic explanation method according to claim 2, characterized in that: Performing text prediction on the frame samples to obtain a text prediction sequence includes: Extracting features from the framed samples to obtain sample features, and performing feature mapping on the sample features to obtain an acoustic sample sequence; Performing feature encoding on the sample features to obtain an acoustic sample vector, and performing word prediction on the acoustic sample vector to obtain a word sample sequence; The text prediction sequence is determined according to the acoustic sample sequence and the word sample sequence.
4. The large-model-based digital human automatic explanation method according to claim 2, characterized in that: Performing feature extraction on the text prediction sequence to obtain text sample features includes: Performing word embedding processing and position encoding on the text prediction sequence to obtain a word embedding vector and a position encoding vector, and combining the word embedding vector and the position encoding vector to obtain a text prediction vector; Performing a multi-head attention mechanism calculation on the text prediction vector to obtain an attention vector, and concatenating the attention vectors to obtain a concatenated vector; Perform a linear transformation on the concatenated vector to obtain the text sample feature.
5. The method for automatic interpretation of digital human based on a large model as claimed in claim 1, characterized in that: Performing style analysis on the user and the target explanation content to obtain the target explanation style includes: Performing semantic recognition on the target explanation content to obtain explanation semantics, and determining explanation emotions according to the explanation semantics; Acquire the user's age, gender and occupation, and construct a presentation style matrix according to the age, gender, occupation and presentation emotion; Performing vector conversion on the explanation style matrix to obtain an explanation style vector, and performing similarity calculation between the explanation style vector and a preset style vector to obtain style similarity; The preset style vector corresponding to the maximum style similarity is determined as a target style vector, and the explanation style corresponding to the target style vector is determined as the target explanation style.
6. The large-model-based digital human automatic explanation method according to claim 1, characterized in that: Controlling the target digital human to explain the content according to the target explanation content and the target explanation style includes: Acquire a phonetic pronunciation dictionary corresponding to the target explanation style, and perform phoneme conversion on the target explanation content to obtain a phoneme string; Dividing the phoneme string according to the content vocabulary in the target explanation content to obtain phoneme pairs, and matching the phoneme pairs with the phonetic pronunciation dictionary to obtain pronunciation information; The target explanation content is converted into audio according to the pronunciation information to obtain the target explanation audio, and the target digital human is controlled to explain the content according to the target explanation audio.
7. The large-model-based digital human automatic explanation method according to claim 1, characterized in that: Performing display analysis on the user and the explanation requirement semantics to obtain a target digital human, including: Determining whether there is a digital human identifier in the semantics of the explanation requirement; If the digital human identifier exists in the explanation requirement semantics, determining the target digital human according to the digital human identifier; If the digital human identifier does not exist in the explanation requirement semantics, then obtaining the historical explanation data of the user, and obtaining the explanation time of the historical digital human in the historical explanation data; The historical digital person corresponding to the maximum explanation time is determined as the target digital person.
8. A digital human automatic explanation system based on a large model, characterized in that: The system comprises: The semantic recognition module is used to obtain the explanation demand voice sent by the user, and input the explanation demand voice into the pre-trained large model for semantic recognition to obtain the explanation demand semantics; A style analysis module, used to determine the target explanation content according to the explanation requirement semantics, and perform style analysis on the user and the target explanation content to obtain the target explanation style; The content explanation module is used to display and analyze the user and the explanation demand semantics to obtain the target digital person, and control the target digital person to explain the content according to the target explanation content and the target explanation style.
9. The large-model-based digital human automatic explanation system according to claim 8, characterized in that: The semantic recognition module is also used for: Acquire a speech sample, and input the speech sample into the large model for endpoint detection; Segmenting the speech sample according to the endpoint detection result to obtain segmented samples, and framing the segmented samples to obtain framed samples; Performing text prediction on the frame samples to obtain a text prediction sequence, and performing feature extraction on the text prediction sequence to obtain text sample features; Performing semantic prediction on the text sample features to obtain sample prediction semantics, and determining a model loss according to the sample prediction semantics, the text prediction sequence and the text sample features; The parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Interactive multi-mode artificial intelligence digital human automatic explanation method and system
CN120596655A
Digital human intelligent interaction method and system based on deep learning
CN120708611A
Digital human explanation and display control method based on voice recognition driving
CN121191518A