Digital human deployment method and system applied to terminal equipment

By deploying transcription modules, language models and voice modules locally on terminal devices, the problems of data leakage, delay and high cost in digital human deployment are solved, and a secure, low-latency personalized digital human interaction is achieved, suitable for small-scale enterprises and individual users.

CN120259492AInactive Publication Date: 2025-07-04SHENZHEN ZHUOYUE ZHIYUN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249090.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing digital human deployments rely on cloud computing, which poses the risk of data leakage, high latency, high operating costs, and is difficult to personalize, making it impossible to meet the needs of multiple application scenarios.

Method used

Transcription modules, language models and speech modules are deployed locally on terminal devices, and digital human interaction is achieved through acquisition, preprocessing, real-time transcription, optimization and animation rendering, using unified interfaces and model quantization technology.

Benefits of technology

Reduce the risk of privacy leakage, reduce delays, reduce operating costs, support personalized customization, adapt to a variety of application scenarios, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259492A_ABST
    Figure CN120259492A_ABST
Patent Text Reader

Abstract

The invention provides a digital human deployment method and system applied to terminal equipment. The method comprises the following steps: acquiring original voice data of a user through the terminal equipment; the original voice data is preprocessed and transmitted to the transcription module; original voice data are transcribed in real time through a transcription module; sending the original voice data and the transcription result to a language model; generating an optimization result according to the transcription result through a language model; converting an optimization result into a voice signal through a voice module and feeding back the voice signal to the terminal equipment; receiving a voice signal, and performing animation rendering on the digital human according to the voice signal; and driving the digital human to output voice signals. The terminal device, the transcription module, the language model and the voice module are deployed locally, so that user data does not need to be uploaded to the cloud, the privacy leakage risk is reduced, the system delay is reduced, and the interaction real-time performance is improved. Personalized customization can be carried out on the digital human according to user requirements and scenes, and various application scenes are met. And the operation cost can be reduced through a localized deployment mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a digital human deployment method and system applied to terminal devices. Background Art

[0002] The implementation of digital human technology relies on a variety of advanced technologies, including artificial intelligence algorithms, natural language processing, speech recognition, and computer vision, etc. The combination of these technologies enables digital humans to simulate human interaction behaviors and provide personalized services and information. However, the current digital human deployment mainly relies on cloud computing platforms and large-scale server clusters to process complex data and provide high-performance services. Although this centralized architecture has powerful computing and storage capabilities and can meet the needs of a large number of users, it still has a series of significant defects.

[0003] Firstly, non-localized systems usually rely on cloud computing, and user data needs to be uploaded to remote servers, which increases the risk of data leakage and privacy infringement. Secondly, since data needs to be transmitted over the Internet, transmission delays may lead to increased response times, thus affecting the real-time interaction experience, especially in application scenarios that require real-time feedback. In addition, building and maintaining large-scale systems usually requires high-cost hardware and network resources, and the operation and maintenance costs are relatively high, which makes it difficult for small enterprises and individual users to bear the relevant expenses. Moreover, existing technologies are difficult to be customized according to specific user needs, with poor flexibility and difficult to meet the needs of various application scenarios. When the system faces a large number of user requests or large-scale applications, due to performance bottlenecks, the service quality will decline and the user experience will be poor. Integrating multiple modules, such as speech recognition and LLM (Large Language Model), into a small-scale system will not only increase the difficulty of development and maintenance but also lead to a decline in system performance in order to achieve complex computing and data processing under limited resources. Summary of the Invention

[0004] The present invention aims to solve the problems in the deployment of digital humans in the above-mentioned prior art, such as low security, high latency, high operation and maintenance costs, inability to perform personalized customization, and difficulty in meeting the needs of various application scenarios, and provides a digital human deployment method and system applied to terminal devices.

[0005] The present invention provides a digital human deployment method applied to terminal devices, including the following steps:

[0006] Collect the original voice data of the user through a locally deployed terminal device;

[0007] Preprocess the original voice data and transmit it to a locally deployed transcription module; wherein, the transcription module includes at least one speech recognition model;

[0008] The original voice data is transcribed in real time by the transcription module to obtain a transcription result;

[0009] The original voice data and the transcription result are sent to a preset language model;

[0010] The language model generates an optimized result according to the transcription result;

[0011] The optimized result is converted into a voice signal by a preset voice module and fed back to the terminal device;

[0012] The voice signal is received, and an animation rendering of a preset digital human is performed according to the voice signal;

[0013] The digital human is driven to output the voice signal.

[0014] Further, in the step of transcribing the original voice data in real time by the transcription module to obtain a transcription result, it includes:

[0015] The voice recognition model is quantized to obtain the transcription result in digital form.

[0016] Further, the interfaces and communication protocols between the transcription module, the language model, and the voice module are the same.

[0017] Further, in the step of generating an optimized result by the language model according to the transcription result, it includes:

[0018] The intonation of the optimized result is recognized by the language model.

[0019] Further, in the step of receiving the voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes:

[0020] Corresponding lip shapes and facial expressions are generated according to the voice signal;

[0021] The tone of the digital human is adjusted according to the intonation; wherein, the tone includes speech rate, intonation, and filler words.

[0022] Further, in the step of receiving the voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes:

[0023] The background of the digital human is set.

[0024] Further, before the step of driving the digital human to output the voice signal, it includes:

[0025] Adjust the features of the digital human, where the features include but are not limited to facial features, hairstyles, and clothing;

[0026] Select the output language of the digital human.

[0027] Further, in the step of preprocessing the original voice data and transmitting it to a locally deployed transcription module, it includes:

[0028] The preprocessing includes language noise reduction, speech enhancement, speech segmentation, or data compression.

[0029] The present invention also provides a digital human deployment system applied to a terminal device, including a local device. The local device includes a terminal device, a transcription module, a language model, and a speech module. A microphone is provided in the terminal device for collecting the user's original voice data; the transcription module is used for real-time transcription of the original voice data to obtain a transcription result; the language model is used for generating an optimized result according to the transcription result; the speech module is used for converting the optimized result into a speech signal and feeding it back to the terminal device; a digital human module, the digital human module is used for receiving the speech signal and performing animation rendering according to the speech signal; an output module, the output module is used for driving the digital human to output the speech signal.

[0030] The present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and the processor executes the computer program to implement the steps in any one of the above methods.

[0031] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above methods are implemented.

[0032] The present invention provides a digital human deployment method and system applied to a terminal device, having the following beneficial effects:

[0033] By locally deploying a terminal device, a transcription module, a language model, and a speech module, the user's data does not need to be uploaded to the cloud, reducing the risk of privacy leakage and enhancing the user's control over the data. The method of real-time transcription and local processing of the original voice data can significantly reduce system latency, improve the real-time and naturalness of interaction, and optimize the user experience. By quantifying the speech recognition module, the energy consumption can be reduced, the service life of the terminal device can be extended, and it can adapt to resource-constrained environments. The digital human can be personalized according to user needs and scenarios, meeting various application scenarios. By means of local deployment, the dependence on cloud resources can be reduced, the hardware and operation costs can be lowered, and it can be widely applied to small-scale enterprises and individual users. Description of the Drawings

[0034] Figure 1 Schematic diagram of the method steps of a digital human deployment method applied to a terminal device in the present invention;

[0035] Figure 2 Block diagram of the structure of a digital human deployment system applied to a terminal device in the present invention;

[0036] Figure 3 Block diagram of the structure of a computer device of the present invention;

[0037] Figure 4 Schematic diagram of the working process of an embodiment of a digital human deployment method applied to a terminal device in the present invention.

[0038] Marking description: Local device 10, terminal device 101, transcription module 102, language model 103, voice module 104, digital human module 20, output module 30. Detailed implementation manners

[0039] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0040] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0041] Refer to the attached Figure 1 , a digital human deployment method applied to a terminal device 101 in an embodiment of the present invention, includes:

[0042] S1, collecting the original voice data of the user through the locally deployed terminal device 101;

[0043] S2, preprocessing the original voice data and transmitting it to the locally deployed transcription module 102; wherein, the transcription module 102 includes at least one speech recognition model;

[0044] S3, real-time transcription of the original voice data through the transcription module 102 to obtain a transcription result;

[0045] S4, sending the original voice data and the transcription result to a preset language model 103;

[0046] S5, generating an optimization result according to the transcription result through the language model 103;

[0047] S6. Convert the optimization result into a voice signal through a preset voice module 104 and feedback it to the terminal device 101;

[0048] S7. Receive the voice signal and perform animation rendering on a preset digital human according to the voice signal;

[0049] S8. Drive the digital human to output a voice signal.

[0050] In the above steps, first collect the user's original voice data through the locally deployed terminal device 101; among them, the terminal device 101 is a device that inputs programs and data to the computer or receives the computer output processing result by a communication facility. Such as mobile phones, computers, and monitors. As Figure 4 shown, in a specific embodiment, a microphone is provided in the terminal device 101, and the user's original voice data is collected through the microphone. Then preprocess the original voice data and transmit it to the locally deployed transcription module 102; in a specific embodiment, the preprocessing method is noise reduction, that is, perform noise reduction processing on the original voice data to remove background noise and other interference signals to ensure the clarity and quality of the original voice data. Further, preprocessing the original voice data can be to compress the original voice data to reduce the amount of data and improve the transmission speed.

[0051] Next, the preprocessed original speech data is sent to the transcription module 102 deployed locally. Among them, the transcription module 102 includes at least one speech recognition model. In a specific embodiment, the speech recognition model is ASR, Automatic Speech Recognition. By deploying it locally, the original speech data can be transcribed in real time to obtain a transcription result. In another specific embodiment, the speech recognition model is whisper, which can perform speech recognition. Specifically, the original speech data is converted into text information. Before transcription, the speech recognition model is quantized and optimized, and the precision of the model parameters is reduced from high precision to lower precision, such as reducing 32-bit floating-point numbers to 16-bit or 8-bit integers, thereby reducing the computational amount and memory occupancy of the model and improving the operating efficiency on resource-constrained devices. Then, the original speech data and the transcription result are sent to the language model 103 on the local device 10. In a specific embodiment, the language model 103 is an LLM, Large Language Model. The large language model 103 can perform semantic understanding on the text information in the transcription result and generate a reply, can handle complex natural language tasks, such as dialogue management, sentiment analysis, and context understanding, and obtain an optimized result. Then, the optimization result is converted into a speech signal through the speech module 104 and fed back to the terminal device 101. The speech module 104 can convert the optimization result into a speech signal. In a specific embodiment, the speech module 104 is TTS, Text To Speech, which can generate a smooth speech output according to the text content. Through the feedback mechanism, the optimization result is transmitted to the front end, thereby improving the response speed, ensuring data privacy, and having a low-latency response.

[0052] After the speech signal returns to the terminal device 101, the preset digital human is animatedly rendered, and the digital human is driven to output a speech signal. Among them, the digital human module 20 generates lip movements and facial expressions according to the speech signal and can realize real-time interaction with the user.

[0053] In one embodiment, in the step of transcribing the original speech data in real time by the transcription module 102 to obtain a transcription result, it includes:

[0054] Quantize the speech recognition model to obtain a digital-form transcription result.

[0055] In this embodiment, when transcribing the original speech data, the speech recognition model is quantized to obtain quantized text information. Among them, performing quantization processing on the speech recognition model can convert non-digital information into digital form for statistical analysis and calculation. Reduce the model complexity and computational requirements to adapt to the operation of small hardware devices. Effectively reduce memory occupancy and computational time, while reducing energy consumption and extending the service life of the device.

[0056] In one embodiment, the interfaces and communication protocols between the transcription module 102, the language model 103, and the speech module 104 are the same.

[0057] In this embodiment, the interfaces and communication protocols between the transcription module 102, the language model 103, and the speech module 104 are the same, which can simplify the integration process between the transcription module 102, the language model 103, and the speech module 104. By defining interface standards, seamless docking between modules is ensured, reducing the complexity and time cost of development, and enhancing flexibility and scalability.

[0058] In one embodiment, in the step of generating an optimized result by the language model 103 based on the transcription result, it includes:

[0059] The language model 103 identifies the intonation of the optimized result.

[0060] In this embodiment, the language model 103 is an LLM, a Large Language Model, which can identify the intonation in the optimized result by analyzing vocabulary, tone words, and sentence structures.

[0061] In one embodiment, in the step of receiving a voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes:

[0062] Generating corresponding lip shapes and facial expressions according to the voice signal;

[0063] Adjusting the tone of the digital human according to the intonation; where the tone includes speech rate, intonation, and tone words.

[0064] In this embodiment, when performing animation rendering on the digital human, the digital human module 20 generates natural lip shapes and expression actions according to the voice signal, and adjusts the tone of the digital human according to the intonation, where the tone includes speech rate, intonation, and tone words, so as to achieve real-time interaction with the user. The digital human module 20 stores multiple virtual characters and can create and manage virtual characters.

[0065] In one embodiment, in the step of receiving a voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes:

[0066] Setting the background of the digital human.

[0067] In this embodiment, the background of the digital human is adjusted, such as a park scene or a library scene. Among them, the setting method can also be to add or delete background elements. For example, in a park scene, the element of a bench is added. Through the custom method, the scene requirements of different users can be met, with high flexibility and applicability.

[0068] In one embodiment, before the step of driving the digital human to output a voice signal, it includes:

[0069] Adjust the characteristics of the digital human, where the characteristics include but are not limited to face, hairstyle, and clothing;

[0070] Select the output language of the digital human.

[0071] In this embodiment, before driving the digital human to output a voice signal, the characteristics of the digital human are adjusted, and the characteristics include but are not limited to face, hairstyle, and clothing. In a specific embodiment, the face of the digital human is set to "smile", the hairstyle of the digital human is adjusted to "curly hair", and the clothing of the digital human is adjusted to "shirt". Then select the output language of the digital human; among them, the output languages include Chinese, English, and Russian. In a specific embodiment, when the output language is adjusted to Chinese, the digital human outputs a Chinese voice signal. The personalized customization mechanism can support users to flexibly configure according to needs and scenarios to meet diverse requirements.

[0072] In one embodiment, in the step of preprocessing the original voice data and transmitting it to the locally deployed transcription module 102, it includes:

[0073] The preprocessing includes language noise reduction, speech enhancement, speech segmentation, or data compression.

[0074] In this embodiment, the preprocessing includes language noise reduction, speech enhancement, and speech segmentation. When the preprocessing method is noise reduction, the original voice data is subjected to noise reduction processing to remove background noise and other interference signals to ensure the clarity and quality of the original voice data. Further, preprocessing the original voice data can be to compress the original voice data to improve the transmission speed.

[0075] In summary, in specific implementation, first, the original voice data of the user is collected through the locally deployed terminal device 101; then the original voice data is preprocessed and transmitted to the locally deployed transcription module 102; next, the original voice data is transcribed in real time by the transcription module 102 to obtain a transcription result; during this period, the speech recognition model is quantized to obtain a transcription result in digital form. The quantization model can reduce energy consumption, extend the service life of the terminal device 101, and adapt to resource-constrained environments. Then the original voice data and the transcription result are sent to the preset language model 103; the language model 103 is used to generate an optimized result; during this period, the language model 103 recognizes the intonation of the optimized result. Next, the optimized result is converted into a voice signal through the voice module 104 and fed back to the terminal device 101; then the voice signal is received, and the preset digital human is animatedly rendered according to the voice signal; specifically, the corresponding lip shape and facial expression are generated according to the voice signal; the background of the digital human is set; the tone of the digital human is adjusted according to the intonation; the features of the digital human are adjusted, and the output language of the digital human is selected. Customizing the digital human in a personalized manner can be flexibly configured according to requirements and scenarios to meet diverse needs. Finally, the digital human is driven to output a voice signal. This application realizes the local deployment of the digital human on the resource-constrained terminal device 101 through the methods of local data storage, unified standardized interface, model quantization, and real-time inference. The local deployment method can ensure that user data does not need to be uploaded to the cloud, reduce the risk of privacy leakage, and enhance data security and privacy protection. The methods of real-time transcription and local data processing can significantly reduce system latency, improve the real-time and naturalness of interaction, and optimize the user experience. Reducing the dependence on cloud resources, lowering hardware and operation costs, and promoting wide application in small-scale enterprises and individual users.

[0076] Reference appendix Figure 2 A digital human deployment system applied to the terminal device 101, including a local device 10. The local device 10 includes a terminal device 101, a transcription module 102, a language model 103, and a voice module 104. A microphone is provided in the terminal device 101 for collecting the original voice data of the user; the transcription module 102 is used for transcribing the original voice data in real time to obtain a transcription result; the language model 103 is used for generating an optimized result according to the transcription result; the voice module 104 is used for converting the optimized result into a voice signal and feeding it back to the terminal device 101; a digital human module 20, the digital human module 20 is used for receiving the voice signal and performing animated rendering according to the voice signal; an output module 30, the output module 30 is used for driving the digital human to output a voice signal.

[0077] In this embodiment, the local device 10 includes a terminal device 101, a transcription module 102, a language model 103, and a voice module 104. The terminal device 101 collects the user's original voice data through a built-in microphone; the terminal device 101 can be a mobile phone or a computer. The transcription module 102 is used to transcribe the original voice data in real time to obtain a transcription result. The language model 103 is used to generate an optimized result based on the transcription result; the voice module 104 is used to convert the optimized result into a voice signal and feedback it to the terminal device 101. The digital human module 20 is used to receive the voice signal and perform animation rendering according to the voice signal; the output module 30 is used to drive the digital human to output the voice signal.

[0078] During this process, the original voice data and the transcription result are sent to the preset language model 103. The language model 103 analyzes the data and generates an optimized text response, and transmits the optimized result to the terminal device 101 through a feedback mechanism, thereby improving the overall performance and response speed of the system. The system realizes efficient voice interaction and digital human driving, while ensuring data privacy and low-latency response.

[0079] Reference appendix Figure 3 , In the embodiment of the present application, a computer device is further provided. The computer device can be a server, and its internal structure can be as Figure 3 shown. The computer device includes a processor, a memory, a network interface, and a database. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system, computer program, and database in the non-volatile storage medium. The database of the computer device is used to store data such as templates, tables, and preset fields. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a digital human deployment method applied to the terminal device 101, including the following steps:

[0080] Collect the user's original voice data through the locally deployed terminal device 101;

[0081] Preprocess the original voice data and transmit it to the locally deployed transcription module 102; wherein, the transcription module 102 includes at least one speech recognition model;

[0082] Transcribe the original voice data in real time through the transcription module 102 to obtain a transcription result;

[0083] Send the original voice data and the transcription result to the preset language model 103;

[0084] Generate an optimized result based on the transcription result by the language model 103;

[0085] Convert the optimized result into a voice signal through a preset voice module 104 and feedback it to the terminal device 101;

[0086] Receive the voice signal and perform animation rendering on a preset digital human according to the voice signal;

[0087] Drive the digital human to output a voice signal.

[0088] In one embodiment, in the step of obtaining the transcription result by real-time transcription of the original voice data through the transcription module 102, it includes:

[0089] Quantize the speech recognition model to obtain a digital-form transcription result.

[0090] In one embodiment, the interfaces and communication protocols between the transcription module 102, the language model 103, and the voice module 104 are the same.

[0091] In one embodiment, in the step of generating an optimized result based on the transcription result by the language model 103, it includes:

[0092] Identify the intonation of the optimized result through the language model 103.

[0093] In one embodiment, in the step of receiving the voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes:

[0094] Generate corresponding lip shapes and facial expressions according to the voice signal;

[0095] Adjust the tone of the digital human according to the intonation; wherein, the tone includes speech rate, intonation, and filler words.

[0096] In one embodiment, in the step of receiving the voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes:

[0097] Set the background of the digital human.

[0098] In one embodiment, before the step of driving the digital human to output a voice signal, it includes:

[0099] Adjust the features of the digital human, and the features include but are not limited to facial features, hairstyles, and clothing;

[0100] Select the output language of the digital human.

[0101] In one embodiment, in the step of preprocessing the original voice data and transmitting it to the locally deployed transcription module 102, it includes:

[0102] The preprocessing includes language noise reduction, speech enhancement, speech segmentation or data compression.

[0103] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied.

[0104] An embodiment of this application also provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a digital human deployment method applied to a terminal device 101, including the following steps:

[0105] Collect the original speech data of the user through the locally deployed terminal device 101;

[0106] Preprocess the original speech data and transmit it to the locally deployed transcription module 102; wherein, the transcription module 102 includes at least one speech recognition model;

[0107] Transcribe the original speech data in real time through the transcription module 102 to obtain a transcription result;

[0108] Send the original speech data and the transcription result to a preset language model 103;

[0109] Generate an optimization result according to the transcription result through the language model 103;

[0110] Convert the optimization result into a speech signal through a preset speech module 104 and feedback it to the terminal device 101;

[0111] Receive the speech signal and perform animation rendering on the preset digital human according to the speech signal;

[0112] Drive the digital human to output a speech signal.

[0113] In one embodiment, in the step of transcribing the original speech data in real time through the transcription module 102 to obtain a transcription result, it includes:

[0114] Quantify the speech recognition model to obtain a transcription result in digital form.

[0115] In one embodiment, the interfaces and communication protocols between the transcription module 102, the language model 103 and the speech module 104 are the same.

[0116] In one embodiment, in the step of generating an optimization result according to the transcription result through the language model 103, it includes:

[0117] Identify the intonation of the optimization result through the language model 103.

[0118] In one embodiment, in the step of receiving a voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes:

[0119] Generate corresponding lip shapes and facial expressions according to the voice signal;

[0120] Adjust the tone of the digital human according to the intonation; wherein, the tone includes speech rate, intonation, and filler words.

[0121] In one embodiment, in the step of receiving a voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes:

[0122] Set the background of the digital human.

[0123] In one embodiment, before the step of driving the digital human to output a voice signal, it includes:

[0124] Adjust the characteristics of the digital human, where the characteristics include but are not limited to facial features, hairstyle, and clothing;

[0125] Select the output language of the digital human.

[0126] In one embodiment, in the step of preprocessing the original voice data and transmitting it to the locally deployed transcription module 102, it includes:

[0127] The preprocessing includes language noise reduction, speech enhancement, speech segmentation, or data compression.

[0128] In summary, a digital human deployment method and system provided in the embodiments of the present application. By locally deploying the terminal device 101, transcription module 102, language model 103, and speech module 104, the user's data does not need to be uploaded to the cloud, reducing the risk of privacy leakage and enhancing the user's control over the data. The real-time transcription and localization processing of the original voice data can significantly reduce system latency, improve the real-time and naturalness of interaction, and optimize the user experience. By quantifying the speech recognition module, the energy consumption can be reduced, the service life of the terminal device 101 can be extended, and it can adapt to resource-constrained environments. The digital human can be customized according to user needs and scenarios to meet various application scenarios. The local deployment method can reduce the dependence on cloud resources, lower hardware and operation costs, and can be widely applied to small-scale enterprises and individual users.

[0129] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium provided in this application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be obtained in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0130] It should be noted that in this document, the terms "include", "comprise", or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, apparatus, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, apparatus, article, or method that includes such element.

[0131] The above are only the preferred embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of this application, or directly or indirectly applied in other related technical fields, shall equally be included in the patent protection scope of this application.

Claims

1. A digital human deployment method applied to a terminal device, characterized in that, It includes the following steps: Collect the original voice data of the user through a locally deployed terminal device; Preprocess the original voice data and transmit it to a locally deployed transcription module; wherein, the transcription module includes at least one speech recognition model; Real-time transcribe the original voice data through the transcription module to obtain a transcription result; Send the original voice data and the transcription result to a preset language model; Generate an optimized result according to the transcription result through the language model; Convert the optimized result into a voice signal through a preset voice module and feedback it to the terminal device; Receive the voice signal and perform animation rendering on a preset digital human according to the voice signal; Drive the digital human to output the voice signal.

2. The digital human deployment method applied to a terminal device according to claim 1, wherein, In the step of real-time transcribing the original voice data through the transcription module to obtain a transcription result, it includes: Quantify the speech recognition model to obtain the transcription result in digital form.

3. The digital human deployment method applied to a terminal device according to claim 1, wherein The interfaces and communication protocols between the transcription module, the language model, and the voice module are the same.

4. The digital human deployment method applied to a terminal device according to claim 1, wherein In the step of generating an optimized result according to the transcription result through the language model, it includes: Identify the intonation of the optimized result through the language model.

5. The digital human deployment method applied to a terminal device according to claim 4, wherein In the step of receiving the voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes: Generate corresponding lip shapes and facial expressions according to the voice signal; Adjust the tone of the digital human according to the intonation; wherein, the tone includes speech rate, intonation, and filler words.

6. The digital human deployment method applied to a terminal device according to claim 1, wherein In the step of receiving the voice signal and performing animation rendering on a preset digital human according to the voice signal, it includes: Set the background of the digital human.

7. The digital human deployment method applied to a terminal device according to claim 1, characterized in that, Before the step of driving the digital human to output the voice signal, it includes: Adjust the features of the digital human, and the features include but are not limited to facial features, hairstyles, and clothing; Select the output language of the digital human.

8. The digital human deployment method applied to a terminal device according to claim 1, wherein In the step of preprocessing the original voice data and transmitting it to a locally deployed transcription module, it includes: The preprocessing includes language noise reduction, speech enhancement, speech segmentation, or data compression.

9. A digital human deployment system applied to a terminal device, characterized in that, It includes a local device, and the local device includes a terminal device, a transcription module, a language model, and a voice module. A microphone is provided in the terminal device for collecting the original voice data of the user; the transcription module is used to real-time transcribe the original voice data to obtain a transcription result; the language model is used to generate an optimized result according to the transcription result; The voice module is used to convert the optimized result into a voice signal and feedback it to the terminal device; A digital human module, the digital human module is used to receive the voice signal and perform animation rendering according to the voice signal; an output module, the output module is used to drive the digital human to output the voice signal.

10. A computer device, comprising a memory and a processor, wherein a computer program is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps in any one of claims 1 to 8 of a digital human deployment method applied to a terminal device.

11. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of a digital human deployment method applied to a terminal device as claimed in any one of claims 1 to 8.