Software control method, system and equipment based on digital human system and medium

Through a large language model, analyzing user multi-modal input information, extracting keywords and operation information, and controlling third-party software, solving the problem of single functions of traditional digital human systems and improving application scenarios and user experience.

CN120045684AActive Publication Date: 2025-05-27HANGZHOU SHUJU CHAIN TECH CO LTD

Patent Information

Application Number
CN202510496416.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-27
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The traditional AI digital human system has a single function, mainly limited to information interaction in text or voice, and cannot effectively control third-party software, which limits its application scenarios and user interaction experience.

Method used

Through a large language model, the multi-modal input information input by the user is analyzed, the user needs are extracted, the keyword information and operation information are obtained, and the third-party software execution operation information is controlled based on the third-party software information, so as to realize the user's operation of the digital human system to control the third-party software.

Benefits of technology

It improves the convenience of using digital human system, enriches the application scenarios of digital human system, enhances the user's interactive experience, and realizes human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045684A_ABST
    Figure CN120045684A_ABST
Patent Text Reader

Abstract

The invention provides a software control method, system and equipment based on a digital human system and a medium, and relates to the technical field of artificial intelligence, and the method comprises the following steps: obtaining multi-modal input information input by a user; sending the multi-modal input information to a large language model, and obtaining keyword information and operation information fed back by the large language model; according to third-party software information in the keyword information, third-party software is controlled to execute the operation information, and output text information is generated; inputting the output text information to a voice synthesis module to obtain output voice; and based on the output voice, controlling a digital person simulated by the digital person system to play the output voice. According to the technical scheme provided by the embodiment of the invention, the digital human system can control the third-party software, the use convenience of the digital human is improved, the application scene of the digital human is enriched, and the interaction experience of the user is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a software control method, system, device and medium based on a digital human system. Background Art

[0002] With the development of artificial intelligence and virtual avatar technology, digital humans have gradually become an important part of the intelligent interaction field. Digital humans can achieve dialogue and communication with users through natural language processing technology, and provide services such as information query and entertainment interaction for users. Currently, digital humans have been applied in multiple scenarios, such as intelligent customer service, virtual assistants, online education, etc.

[0003] At the present stage, traditional AI digital human systems mainly rely on pre-trained language models to understand users' questions and make responses, and their core concern is to provide users with a high-quality dialogue experience. However, this mode also leads to the relatively single functions of such systems, mainly limited to information interaction methods in text or voice forms. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a software control method based on a digital human system, which, through the way of human-computer interaction, uses the digital human system to control third-party software to perform relevant operations, improves the usability of the digital human system, enriches the application scenarios of digital humans, and enhances the interactive experience of users.

[0005] In a first aspect, an embodiment of the present invention provides a software control method based on a digital human system, which is applied to a digital human system. The method includes: Obtain multimodal input information input by a user; Send the multimodal input information to a large language model, and obtain keyword information and operation information fed back by the large language model; According to the third-party software information in the keyword information, control the third-party software to execute the operation information, and generate output text information; Input the output text information into a speech synthesis module to obtain output speech; Based on the output speech, control the digital human simulated by the digital human system to play the output speech.

[0006] In a preferred embodiment of the present invention, the above multimodal input information includes voice information; the sending the multimodal input information to a large language model, and obtaining the keyword information and operation information fed back by the large language model includes: Input the voice information into a text recognition model to obtain input text information corresponding to the multimodal input information; Send the input text information to the large language model to obtain the keyword information and operation information feedback by the large language model.

[0007] In a preferred embodiment of the present invention, before obtaining the multi-modal input information input by the user, it includes: Receive the wake-up voice input by the user; Extract features from the wake-up voice to identify the user identity information corresponding to the wake-up voice; When the user identity information is successfully obtained, wake up the digital human system.

[0008] In a preferred embodiment of the present invention, the above-mentioned waking up the digital human system includes: Query the wake-up permission corresponding to the user identity information according to the black and white list; When the wake-up permission is to allow waking up, wake up the digital human system.

[0009] In a preferred embodiment of the present invention, after sending the multi-modal input information to the large language model to obtain the keyword information and operation information feedback by the large language model, it includes: According to the query information in the keyword information, obtain at least one document corresponding to the keyword information from the knowledge base; Extract and classify the content in each document according to the operation information to obtain the output text information.

[0010] In a preferred embodiment of the present invention, the above-mentioned controlling the digital human simulated by the digital human system to play the output voice based on the output voice includes: Input the output voice into the emotion simulation model to obtain the emotion corresponding to the output voice; According to the emotion, control the digital human simulated by the digital human system to play the output voice.

[0011] In a second aspect, an embodiment of the present invention further provides a digital human system that executes the software control method based on the digital human system as described in the first aspect. The system includes: An interaction module, configured to obtain multi-modal input information input by a user, send the multi-modal input information to a large language model, and obtain keyword information and operation information feedback by the large language model; A linkage control module, configured to control the third-party software to execute the operation information according to the third-party software information in the keyword information and generate output text information; A voice synthesis module, configured to generate the output voice according to the received output text information; An image generation module, configured to simulate a digital human and play the output voice through the digital human.

[0012] In a preferred embodiment of the present invention, the above digital human system further includes: A voice wake-up module, configured to receive a wake-up voice input by a user, extract features of the wake-up voice, identify user identity information corresponding to the wake-up voice, and wake up the digital human system when the user identity information is successfully obtained; A knowledge graph module, configured to obtain at least one document corresponding to the keyword information from a knowledge base according to the query information in the keyword information, and extract and classify the content in each document according to the operation information to obtain output text information.

[0013] In a third aspect, an embodiment of the present invention further provides an electronic device, including a processor and a memory, where the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the software control method based on the digital human system in the first aspect above.

[0014] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the software control method based on the digital human system in the first aspect above.

[0015] The embodiments of the present invention bring the following beneficial effects: The embodiments of the present invention provide a software control method based on a digital human system. By analyzing multi-modal input information input by a user through a large language model, extracting user requirements included in the multi-modal input information to obtain keyword information and operation information, and controlling a third-party software to execute the operation information according to the third-party software information in the keyword, the user can control the third-party software using the digital human system, improving the usability of the digital human system and enriching the application scenarios of the digital human system. After the third-party software executes the operation information, output voice is generated and played using the digital human system, realizing human-computer interaction and enhancing the user's human-computer interaction experience.

[0016] Other features and advantages of the present invention will be described in the following description, or some features and advantages can be inferred from the description without doubt, or can be known by implementing the above technologies of the present invention.

[0017] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given below in conjunction with the accompanying drawings for detailed description. Brief Description of the Drawings

[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of a software control method based on a digital human system provided by an embodiment of the present invention; Figure 2 It is a flowchart of another software control method based on a digital human system provided by an embodiment of the present invention; Figure 3 It is a flowchart of another software control method based on a digital human system provided by an embodiment of the present invention; Figure 4 It is a schematic structural diagram of a digital human system provided by an embodiment of the present invention; Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Description of the Embodiments

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0021] Today, with the digital wave sweeping through all industries, customer service and interactive experience have become the key to the digital transformation of enterprises. In order to meet the market demand for intelligent software and the needs of customers for more convenient, intelligent, and simpler human-computer interaction, by developing an AI digital human system, the work efficiency of users has been improved, the human-computer interaction experience has been enhanced, and personalized services can be provided for different customers, reducing costs and increasing efficiency. However, traditional digital humans only provide dialogue functions and lack the ability to control third-party software. We hope that the AI digital human system can be connected to third-party software to achieve software control through dialogue with the AI digital human, such as querying energy consumption, switching software, switching different panel interfaces of software, demonstrating software functions, etc. In addition, the AI digital human system also retains the functions of providing dialogue and search, and can obtain relevant information from the local database to provide question-and-answer services.

[0022] Based on this, a software control method based on a digital human system provided by an embodiment of the present invention analyzes multimodal input information input by a user through a large language model, extracts the user requirements included in the multimodal input information to obtain keyword information and operation information, and controls a third-party software to execute the operation information according to the third-party software information in the keywords, so as to enable the user to control the third-party software using the digital human system, improve the usability of the digital human system, enrich the application scenarios of the digital human system, and enhance the user's human-computer interaction experience.

[0023] For ease of understanding of this embodiment, a software control method based on a digital human system disclosed by an embodiment of the present invention will be introduced in detail first.

[0024] Embodiment 1 An embodiment of the present invention provides a software control method based on a digital human system. The digital human system is a highly integrated artificial intelligence and virtual reality technology that can create and drive virtual avatars with human characteristics. These virtual avatars can perform tasks in various application scenarios, such as customer service, education tutoring, entertainment interaction, etc. In the embodiment of the present invention, the digital human system is installed in an electronic device, and the electronic device at least includes a camera, a display, a keyboard, a microphone, a speaker, etc. The user can input multimodal input information to the digital human system through devices such as a camera, a display, a keyboard, and a microphone. The digital human system can display the digital human avatar simulated by the digital human system through the display and output voice to the user in combination with the speaker, so as to achieve human-computer interaction. A third-party software can also be installed on the electronic device, and the digital human system can control the third-party software to execute relevant operations according to the multimodal input information input by the user to complete the user requirements.

[0025] Figure 1 It is a flowchart of a software control method based on a digital human system provided by an embodiment of the present invention. As Figure 1 shown, the software control method based on the digital human system may include the following steps: Step S101, obtain multimodal input information input by a user.

[0026] Multimodal input information refers to at least one type of information input by a user to the digital human system. Among them, the types of multimodal input information include voice, image, text, gesture, etc. Exemplarily, the user can input text through a keyboard, input voice through a microphone, and capture gestures, videos or images through a camera.

[0027] Specifically, the digital human system can receive multimodal input information input by the user through a variety of input devices. Exemplarily, in an interface with voice and text input functions, the user can either speak their needs into the microphone, such as "Recommend some good movies for me", and at the same time can also enter supplementary information in the text box, such as "Action type"; such as "Use drawing software to help me draw a simple cartoon kitten", and at the same time can further supplement information in the text box, such as "The kitten's eyes should be bigger and the color is blue". At this time, the digital human system will collect voice information and text information respectively and combine them into multimodal input information.

[0028] Step S102, send the multimodal input information to the large language model to obtain the keyword information and operation information fed back by the large language model.

[0029] The large language model refers to a deep learning model in the field of natural language processing. It is trained through a large amount of text data and can understand and generate human natural language. Keyword information refers to the information extracted from the multimodal input information. Expressing the user's needs through keyword information can be key nouns, verbs, etc. that represent the user's intention. Operation information refers to the operations that the digital human system needs to perform, that is, the specific actions that the user wants to perform. It can be understood that the operation information is used to describe the operations that the digital human system needs to perform when meeting the user's needs, such as opening a certain software, querying a certain information, etc.

[0030] Specifically, the digital human system sends the multimodal input information to the large language model. The large language model analyzes and understands the multimodal input information, extracts the keyword information contained in the multimodal input information according to its pre-trained knowledge and algorithms, and determines the operation information that needs to be performed when completing the user's needs expressed by the keyword information, and then feeds back the keyword information and operation information to the digital human system.

[0031] Exemplarily, the multimodal input information input by the user is the text information "Use drawing software to help me draw a simple cartoon kitten." The digital human inputs the text information into the large language model. After analysis, the large language model extracts keyword information such as "drawing software" and "cartoon kitten", and the operation information may be "Open the drawing software and draw a cartoon kitten graphic".

[0032] Further, step A can be adopted 1 - Step A 2 Obtain the keyword information and operation information fed back by the large language model.

[0033] Step A 1 , input the voice information into the speech recognition model to obtain the input text information corresponding to the multimodal input information.

[0034] Step A 2 , send the input text information to the large language model to obtain the keyword information and operation information feedback by the large language model.

[0035] The text recognition model is used to recognize speech information and convert it into text information.

[0036] Before sending the multi-modal input information to the large language model, these information need to be preprocessed. For speech information, speech recognition can be performed first to convert it into text form to form the input text information. Send the input text information to the large language model. After receiving the input text information, the large language model analyzes the semantics, intentions and other content of the input information based on the trained model parameters inside it, and extracts keyword information (such as key nouns, verbs, etc. representing the user's intention) and operation information (such as specific actions the user wants to perform, like opening a certain software, querying a certain type of information, etc.). Then these information are returned to the client calling it in a pre-agreed format (such as JSON format or other custom formats).

[0037] Furthermore, the multi-modal input information can also include image data. For image data, its feature vectors or other descriptive information (such as color histograms, texture features, etc.) can be extracted, and then the preprocessed information is combined into a data format suitable for the input of the large language model according to a certain format. For example, the text after speech conversion, image features and the original text (if there is text input) can be spliced into a long text sequence, or the relationship between these multi-modal information can be represented in the form of structured data (such as JSON format). Use standard communication protocols (such as HTTP / HTTPS protocol, gRPC, etc.) to communicate with the server where the large language model is located. Build a request message body and encapsulate the preprocessed multi-modal information into the request. After receiving the request, the large language model will process the input information using deep learning algorithms and analyze the semantics, intentions and other content of the input information based on the trained model parameters inside it.

[0038] Preprocessing the speech information through the text recognition model can unify the format of the input multi-modal input information, convert it into a data format suitable for the input of the large language model, and improve the efficiency and accuracy of the large language model when analyzing the multi-modal input information.

[0039] Step S103, according to the third-party software information in the keyword information, control the third-party software to execute the operation information and generate output text information.

[0040] A third-party software refers to software that is independent of the digital human system. It is understandable that the third-party software can be installed on the same electronic device as the digital human system or on different electronic devices. Third-party software information is used to describe the third-party software and can include information such as the name, interface, and function identifier of the third-party software. The keyword information can include third-party software information, user sentiment, and other information. The output text information refers to the text information that the digital human system feeds back to the user after controlling the third-party software to execute the operation information. The output text information can be pre-set text information, such as "completed"; it can also be a text description formed based on the execution result.

[0041] Specifically, the digital human system transmits the operation information through the interface of the third-party software (such as the SDK or API provided by the software) according to the third-party software information involved in the keyword information, and controls the third-party software to execute the operation information. After the third-party software executes the operation information, it returns the result, and the digital human system organizes the result into output text information. Exemplarily, the digital human system identifies that the third-party software mentioned in the keyword is a certain professional drawing software. Based on the operation information "open the drawing software and draw a cartoon kitten", it sends an instruction through the API of the software to create a new canvas and draw a cartoon kitten with big blue eyes. After the drawing software completes the drawing, the digital human system obtains the drawn graphic and organizes it into the text information "the cartoon kitten has been drawn as required"; the digital human system can also directly obtain the pre-set output text information after the drawing software completes the drawing.

[0042] Further, according to the third-party software information (such as software name, function identifier, etc.) in the keyword information obtained in step S102, a list of locally installed software or a pre-configured software mapping table can be searched. For example, if "WeChat" is mentioned in the keyword information, the identifier corresponding to WeChat (such as program path, package name, etc.) is found in the local software list. The operation information is matched with the function of the third-party software, which can be realized by reading the software's documentation, API interface description, or an operation mapping model pre-established by a machine learning algorithm. For example, if the operation information is "send a message", for the WeChat software, it is necessary to know how to implement the function of sending a message through its API or automated script (such as simulating the user clicking the send button, etc.). According to the above mapping relationship, the corresponding function of the third-party software is called. If it is through the API interface, the request parameters are sent according to the interface specification; if it is by simulating the user operation, the operation is performed according to the predetermined process (such as starting the software, navigating to the specified interface, inputting content, etc.). When the third-party software successfully executes the operation, an output text message is generated according to the result of the operation. For example, if a message is successfully sent in WeChat, the output text message can be "The message has been successfully sent through WeChat". If the operation fails, a corresponding prompt message may be generated based on the error code or abnormal situation, such as "failed to send message".

[0043] Step S104: input the output text information into a speech synthesis module to obtain output speech.

[0044] The speech synthesis module is used to convert the output text information into speech information. The speech synthesis module can be divided into multiple speech models according to different timbres, such as girl voice, boy voice, young voice, uncle voice, lady voice, old voice, etc.

[0045] Specifically, the speech synthesis module generally adopts text-to-speech (TTS) technology, receives the output text information generated in step S103, and converts the text into a corresponding speech signal according to a preset speech model. For example, the user can pre-select a speech model with a suitable timbre, and input the text information "a cartoon kitten has been drawn as required" into the speech synthesis module. According to the pre-selected speech model (such as a girl's voice), the module will convert the text into speech and generate an output speech with a corresponding timbre. In the process of converting text into speech, the speech model can obtain the user's emotional tendency in the keyword information, and synthesize the output speech of the corresponding tone according to the user's emotional tendency. For example, if the user's emotional tendency is happy, the tone of the synthesized output speech rises, indicating a happy state; if the user's emotional tendency is serious, the tone of the synthesized output speech is stable, indicating a stable state, etc.

[0046] Step S105: Based on the output speech, control the digital human simulated by the digital human system to play the output speech.

[0047] The digital human system synchronizes the output speech with the actions of the digital human simulated by the digital human system, such as mouth shapes and expressions, through its internal image generation module. The digital human system controls the facial expressions and mouth shape changes of the digital human according to the content and rhythm of the output speech, making it look like it is truly speaking these speech contents. Exemplarily, when the output speech is "The recommended action movies for you are: the 'Fast and Furious' series and the 'Mission: Impossible' series", the digital human system controls the digital human to make appropriate expressions (such as smiling, presenting gesture), and at the same time, the mouth shape is synchronized with the speech content to play this output speech.

[0048] A software control method based on a digital human system provided by an embodiment of the present invention analyzes the multimodal input information input by a user through a large language model, extracts the user requirements included in the multimodal input information to obtain keyword information and operation information, and controls a third-party software to execute the operation information according to the third-party software information in the keywords, realizing the control of the third-party software by the user using the digital human system, improving the usability of the digital human system, and enriching the application scenarios of the digital human system. After the third-party software finishes executing the operation information, output speech is generated and played using the digital human system, realizing human-computer interaction and enhancing the user's human-computer interaction experience.

[0049] Embodiment 2 The embodiment of the present invention also provides another software control method based on a digital human system; this method is implemented on the basis of the method in the above embodiment; this method focuses on the specific implementation before obtaining the multimodal input information input by the user.

[0050] Figure 2 The flowchart of another software control method based on a digital human system provided by an embodiment of the present invention is as Figure 2 shown, and this software control method based on a digital human system may include the following steps: Step S201: Receive the wake-up word speech input by the user.

[0051] The wake-up word is used to wake up the digital human system, and the wake-up word speech refers to the speech received by the digital human system when the user says the wake-up word. Only after the digital human system is woken up can it continue to perform human-computer interaction with the user. The user can customize the wake-up word speech of the digital human system before using the digital human system. Specifically, the user can say the wake-up word to the digital human system, and the digital human system will receive the wake-up word said by the user as the wake-up word speech.

[0052] Step S202: Extract features from the wake-up voice and identify the user identity information corresponding to the wake-up voice.

[0053] User identity information is used to describe the users of the digital human system. It can be understood that the digital human system can provide human-computer interaction functions only for users whose identity information is stored in the digital human system. The digital human system can pre-store the voice features of at least one user as pre-stored voice features. One user corresponds to one pre-stored voice feature, and the user identity information can be used as the unique label of the pre-stored voice feature. The pre-stored voice features include at least voice features such as timbre, pitch, and accent.

[0054] Specifically, after receiving the wake-up voice, the digital human system can extract features from the wake-up voice to obtain voice features such as the timbre, pitch, and accent of the wake-up voice. The extracted voice features are matched with the pre-stored voice features stored in the digital human system. If there is a pre-stored voice feature in the at least one pre-stored voice features stored in the digital human system that is the same as or has a similarity exceeding the threshold with the voice features of the wake-up voice, it indicates a successful match, and the unique label corresponding to the pre-stored voice feature that is the same as the voice features of the wake-up voice is used as the user identity information.

[0055] Exemplarily, for the wake-up voice of "Xiaomi AI Assistant" spoken by the user, the system first extracts the voice features of the voice through the MFCC algorithm to obtain a digital sequence containing voice spectrum information. Then, the voice features are compared with multiple pre-stored voice features of users pre-stored in the system. Assume that the pre-stored voice features of user A and user B are stored in the system. By calculating the similarity, it is found that the features of this wake-up voice have the highest similarity with the pre-stored voice features of user A, reaching 80% (the set threshold is 70%), then the system identifies the user identity corresponding to this wake-up voice as user A.

[0056] Step S203: Wake up the digital human system when the user identity information is successfully obtained.

[0057] The successful acquisition of user identity information indicates that there is a pre-stored voice feature in the at least one pre-stored voice features stored in the digital human system that is the same as or has a similarity exceeding the threshold with the voice features of the wake-up voice. At this time, the digital human system is woken up, and step S204 is continued to be executed.

[0058] Further, the digital human system can be woken up through step B 1 - Step B 2 Wake up the digital human system.

[0059] Step B 1 , query the wake-up permission corresponding to the user identity information according to the black and white list.

[0060] Step B 2 When the wake-up permission is set to allow wake-up, wake up the digital human system.

[0061] The black and white list is used to record the wake-up permissions corresponding to different user identity information. The wake-up permissions include allow wake-up and prohibit wake-up. Allow wake-up means that the user is allowed to wake up the digital human system, and prohibit wake-up means that the user is prohibited from waking up the digital human system. Among them, the black and white list includes a black list and a white list. The black list is used to record the user identity information of users prohibited from waking up, and the white list is used to record the user identity information of users allowed to wake up.

[0062] Specifically, when the user identity information is successfully obtained, query the user identity information in the black list and the white list respectively. If the user identity information is found in the black list, it indicates that the wake-up permission corresponding to the user identity information is prohibit wake-up, and at this time, the digital human system is prohibited from waking up. If the user identity information is found in the white list, it indicates that the wake-up permission corresponding to the user identity information is allow wake-up, and at this time, the digital human system is woken up.

[0063] Step S204, obtain the multi-modal input information input by the user.

[0064] Step S205, send the multi-modal input information to the large language model, and obtain the keyword information and operation information fed back by the large language model.

[0065] Step S206, according to the third-party software information in the keyword information, control the third-party software to execute the operation information, and generate output text information.

[0066] Step S207, input the output text information into the speech synthesis module to obtain the output speech.

[0067] Step S208, based on the output speech, control the digital human simulated by the digital human system to play the output speech.

[0068] The software control method based on the digital human system provided by the embodiments of the present invention extracts the features of the wake-up voice to identify the user identity information. When the user identity information is successfully obtained, the digital human system is woken up, and the usage permission of the digital human system can be set, which can prevent unauthorized users from using the digital human system, so that the system resources can be concentrated on serving users with needs, and ensure the stable and efficient operation of the overall performance of the digital human system.

[0069] Embodiment 3 The embodiment of the present invention also provides another software control method based on a digital human system; this method is implemented on the basis of the method in the above embodiment; this method focuses on describing the specific implementation manner after sending the multi-modal input information to a large language model and obtaining the keyword information and operation information feedback by the large language model.

[0070] Figure 3 It is a flowchart of another software control method based on a digital human system provided by the embodiment of the present invention, as Figure 3 shown, this software control method based on a digital human system may include the following steps: Step S301, obtain multi-modal input information input by a user.

[0071] Step S302, send the multi-modal input information to a large language model, and obtain the keyword information and operation information feedback by the large language model.

[0072] Step S303, according to the query information in the keyword information, obtain at least one document corresponding to the keyword information from a knowledge base.

[0073] The query information is used to describe the content that the user wants to query. The knowledge base refers to a pre-constructed repository storing a large number of document resources in various fields, and these documents can be in various forms, such as text files, reports, papers, etc.

[0074] Specifically, in step S302, after the large language model processes the multi-modal input information, it feedbacks the keyword information and operation information. The keyword information contains the query information, and according to the query information, at least one document corresponding to the query information is accurately located and obtained in the knowledge base.

[0075] Exemplarily, assume that the user inputs the text "Introduce the characteristics and historical background of the architectural style of the Forbidden City". After step S302, the query information in the keyword information feedback by the large language model is "Characteristics and historical background of the architectural style of the Forbidden City". At this time, there are many documents about Chinese ancient architecture, history of the Forbidden City, etc. stored in the knowledge base, and the system will search for relevant documents in the knowledge base according to this query information.

[0076] Step S304, according to the operation information, extract and classify the content in each of the documents to obtain output text information.

[0077] The operation information is used to describe the process of processing the content in the document. It is understandable that the document obtained in step S303 may be long, rich in content, and contain a lot of unnecessary information. The operation information provides a clear direction for processing these documents. It specifies which key content to extract from the document and how to classify these contents. Through step S304, the useful information in the document can be refined and organized so that it can be presented in a clearer and more organized form.

[0078] Based on the previous example, assume that the operation information fed back by the large language model is "extract the structural features and decorative features of the architectural style, as well as the construction time and purpose of construction in the historical background, and classify them according to the information category". For the acquired documents, the digital human system will conduct a detailed analysis of each one. For example, in Document 1, the following information may be extracted: "Structural features: the use of a raised beam wooden frame, the architectural layout is symmetrical and rigorous; decorative features: a large number of decorative techniques such as color paintings and carvings are used; construction time: It was built in the fourth year of Yongle in the Ming Dynasty (1406) and took 14 years to complete; construction purpose: as the royal palace of the Ming and Qing dynasties, it demonstrates the majesty of imperial power". Then, the system will classify according to the categories required by the operation information, and the final output text information may be as follows: Architectural style features: Structural features: It adopts a raised-beam wooden frame and the architectural layout is symmetrical and rigorous.

[0079] Decorative features: Use a lot of painting, carving and other decorative techniques.

[0080] Historical background: Construction time: Construction began in the fourth year of Yongle in the Ming Dynasty (1406) and took 14 years to complete.

[0081] Purpose of construction: Served as the royal palace of the Ming and Qing dynasties to demonstrate imperial power.

[0082] Step S305: input the output text information into a speech synthesis module to obtain output speech.

[0083] Step S306: Based on the output voice, control the digital human simulated by the digital human system to play the output voice.

[0084] Specifically, step C 1 -Step C 2 The output voice is played.

[0085] Step C 1 , input the output speech into the emotion simulation model to obtain the emotion corresponding to the output speech.

[0086] Step C2 , according to the emotion, controlling the digital human simulated by the digital human system to play the output voice.

[0087] The emotion simulation model is used to identify the emotion of the output speech. It is understandable that the emotion simulation model is an artificial intelligence model that has been trained with a large amount of data. It can recognize various acoustic features in speech, such as pitch, timbre, speaking speed, volume changes, etc., and judge the emotions contained in the speech based on these features. These acoustic features contained in the speech are closely related to human emotional expression, and different emotions are often expressed through different acoustic features. For example, when excited, the speaking speed may increase and the pitch may increase, and when sad, the speaking speed may decrease and the pitch may decrease. By inputting the output speech into the emotion simulation model, the emotion simulation model can accurately identify the emotion corresponding to the output speech. The digital human system can adjust the digital human's expression, body movements and voice expression according to different emotions. For example, if the recognized emotion is happiness, the digital human may smile, shake his body slightly, and the tone of his voice will be more cheerful and bright; if it is a sad emotion, the digital human may show a melancholy expression, his body posture is relatively low, and his voice will be lower and slower. Such processing makes the digital human's expression more vivid and realistic, and enables better emotional interaction with users, thus enhancing the user experience.

[0088] The software control method based on the digital human system provided by the embodiment of the present invention obtains the corresponding document from the knowledge base through the query information in the keyword information, and classifies the content in the document according to the operation information to obtain output text information, thereby realizing the query function of the digital human system and greatly saving the user's time in collating and analyzing information.

[0089] Example 4 Corresponding to the above method embodiment, the embodiment of the present invention provides a digital human system, Figure 4 A schematic diagram of the structure of a digital human system provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the digital human system may include: The interaction module 401 is used to obtain multimodal input information input by the user, and send the multimodal input information to the large language model to obtain keyword information and operation information fed back by the large language model; A linkage control module 402, configured to control the third-party software to execute the operation information and generate output text information according to the third-party software information in the keyword information; The speech synthesis module 403 is used to generate the output speech according to the received output text information; The image generation module 404 is used to simulate a digital human and play the output voice through the digital human.

[0090] An embodiment of the present invention provides a digital human system that analyzes multimodal input information input by a user through a large language model, extracts the user requirements contained in the multimodal input information, obtains keyword information and operation information, and controls a third-party software to execute the operation information according to the third-party software information in the keywords, realizing the control of the third-party software by the user using the digital human system, improving the usability of the digital human system, and enriching the application scenarios of the digital human system. After the third-party software finishes executing the operation information, an output voice is generated and played using the digital human system, realizing human-computer interaction and enhancing the user's human-computer interaction experience.

[0091] In some embodiments, the system further includes: A voice wake-up module for receiving the wake-up voice input by the user, extracting features of the wake-up voice, identifying the user identity information corresponding to the wake-up voice, and waking up the digital human system when the user identity information is successfully obtained; A knowledge graph module for obtaining at least one document corresponding to the keyword information from a knowledge base according to the query information in the keyword information, and extracting and classifying the content in each document according to the operation information to obtain output text information.

[0092] In some embodiments, the interaction module is further configured to: Input the voice information into a text recognition model to obtain input text information corresponding to the multimodal input information; Send the input text information to the large language model to obtain the keyword information and operation information fed back by the large language model.

[0093] In some embodiments, the voice wake-up module is further configured to: Query the wake-up permission corresponding to the user identity information according to a black and white list; Wake up the digital human system when the wake-up permission is to allow waking up.

[0094] In some embodiments, the linkage control module is further configured to: Obtain at least one document corresponding to the keyword information from a knowledge base according to the query information in the keyword information; Extract and classify the content in each document according to the operation information to obtain output text information.

[0095] In some embodiments, the image generation module is further configured to: Input the output voice into an emotion simulation model to obtain the emotion corresponding to the output voice; According to the emotion, control the digital human simulated by the digital human system to play the output voice.

[0096] The digital human system provided by the embodiments of the present invention has the same implementation principle and the same technical effects as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference may be made to the corresponding content in the foregoing method embodiments.

[0097] Embodiment 5 The embodiments of the present invention further provide an electronic device for running the foregoing software control method based on the digital human system; refer to Figure 5 the structural schematic diagram of an electronic device shown in the figure. The electronic device includes a memory 500 and a processor 501. Among them, the memory 500 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor 501 to implement the foregoing software control method based on the digital human system.

[0098] Furthermore, Figure 5 the electronic device shown in the figure further includes a bus 502 and a communication interface 503, and the processor 501, the communication interface 503 and the memory 500 are connected through the bus 502.

[0099] Among them, the memory 500 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 503 (which may be wired or wireless), a communication connection between the system network element and at least one other network element is realized, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used. The bus 502 may be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0100] The processor 501 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 501 or the instructions in the form of software. The above-mentioned processor 501 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 500, and the processor 501 reads the information in the memory 500 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0101] The embodiments of the present invention also provide a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the above-mentioned software control method based on the digital human system. For the specific implementation, reference may be made to the method embodiments, and details are not described herein again.

[0102] The computer program product for implementing the software control method based on the digital human system provided by the embodiments of the present invention includes a computer-readable storage medium storing non-volatile program code executable by a processor. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For the specific implementation, reference may be made to the method embodiments, and details are not described herein again.

[0103] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and details are not described herein again.

[0104] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.

[0105] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0106] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0107] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0108] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A software control method based on a digital human system, applied to a digital human system, characterized in that: include: Get multimodal input information entered by the user; Sending the multimodal input information to the large language model, and obtaining keyword information and operation information fed back by the large language model; According to the third-party software information in the keyword information, controlling the third-party software to execute the operation information and generate output text information; Inputting the output text information into a speech synthesis module to obtain output speech; Based on the output voice, the digital human simulated by the digital human system is controlled to play the output voice.

2. The method according to claim 1, characterized in that The multimodal input information includes voice information; sending the multimodal input information to the large language model to obtain keyword information and operation information fed back by the large language model includes: Inputting the speech information into a text recognition model to obtain input text information corresponding to the multimodal input information; The input text information is sent to the large language model to obtain keyword information and operation information fed back by the large language model.

3. The method according to claim 1, characterized in that Before obtaining the multimodal input information entered by the user, it also includes: Receive the wake-up word voice input by the user; Extracting features of the wake-up word voice and identifying user identity information corresponding to the wake-up word voice; When the user identity information is successfully obtained, the digital human system is awakened.

4. The method according to claim 3, characterized in that The step of waking up the digital human system comprises: According to the black and white lists, query the wake-up permission corresponding to the user identity information; When the wake-up authority allows wake-up, the digital human system is woken up.

5. The method according to claim 1, characterized in that After sending the multimodal input information to the large language model and obtaining the keyword information and operation information fed back by the large language model, the method further includes: According to the query information in the keyword information, obtaining at least one document corresponding to the keyword information from a knowledge base; According to the operation information, the contents in each of the documents are extracted and classified to obtain output text information.

6. The method according to claim 1, characterized in that The step of controlling the digital human simulated by the digital human system to play the output voice based on the output voice comprises: Inputting the output speech into an emotion simulation model to obtain the emotion corresponding to the output speech; According to the emotion, the digital human simulated by the digital human system is controlled to play the output voice.

7. A digital human system, characterized in that: Executing the software control method based on the digital human system according to any one of claims 1 to 6, the system comprises: An interaction module, used to obtain multimodal input information input by a user, and send the multimodal input information to a large language model, and obtain keyword information and operation information fed back by the large language model; A linkage control module, used to control the third-party software to execute the operation information and generate output text information according to the third-party software information in the keyword information; A speech synthesis module, used for generating the output speech according to the received output text information; The image generation module is used to simulate a digital human and play the output voice through the digital human.

8. The system according to claim 7, characterized in that Also includes: A voice wake-up module, used to receive a wake-up word voice input by a user, perform feature extraction on the wake-up word voice, identify the user identity information corresponding to the wake-up word voice, and wake up the digital human system if the user identity information is successfully obtained; The knowledge graph module is used to obtain at least one document corresponding to the keyword information from the knowledge base according to the query information in the keyword information, and to extract and classify the content in each of the documents according to the operation information to obtain output text information.

9. An electronic device, characterized in that: It comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the software control method based on the digital human system as described in any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the software control method based on the digital human system as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Virtual human interaction software bus system and an implementation method thereof

    CN109917917A

  • Control method of smart home system and smart home system

    CN114220442A

  • Human-computer interaction method, system and equipment based on AI (Artificial Intelligence) and medium

    CN118466751A

  • Efficient digital human interaction system fusing multi-modal information

    CN118897887A

  • Content generation method and device based on large model, and electronic equipment

    CN119180890A

Cited By

  • Digital human video generation method, training method and device of digital human generation model, and computer equipment

    CN121330444A

  • Digital human video generation method, digital human generation model training method, device, and computer equipment

    CN121330444B