Software control method, system, device and medium based on digital human system

By obtaining multimodal input information through the digital human system and using a large language model to control third-party software operations, the problem of the digital human system's single function is solved, the control of third-party software and rich application scenarios are realized, and the user's interactive experience is enhanced.

CN120045684BActive Publication Date: 2025-10-21HANGZHOU SHUJU CHAIN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510496416.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-10-21
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing digital human system has relatively simple functions, mainly limited to information interaction in the form of text or voice, and lacks the ability to control third-party software, resulting in limited ease of use and application scenarios.

Method used

The digital human system obtains the user's multimodal input information, uses a large language model to analyze and extract keyword information and operation information, controls third-party software to perform corresponding operations, and generates output speech to achieve human-computer interaction.

Benefits of technology

It improves the ease of use of the digital human system, enriches application scenarios, enhances the user's human-computer interaction experience, and realizes the control of third-party software and information query functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045684B_ABST
    Figure CN120045684B_ABST
Patent Text Reader

Abstract

The application provides a software control method, system, device and medium based on a digital human system, relates to the technical field of artificial intelligence, and comprises the following steps: acquiring multi-modal input information input by a user; sending the multi-modal input information to a large language model, acquiring keyword information and operation information fed back by the large language model; controlling a third-party software to execute the operation information according to third-party software information in the keyword information, and generating output text information; inputting the output text information into a speech synthesis module to obtain output speech; and controlling a digital human simulated by the digital human system to play the output speech based on the output speech. The technical scheme of the embodiment of the application can realize control of the digital human system on the third-party software, improve the use convenience of the digital human, enrich the application scenarios of the digital human, and enhance the interactive experience of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a software control method, system, device and medium based on a digital human system. Background Art

[0002] With the development of artificial intelligence and avatar technology, digital humans have gradually become an important component of intelligent interaction. Using natural language processing, digital humans can communicate with users, providing services such as information query and entertainment interaction. Currently, digital humans are being applied in various scenarios, such as intelligent customer service, virtual assistants, and online education.

[0003] Currently, traditional AI digital human systems rely primarily on pre-trained language models to understand and respond to user questions, with a core focus on providing a high-quality conversational experience. However, this model also results in relatively limited functionality for these systems, primarily limited to text or voice interaction. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a software control method based on a digital human system. Through human-computer interaction, the digital human system is used to control third-party software to perform related operations, thereby improving the ease of use of the digital human system, enriching the application scenarios of digital humans, and enhancing the user's interactive experience.

[0005] In a first aspect, an embodiment of the present invention provides a software control method based on a digital human system, which is applied to the digital human system. The method includes:

[0006] Get multimodal input information entered by the user;

[0007] Sending the multimodal input information to a large language model to obtain keyword information and operation information fed back by the large language model;

[0008] According to the third-party software information in the keyword information, controlling the third-party software to execute the operation information and generate output text information;

[0009] Inputting the output text information into a speech synthesis module to obtain output speech;

[0010] Based on the output voice, the digital human simulated by the digital human system is controlled to play the output voice.

[0011] In a preferred embodiment of the present invention, the multimodal input information includes voice information; sending the multimodal input information to the large language model and obtaining keyword information and operation information fed back by the large language model includes:

[0012] Inputting the speech information into a text recognition model to obtain input text information corresponding to the multimodal input information;

[0013] The input text information is sent to the large language model to obtain keyword information and operation information fed back by the large language model.

[0014] In a preferred embodiment of the present invention, before obtaining the multimodal input information input by the user, the following steps are included:

[0015] Receive the wake-up word voice input by the user;

[0016] Extract features of the wake-up word voice and identify user identity information corresponding to the wake-up word voice;

[0017] When the user identity information is successfully obtained, the digital human system is awakened.

[0018] In a preferred embodiment of the present invention, the above-mentioned awakening of the digital human system includes:

[0019] According to the blacklist and whitelist, query the wake-up permission corresponding to the user identity information;

[0020] When the wake-up authority allows wake-up, the digital human system is woken up.

[0021] In a preferred embodiment of the present invention, after sending the multimodal input information to the large language model and obtaining the keyword information and operation information fed back by the large language model, the following steps are included:

[0022] Retrieving at least one document corresponding to the keyword information from a knowledge base according to query information in the keyword information;

[0023] According to the operation information, the contents of each document are extracted and classified to obtain output text information.

[0024] In a preferred embodiment of the present invention, the above-mentioned controlling the digital human simulated by the digital human system to play the output voice based on the output voice includes:

[0025] Inputting the output speech into an emotion simulation model to obtain the emotion corresponding to the output speech;

[0026] According to the emotion, the digital human simulated by the digital human system is controlled to play the output voice.

[0027] In a second aspect, an embodiment of the present invention further provides a digital human system that executes the software control method based on the digital human system as described in the first aspect, the system comprising:

[0028] An interaction module is used to obtain multimodal input information input by the user, send the multimodal input information to the large language model, and obtain keyword information and operation information fed back by the large language model;

[0029] A linkage control module, configured to control the third-party software to execute the operation information and generate output text information according to the third-party software information in the keyword information;

[0030] A speech synthesis module, configured to generate the output speech according to the received output text information;

[0031] The image generation module is used to simulate a digital human and play the output voice through the digital human.

[0032] In a preferred embodiment of the present invention, the above-mentioned digital human system further includes:

[0033] A voice wake-up module is used to receive a wake-up word voice input by the user, perform feature extraction on the wake-up word voice, identify the user identity information corresponding to the wake-up word voice, and wake up the digital human system if the user identity information is successfully obtained;

[0034] The knowledge graph module is used to obtain at least one document corresponding to the keyword information from the knowledge base according to the query information in the keyword information, and to extract and classify the content in each of the documents according to the operation information to obtain output text information.

[0035] In a third aspect, an embodiment of the present invention further provides an electronic device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the software control method based on the digital human system of the first aspect mentioned above.

[0036] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the software control method based on the digital human system of the first aspect mentioned above.

[0037] The embodiments of the present invention bring the following beneficial effects:

[0038] Embodiments of the present invention provide a software control method based on a digital human system. This method uses a large language model to analyze multimodal input information from users, extracting user needs contained in the multimodal input information to obtain keyword information and operation information. Based on the third-party software information contained in the keywords, the method controls third-party software to execute the operation information. This enables users to control third-party software using the digital human system, improving the system's ease of use and enriching its application scenarios. After the third-party software completes the operation information, it generates output speech and plays it using the digital human system, enabling human-computer interaction and enhancing the user's human-computer interaction experience.

[0039] Other features and advantages of the present invention will be set forth in the following description, or some features and advantages may be inferred or unambiguously determined from the description, or may be learned by implementing the above-mentioned technology of the present invention.

[0040] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0042] Figure 1 A flowchart of a software control method based on a digital human system provided by an embodiment of the present invention;

[0043] Figure 2 A flowchart of another software control method based on a digital human system provided by an embodiment of the present invention;

[0044] Figure 3 A flowchart of another software control method based on a digital human system provided by an embodiment of the present invention;

[0045] Figure 4 A schematic structural diagram of a digital human system provided by an embodiment of the present invention;

[0046] Figure 5 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0048] As digitalization sweeps across industries, customer service and interactive experiences have become crucial for businesses' digital transformation. To meet market demand for intelligent software and satisfy customer demands for more convenient, intelligent, and simplified human-computer interaction, we are developing AI digital human systems to improve user efficiency, enhance the human-computer interaction experience, and provide personalized services to diverse customers, reducing costs and increasing efficiency. However, traditional digital humans only offer conversational functionality and lack the ability to control third-party software. We aim to integrate the AI ​​digital human system with third-party software, enabling software control through conversations with the AI ​​digital human. This allows for tasks such as querying energy consumption, switching between software interfaces, and demonstrating software functionality. Furthermore, the AI ​​digital human system retains conversational and search capabilities, retrieving relevant information from a local database, and providing question-and-answer services.

[0049] Based on this, an embodiment of the present invention provides a software control method based on a digital human system. The method analyzes the multimodal input information input by the user through a large language model, extracts the user needs contained in the multimodal input information, obtains keyword information and operation information, and controls the third-party software to execute the operation information based on the third-party software information in the keyword, thereby enabling the user to control the third-party software using the digital human system. This improves the ease of use of the digital human system, enriches the application scenarios of the digital human system, and enhances the user's human-computer interaction experience.

[0050] To facilitate understanding of this embodiment, a software control method based on a digital human system disclosed in an embodiment of the present invention is first introduced in detail.

[0051] Example 1

[0052] An embodiment of the present invention provides a software control method based on a digital human system. The digital human system is a highly integrated artificial intelligence and virtual reality technology that can create and drive virtual images with human characteristics. These virtual images can perform tasks in various application scenarios, such as customer service, educational guidance, and entertainment interaction. In an embodiment of the present invention, the digital human system is installed in an electronic device that includes at least a camera, a display, a keyboard, a microphone, and a speaker. A user can input multimodal input information into the digital human system through devices such as a camera, a display, a keyboard, and a microphone. The digital human system can display a digital human image simulated by the digital human system on the display and output voice to the user through the speaker, thereby achieving human-computer interaction. Third-party software can also be installed on the electronic device. The digital human system can control the third-party software to perform related operations based on the multimodal input information entered by the user to meet the user's needs.

[0053] Figure 1 This is a flow chart of a software control method based on a digital human system provided by an embodiment of the present invention. Figure 1 As shown, the software control method based on the digital human system may include the following steps:

[0054] Step S101: Acquire multimodal input information input by a user.

[0055] Multimodal input information refers to at least one type of information input by a user into a digital human system. The types of multimodal input information include voice, image, text, gesture, etc. For example, a user can input text via a keyboard, input voice via a microphone, and capture gestures, videos, or images via a camera.

[0056] Specifically, the Digihuman system can receive multimodal input information from users through various input devices. For example, in an interface with voice and text input, a user can speak their needs into a microphone, such as "recommend some good movies for me," and also enter supplementary information in a text box, such as "action-based"; or "use drawing software to help me draw a simple cartoon kitten," and further supplementary information in the text box, such as "the kitten's eyes should be larger and blue." At this point, the Digihuman system will collect the voice and text information separately and combine them into multimodal input information.

[0057] Step S102: Send the multimodal input information to the large language model to obtain keyword information and operation information fed back by the large language model.

[0058] A large language model is a deep learning model in the field of natural language processing. It is trained with large amounts of text data and can understand and generate natural human language. Keyword information refers to information derived from multimodal input. Keyword information expresses user needs and can be key nouns, verbs, or other expressions of user intent. Operation information refers to the operations that the digital human system must perform—specific actions the user wants to perform. It can be understood that operation information describes the actions the digital human system must perform to meet user needs, such as opening a specific software application or searching for information.

[0059] Specifically, the digital human system sends multimodal input information to the large language model, which analyzes and understands the multimodal input information. Based on its pre-trained knowledge and algorithms, the large language model extracts the keyword information contained in the multimodal input information, and based on the key system information, determines the operation information that needs to be performed to complete the user's needs expressed by the keyword information, and then feeds back the keyword information and operation information to the digital human system.

[0060] For example, the multimodal input information entered by the user is the text message "Use drawing software to help me draw a simple cartoon kitten." The digital human inputs the text message into the large language model. After analysis, the large language model extracts keywords such as "drawing software" and "cartoon kitten." The operation information may be "Open the drawing software and draw a cartoon kitten."

[0061] Furthermore, the keyword information and operation information fed back by the large language model can be obtained through steps A1 and A2.

[0062] Step A1: input the voice information into a text recognition model to obtain input text information corresponding to the multimodal input information.

[0063] Step A2: sending the input text information to the large language model to obtain keyword information and operation information fed back by the large language model.

[0064] The text recognition model is used to recognize voice information and convert it into text information.

[0065] Before sending multimodal input to the large language model, it needs to be preprocessed. For voice information, speech recognition can be performed first to convert it into text, forming the input text information. After receiving the input text information, the large language model analyzes the semantics and intent of the input information based on its internally trained model parameters. It extracts keyword information (such as key nouns and verbs that represent user intent) and operation information (such as the specific action the user wants to perform, such as opening a certain software or searching for information). This information is then returned to the calling client in a predefined format (such as JSON or other custom formats).

[0066] Furthermore, multimodal input information can also include image data. For image data, feature vectors or other descriptive information (such as color histograms and texture features) can be extracted. This preprocessed information is then combined into a data format suitable for large language model input. For example, the converted text, image features, and original text (if text input is provided) can be concatenated into a long text sequence, or the relationships between these multimodal information can be represented as structured data (such as JSON). Standard communication protocols (such as HTTP / HTTPS or gRPC) are used to communicate with the server hosting the large language model. A request message body is constructed, encapsulating the preprocessed multimodal information within the request. Upon receiving the request, the large language model processes the input information using a deep learning algorithm. Based on its internally trained model parameters, it analyzes the semantics, intent, and other aspects of the input information.

[0067] By preprocessing the speech information through the text recognition model, the format of the input multimodal input information can be unified and converted into a data format suitable for the input of the large language model, thereby improving the efficiency and accuracy of the large language model in analyzing the multimodal input information.

[0068] Step S103: According to the third-party software information in the keyword information, the third-party software is controlled to execute the operation information and generate output text information.

[0069] Third-party software refers to software that is independent of the Digital Human system. It is understood that third-party software can be installed on the same electronic device as the Digital Human system, or on a different electronic device. Third-party software information is used to describe the third-party software and can include information such as the third-party software's name, interface, and function identifier. Keyword information can include information such as third-party software information and user emotional tendencies. Output text information refers to the textual information that the Digital Human system provides to the user regarding the execution results after controlling the third-party software to execute operation information. Output text information can be pre-set textual information, such as "Completed," or it can be a textual description compiled based on the execution results.

[0070] Specifically, the Digihuman system, based on the third-party software information mentioned in the keyword information, transmits operational information through an interface with the third-party software (such as the software's SDK or API), controlling the third-party software to execute the operational information. After executing the operational information, the third-party software returns a result, which the Digihuman system then organizes into output text. For example, the Digihuman system identifies the third-party software mentioned in the keyword as professional drawing software. Based on the operational information "Open the drawing software and draw a cartoon kitten," it sends instructions through the software's API to create a new canvas and draw a cartoon kitten with large blue eyes. After the drawing software completes the drawing, the Digihuman system retrieves the resulting image and organizes it into a text message: "The cartoon kitten has been drawn as requested." The Digihuman system can also directly retrieve the preset output text message after the drawing software completes the drawing.

[0071] Furthermore, based on the third-party software information (e.g., software name, function identifier, etc.) in the keyword information obtained in step S102, a local list of installed software or a pre-configured software mapping table can be searched. For example, if the keyword information mentions "WeChat," the corresponding identifier (e.g., program path, package name, etc.) for WeChat is found in the local software list. Matching the operation information with the third-party software's function can be achieved by reading the software's documentation, API interface specifications, or using a pre-established operation mapping model using a machine learning algorithm. For example, if the operation information is "send message," WeChat needs to know how to implement the message sending function through its API or automated script (e.g., simulating a user clicking a send button). Based on the aforementioned mapping, the corresponding function of the third-party software is called. If the operation is performed through the API, the request parameters are sent according to the interface specifications. If the operation is performed by simulating a user operation, the operation is performed according to the predetermined process (e.g., launching the software, navigating to a specified interface, entering content, etc.). When the third-party software successfully executes the operation, it generates output text based on the result of the operation. For example, if a message is successfully sent through WeChat, the output text may be "Message successfully sent via WeChat." If the operation fails, a corresponding prompt message may be generated based on the error code or abnormal situation, such as "failed to send message".

[0072] Step S104: input the output text information into a speech synthesis module to obtain output speech.

[0073] The speech synthesis module converts text output into speech. The speech synthesis module can be divided into multiple voice models based on different timbre, such as female voice, male voice, young voice, uncle voice, mature woman voice, and elderly voice.

[0074] Specifically, the speech synthesis module typically utilizes text-to-speech (TTS) technology to receive the output text information generated in step S103 and convert it into a corresponding speech signal based on a preset speech model. For example, a user can pre-select a speech model with an appropriate timbre and input the text message "A cartoon kitten has been drawn as requested" into the speech synthesis module. Based on the pre-selected speech model (e.g., a girl's voice), the module converts the text into speech and generates output speech with the appropriate timbre. During the text-to-speech conversion process, the speech model can extract the user's emotional orientation from keyword information and synthesize the output speech with the appropriate tone based on the user's emotional orientation. For example, if the user's emotional orientation is happiness, the synthesized output speech will have an up-pitched tone, indicating a happy state; if the user's emotional orientation is serious, the synthesized output speech will have a steady tone, indicating a calm state.

[0075] Step S105: Based on the output voice, controlling the digital human simulated by the digital human system to play the output voice.

[0076] The Digital Human system, through its internal image generation module, synchronizes the output speech with the lip movements, facial expressions, and other movements of the digital human simulated by the system. Based on the content and rhythm of the output speech, the Digital Human system controls the facial expressions and lip movements of the digital human, making it appear as if the speech is actually being spoken. For example, when the output speech is "Recommended action-themed movies include the Fast and Furious series and the Mission: Impossible series," the Digital Human system controls the digital human to produce appropriate expressions (such as a smile or an introduction gesture), synchronizes the lip movements with the speech content, and plays the output speech.

[0077] An embodiment of the present invention provides a software control method based on a digital human system. This method uses a large language model to analyze multimodal user input, extracting user needs contained in the multimodal input to obtain keyword information and operation information. The method then controls third-party software to execute the operation information based on the third-party software information contained in the keyword information. This method enables users to control third-party software using the digital human system, improving the system's ease of use and enriching its application scenarios. After the third-party software completes the operation information, it generates output speech and plays it back using the digital human system, enabling human-computer interaction and enhancing the user's interactive experience.

[0078] Example 2

[0079] The embodiment of the present invention also provides another software control method based on the digital human system; this method is implemented on the basis of the method in the above embodiment; this method focuses on describing the specific implementation method before obtaining the multimodal input information input by the user.

[0080] Figure 2 A flowchart of another software control method based on a digital human system provided by an embodiment of the present invention is shown in FIG. Figure 2 As shown, the software control method based on the digital human system may include the following steps:

[0081] Step S201: Receive a wake-up word voice input by a user.

[0082] The wake-up word is used to wake up the Digital Human system. The wake-up word audio refers to the audio received by the Digital Human system when the user speaks the wake-up word. Only after the wake-up word is spoken can the Digital Human system continue human-computer interaction with the user. Before using the Digital Human system, the user can customize the wake-up word audio for the Digital Human system. Specifically, the user speaks the wake-up word to the Digital Human system, and the Digital Human system will receive the user's wake-up word as the wake-up word audio.

[0083] Step S202: extract features of the wake-up word voice and identify user identity information corresponding to the wake-up word voice.

[0084] User identity information is used to describe the user using the Digital Human system. It is understood that the Digital Human system can only provide human-computer interaction functions to users whose identity information is stored in the Digital Human system. The Digital Human system can pre-store the voice characteristics of at least one user as pre-stored voice characteristics, with one pre-stored voice characteristic corresponding to each user. User identity information can serve as a unique label for the pre-stored voice characteristics. Pre-stored voice characteristics include at least timbre, pitch, accent, and other voice characteristics.

[0085] Specifically, after receiving the wake-up word voice, the Digihuman system can extract features from the wake-up word voice to obtain voice features such as timbre, pitch, and accent. The extracted voice features are matched with pre-stored voice features stored in the Digihuman system. If at least one pre-stored voice feature stored in the Digihuman system is identical to the voice feature of the wake-up word voice or has a similarity exceeding a threshold, it indicates a successful match, and the unique tag corresponding to the pre-stored voice feature identical to the voice feature of the wake-up word voice is used as the user identity information.

[0086] For example, for a user speaking the wake-up word "Xiao Ai," the system first extracts the speech features using the MFCC algorithm, generating a digital sequence containing speech spectrum information. These features are then compared with pre-stored speech features of multiple users stored in the system. Assuming the system has pre-stored speech features for users A and B, and calculating similarity reveals that the wake-up word's features have the highest similarity with those of user A, reaching 80% (the threshold is 70%), the system then identifies the user corresponding to the wake-up word as user A.

[0087] Step S203: When the user identity information is successfully obtained, the digital human system is awakened.

[0088] Successful acquisition of user identity information indicates that among at least one pre-stored voice feature stored in the digital human system, there is a pre-stored voice feature that is identical to the voice feature of the wake-up word voice or whose similarity exceeds a threshold. At this time, the digital human system is woken up and step S204 is continued.

[0089] Furthermore, the digital human system can be awakened through steps B1 and B2.

[0090] Step B1: query the wake-up permission corresponding to the user identity information according to the blacklist and whitelist.

[0091] Step B2: When the wake-up authority allows wake-up, wake up the digital human system.

[0092] The blacklist and whitelist are used to record the wake-up permissions corresponding to different user identity information. Wake-up permissions include allowing wake-up and prohibiting wake-up. Allowing wake-up means that the user is allowed to wake up the Digital Human system, while prohibiting wake-up means that the user is prohibited from waking up the Digital Human system. The blacklist and whitelist include blacklists and whitelists. The blacklist records the user identity information of users who are prohibited from waking up, while the whitelist records the user identity information of users who are allowed to wake up.

[0093] Specifically, when the user identity information is successfully obtained, the user identity information is searched in the blacklist and whitelist respectively. If the user identity information is found in the blacklist, it indicates that the wake-up permission corresponding to the user identity information is prohibited, and the Digihuman system is prohibited from being awakened. If the user identity information is found in the whitelist, it indicates that the wake-up permission corresponding to the user identity information is allowed, and the Digihuman system is awakened.

[0094] Step S204: Acquire multimodal input information input by the user.

[0095] Step S205: Send the multimodal input information to the large language model to obtain keyword information and operation information fed back by the large language model.

[0096] Step S206 : According to the third-party software information in the keyword information, the third-party software is controlled to execute the operation information and generate output text information.

[0097] Step S207: input the output text information into a speech synthesis module to obtain output speech.

[0098] Step S208: Based on the output voice, controlling the digital human simulated by the digital human system to play the output voice.

[0099] The software control method based on the Digihuman system provided in the embodiment of the present invention extracts features from the wake-up word voice and identifies the user's identity information. When the user's identity information is successfully obtained, the Digihuman system is awakened. The usage permissions of the Digihuman system can be set to prevent unauthorized users from using the Digihuman system, so that system resources can be concentrated on serving users in need, ensuring the overall stable performance and efficient operation of the Digihuman system.

[0100] Example 3

[0101] An embodiment of the present invention also provides another software control method based on a digital human system; this method is implemented on the basis of the method in the above embodiment; this method focuses on describing the specific implementation method after sending the multimodal input information to the large language model and obtaining the keyword information and operation information fed back by the large language model.

[0102] Figure 3 A flowchart of another software control method based on a digital human system provided by an embodiment of the present invention is shown in FIG. Figure 3 As shown, the software control method based on the digital human system may include the following steps:

[0103] Step S301: Acquire multimodal input information input by the user.

[0104] Step S302: Send the multimodal input information to the large language model to obtain keyword information and operation information fed back by the large language model.

[0105] Step S303: acquiring at least one document corresponding to the keyword information from a knowledge base according to the query information in the keyword information.

[0106] Query information is used to describe the content that the user wants to query. A knowledge base is a pre-built resource that stores a large number of documents in different fields. These documents can be in various forms, such as text files, reports, papers, etc.

[0107] Specifically, in step S302, the large language model processes the multimodal input information and feeds back keyword information and operation information. The keyword information includes query information, and based on the query information, at least one document corresponding to the query information is accurately located and retrieved in the knowledge base.

[0108] For example, suppose a user enters the text "Describe the architectural style and historical background of the Forbidden City." In step S302, the large language model returns the query information "Architectural style and historical background of the Forbidden City." The knowledge base now contains numerous documents on ancient Chinese architecture and the history of the Forbidden City. Based on this query, the system searches for relevant documents within the knowledge base.

[0109] Step S304: extracting and classifying the contents of each document according to the operation information to obtain output text information.

[0110] Operational information describes the process of processing the content in the document. It's understandable that the document obtained in step S303 may be long and rich in content, containing a lot of unnecessary information. Operational information provides clear direction for processing these documents, specifying which key content to extract from the document and how to categorize it. Step S304 allows the useful information in the document to be refined and organized, presenting it in a clearer and more organized form.

[0111] Based on the previous example, let's assume that the operational information fed back by the large language model is "Extract the structural and decorative features of the architectural style, as well as the construction time and purpose from the historical context, and classify them according to information categories." For each acquired document, the Digital Human system will perform a detailed analysis. For example, in Document 1, it might extract information such as "Structural features: uses a raised-beam wooden frame, and the architectural layout is symmetrical and rigorous; Decorative features: uses a large number of decorative techniques such as paintings and carvings; Construction time: Construction began in the fourth year of Yongle in the Ming Dynasty (1406) and took 14 years to complete; Construction purpose: as the royal palace of the Ming and Qing dynasties, it demonstrated the majesty of imperial power." The system will then classify the information according to the categories required by the operational information, and the final output text information may be as follows:

[0112] Architectural style features:

[0113] Structural features: It adopts a raised-beam wooden structure and the architectural layout is symmetrical and rigorous.

[0114] Decorative features: Use a lot of decorative techniques such as painting and carving.

[0115] Historical Background:

[0116] Construction time: Construction began in the fourth year of Yongle in the Ming Dynasty (1406) and took 14 years to complete.

[0117] Purpose of construction: Served as the royal palace of the Ming and Qing dynasties to demonstrate imperial power.

[0118] Step S305: input the output text information into a speech synthesis module to obtain output speech.

[0119] Step S306: Based on the output voice, control the digital human simulated by the digital human system to play the output voice.

[0120] Specifically, the output voice can be played through steps C1 and C2.

[0121] Step C1: input the output speech into an emotion simulation model to obtain the emotion corresponding to the output speech.

[0122] Step C2: controlling the digital human simulated by the digital human system to play the output voice according to the emotion.

[0123] The emotion simulation model is used to identify the emotion in the output speech. It's understood that the emotion simulation model is an AI model trained with extensive data. It can identify various acoustic features in speech, such as pitch, timbre, speaking rate, and volume variations, and use these features to determine the emotion implied by the speech. These acoustic features in speech are closely related to human emotional expression, and different emotions are often expressed through different acoustic features. For example, excitement may lead to faster speech and higher pitch, while sadness may lead to slower speech and lower pitch. By inputting the output speech into the emotion simulation model, the emotion simulation model can accurately identify the corresponding emotion in the output speech. The digital human system can adjust the digital human's facial expressions, body movements, and vocal expressions based on different emotions. For example, if the recognized emotion is happiness, the digital human may smile, sway slightly, and the voice tone may become more cheerful and bright. If the emotion is sadness, the digital human may appear melancholic, with a relatively slumped posture and a lowered, slower voice. This processing makes the digital human's expression more vivid and realistic, and can better interact with users emotionally and enhance the user experience.

[0124] The software control method based on the digital human system provided by the embodiment of the present invention obtains corresponding documents from the knowledge base through the query information in the keyword information, and classifies the content in the document according to the operation information to obtain output text information, thereby realizing the query function of the digital human system and greatly saving the user's time in organizing and analyzing information.

[0125] Example 4

[0126] Corresponding to the above method embodiment, the embodiment of the present invention provides a digital human system, Figure 4 A schematic diagram of the structure of a digital human system provided by an embodiment of the present invention is shown in FIG. Figure 4 As shown, the digital human system may include:

[0127] Interaction module 401 is used to obtain multimodal input information input by the user, send the multimodal input information to the large language model, and obtain keyword information and operation information fed back by the large language model;

[0128] A linkage control module 402 is configured to control the third-party software to execute the operation information and generate output text information based on the third-party software information in the keyword information;

[0129] The speech synthesis module 403 is used to generate the output speech according to the received output text information;

[0130] The image generation module 404 is used to simulate a digital human and play the output voice through the digital human.

[0131] The embodiments of the present invention provide a digital human-based system that uses a large language model to analyze multimodal input information from users, extracting user needs contained in the multimodal input information to obtain keyword information and operation information. Based on the third-party software information contained in the keywords, the system controls third-party software to execute the operation information, enabling users to control the third-party software using the digital human system. This improves the system's ease of use and enriches its application scenarios. After the third-party software completes the operation information, it generates output speech and plays it through the digital human system, enabling human-computer interaction and enhancing the user's human-computer interaction experience.

[0132] In some embodiments, the system further comprises:

[0133] A voice wake-up module is used to receive a wake-up word voice input by the user, perform feature extraction on the wake-up word voice, identify the user identity information corresponding to the wake-up word voice, and wake up the digital human system if the user identity information is successfully obtained;

[0134] The knowledge graph module is used to obtain at least one document corresponding to the keyword information from the knowledge base according to the query information in the keyword information, and to extract and classify the content in each of the documents according to the operation information to obtain output text information.

[0135] In some embodiments, the interaction module is further configured to:

[0136] Inputting the speech information into a text recognition model to obtain input text information corresponding to the multimodal input information;

[0137] The input text information is sent to the large language model to obtain keyword information and operation information fed back by the large language model.

[0138] In some embodiments, the voice wake-up module is further configured to:

[0139] According to the blacklist and whitelist, query the wake-up permission corresponding to the user identity information;

[0140] When the wake-up authority allows wake-up, the digital human system is woken up.

[0141] In some embodiments, the linkage control module is further configured to:

[0142] Retrieving at least one document corresponding to the keyword information from a knowledge base according to query information in the keyword information;

[0143] According to the operation information, the contents of each document are extracted and classified to obtain output text information.

[0144] In some embodiments, the image generation module is further configured to:

[0145] Inputting the output speech into an emotion simulation model to obtain the emotion corresponding to the output speech;

[0146] According to the emotion, the digital human simulated by the digital human system is controlled to play the output voice.

[0147] The digital human system provided in the embodiment of the present invention has the same implementation principle and technical effects as those in the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference may be made to the corresponding content in the aforementioned method embodiment.

[0148] Example 5

[0149] The embodiment of the present invention further provides an electronic device for running the above-mentioned software control method based on the digital human system; Figure 5 The structural diagram of an electronic device shown in the figure includes a memory 500 and a processor 501, wherein the memory 500 is used to store one or more computer instructions, and the one or more computer instructions are executed by the processor 501 to implement the above-mentioned software control method based on the digital human system.

[0150] Furthermore, Figure 5 The electronic device shown further includes a bus 502 and a communication interface 503 , and the processor 501 , the communication interface 503 and the memory 500 are connected via the bus 502 .

[0151] The memory 500 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is achieved through at least one communication interface 503 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 502 may be an ISA bus, a PCI bus, or an EISA bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0152] The processor 501 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 501 or software instructions. The above processor 501 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as a random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or register. The storage medium is located in the memory 500, and the processor 501 reads the information in the memory 500 and, in conjunction with its hardware, completes the steps of the method of the aforementioned embodiment.

[0153] An embodiment of the present invention further provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the above-mentioned software control method based on the digital human system. The specific implementation can be found in the method embodiment and will not be repeated here.

[0154] The computer program product for performing the software control method based on the digital human system provided in the embodiment of the present invention includes a computer-readable storage medium storing a non-volatile program code executable by a processor. The instructions included in the program code can be used to execute the method described in the previous method embodiment. The specific implementation can be found in the method embodiment and will not be repeated here.

[0155] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0156] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. There may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, indirect coupling or communication connection of devices or units, which may be electrical, mechanical or other forms.

[0157] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0158] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0159] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0160] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A software control method based on a digital human system, applied to a digital human system, characterized in that: include: Get multimodal input information entered by the user; Sending the multimodal input information to a large language model to obtain keyword information and operation information fed back by the large language model; Based on the third-party software information in the keyword information, control the third-party software to execute the operation information and generate output text information; wherein, the third-party software refers to software independent of the digital human system; the operation information refers to the operation that needs to be performed by the digital human system, which is used to match the functions of the third-party software by reading the software's documentation, API interface instructions, or the operation mapping model pre-established by the machine learning algorithm, and realize the functions of the third-party software through the API or simulating user operations; Inputting the output text information into a speech synthesis module to obtain output speech; Based on the output voice, the digital human simulated by the digital human system is controlled to play the output voice.

2. The method according to claim 1, characterized in that The multimodal input information includes voice information; and sending the multimodal input information to the large language model to obtain keyword information and operation information fed back by the large language model includes: Inputting the speech information into a text recognition model to obtain input text information corresponding to the multimodal input information; The input text information is sent to the large language model to obtain keyword information and operation information fed back by the large language model.

3. The method according to claim 1, characterized in that Before obtaining the multimodal input information entered by the user, it also includes: Receive the wake-up word voice input by the user; Extract features of the wake-up word voice and identify user identity information corresponding to the wake-up word voice; When the user identity information is successfully obtained, the digital human system is awakened.

4. The method according to claim 3, characterized in that The step of waking up the digital human system comprises: According to the blacklist and whitelist, query the wake-up permission corresponding to the user identity information; When the wake-up authority allows wake-up, the digital human system is woken up.

5. The method according to claim 1, wherein After sending the multimodal input information to the large language model and obtaining keyword information and operation information fed back by the large language model, the method further includes: Retrieving at least one document corresponding to the keyword information from a knowledge base according to query information in the keyword information; According to the operation information, the contents of each document are extracted and classified to obtain output text information.

6. The method according to claim 1, characterized in that The step of controlling the digital human simulated by the digital human system to play the output voice based on the output voice includes: Inputting the output speech into an emotion simulation model to obtain the emotion corresponding to the output speech; According to the emotion, the digital human simulated by the digital human system is controlled to play the output voice.

7. A digital human system, characterized in that: Executing the software control method based on the digital human system according to any one of claims 1 to 6, the system comprises: An interaction module is used to obtain multimodal input information input by the user, send the multimodal input information to the large language model, and obtain keyword information and operation information fed back by the large language model; A linkage control module, configured to control the third-party software to execute the operation information based on the third-party software information in the keyword information, and to generate output text information; wherein the third-party software refers to software independent of the digital human system; the operation information refers to the operation that needs to be executed by the digital human system, and is configured to match the functions of the third-party software by reading the software's documentation, API interface instructions, or an operation mapping model pre-established by a machine learning algorithm, and to implement the functions of the third-party software through the API or by simulating user operations; A speech synthesis module, configured to generate the output speech according to the received output text information; The image generation module is used to simulate a digital human and play the output voice through the digital human.

8. The system according to claim 7, characterized in that Also includes: A voice wake-up module is used to receive a wake-up word voice input by the user, perform feature extraction on the wake-up word voice, identify the user identity information corresponding to the wake-up word voice, and wake up the digital human system if the user identity information is successfully obtained; The knowledge graph module is used to obtain at least one document corresponding to the keyword information from the knowledge base according to the query information in the keyword information, and to extract and classify the content in each of the documents according to the operation information to obtain output text information.

9. An electronic device, characterized in that: The system comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the software control method based on the digital human system according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions. When the computer-executable instructions are called and executed by the processor, the computer-executable instructions prompt the processor to implement the software control method based on the digital human system according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Control method of smart home system and smart home system

    CN114220442A

  • Efficient digital human interaction system fusing multi-modal information

    CN118897887A

  • Content generation method and device based on large model, and electronic equipment

    CN119180890A