Interaction method and related equipment

By optimizing resource allocation through a cloud management platform and utilizing CDN nodes to provide audio and video output, the problem of wasted computing resources in user interaction with digital humans has been solved, improving resource utilization efficiency and user experience.

CN122064264APending Publication Date: 2026-05-19HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
Filing Date
2024-11-18
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

In the process of users interacting with digital humans, computing resources are wasted in large quantities, resulting in low resource utilization efficiency.

Method used

By managing computing resources through a cloud management platform, audio and video output can be provided using CDN nodes that are close to the user, reducing dependence on computing nodes and calling computing nodes for inference only when necessary.

Benefits of technology

It improved resource utilization efficiency, reduced the waste of computing resources, and enhanced the user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122064264A_ABST
    Figure CN122064264A_ABST
Patent Text Reader

Abstract

The invention provides an interaction method and related equipment. The method is applied to a cloud management platform. The method comprises the following steps: receiving first input information of a user; in response to the first input information, determining first text information; sending indication information used for indicating the first node to provide first output information for the user to the first node, wherein a first database is stored in the first node; and when the first database does not include the first output information, receiving second indication information from the first node, and sending indication information used for indicating the second node to determine the first output information according to the first text information to the second node so as to provide the first output information for a user. According to the method, the first node can be instructed to query the database to obtain the corresponding audio and / or video, and calculation resources in the second node do not need to be used for reasoning each time to obtain the corresponding audio and / or video, so that resource waste in the second node can be reduced, and the efficiency of providing answers for users is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing, and more specifically, to an interaction method, a cloud management platform, a cluster of computing devices, a computer program product, and a computer-readable storage medium. Background Technology

[0002] With the development of artificial intelligence (AI) technology, the application scenarios of digital humans are becoming increasingly widespread. A digital human is a digitized human figure created using digital technology, closely resembling a human. Digital human interaction services allow users to interact with digital humans in real time, providing personalized and intelligent services. There are various ways for digital humans to interact with users, such as through voice, text, and images. During the interaction, the computing node receives user input and provides responses based on it. This computing node determines the corresponding answer based on the user's input and uses the digital human's voice and / or image to perform reasoning, thereby providing the user with an audio or video response corresponding to the user's input. That is, each user consumes the computing resources of the computing node corresponding to the digital human they are interacting with. Furthermore, even if the user does not input a question or the computing node does not perform reasoning based on the question, the computing resources of that node are still occupied by the user, leading to significant resource waste.

[0003] Therefore, how to reduce resource waste during user interaction with digital humans has become an urgent problem to be solved. Summary of the Invention

[0004] This application provides an interaction method, a cloud management platform, a computing device cluster, a computer program product, and a computer-readable storage medium that can reduce resource waste during user interaction with digital humans.

[0005] Firstly, an interaction method is provided. This method is applied to a cloud management platform for managing infrastructure providing cloud services, the infrastructure including at least one computing node. The method includes: receiving first input information from a user, the first input information being used to request an answer corresponding to the user's input; responding to the first input information, determining first text information, the first text information being used to indicate the answer text corresponding to the user's input; sending first instruction information to a first node, the first instruction information including the first text information, the first instruction information being used to instruct the first node to provide first output information to the user, the first node storing a first database, the first database including at least one audio segment and / or at least one video segment of a first digital human, the first output information including at least one of the following: audio corresponding to the first text information obtained using the voice of the first digital human, video corresponding to the first text information obtained using the image of the first digital human; or, if the first output information is not included in the first database, receiving second instruction information from the first node and sending third instruction information to a second node to provide the first output information to the user using an interactive service, the second node running the interactive service, the second instruction information being used to indicate that the first output information is not included in the first database, and the third instruction information being used to instruct the second node to determine the first output information based on the first text information.

[0006] In this embodiment, during user interaction with the digital human, after the user inputs a question, the cloud management platform generates a corresponding answer text based on the question and sends an instruction to the first node based on the answer text. This allows the first node to query a first database to obtain the corresponding audio and / or video by the first node, and then provide it to the user. Since the cloud management platform can instruct the first node to provide the user with the corresponding audio and / or video without needing to use the computing resources of the second node to infer the answer text each time, it reduces the waste of computing resources in the second node and improves the efficiency of providing answers to the user, thereby enhancing the user experience.

[0007] In some embodiments, the audio obtained using the voice of the first digital human corresponding to the first text information includes: audio of the first text information being read aloud using the voice of the first digital human.

[0008] In some embodiments, a video corresponding to the first text information obtained using the image of the first digital human includes the image of the first digital human, and the lip movements of the first digital human in the video match the lip movements of the pronunciation of the first text information. Alternatively, a video corresponding to the first text information obtained using the image of the first digital human includes the image of the first digital human and audio corresponding to the first text information obtained using the voice of the first digital human, and the lip movements of the first digital human in the video match the lip movements of the pronunciation of the first text information.

[0009] In conjunction with the first aspect, in some implementations, the first node belongs to the content delivery network (CDN).

[0010] In this embodiment, the cloud management platform can send instruction information to CDN nodes that are close to the user's device, so that the CDN nodes can directly provide the user with the first output information, thereby improving the efficiency of providing answers to the user.

[0011] In conjunction with the first aspect, in some implementations, if the first output information is not included in the first database, a third instruction information is sent to the second node, which belongs to at least one computing node; the first output information is received from the second node; and the first output information is provided to the user.

[0012] In this embodiment of the application, when the first database does not include output information, the cloud management platform can call the computing node managed by the cloud management platform, so that the computing node determines the corresponding first output information based on the first text information, thereby providing the first output information to the user.

[0013] In conjunction with the first aspect, in some implementations, second output information is provided to the user before and / or after the first output information is provided. This second output information includes at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human.

[0014] In some embodiments, the default audio includes audio when the first digital human is not interacting with the user, such as the voice of the first digital human in a silent state (i.e., not speaking).

[0015] In some embodiments, the default video includes at least one image frame when the first digital human is not interacting with the user, the at least one image frame including, for example, the facial expressions and / or postures of the first digital human in a silent state (i.e., a non-speaking state).

[0016] In this embodiment, when the first digital human does not answer the user's question, the cloud management platform can provide the user with audio and / or video of the first digital human in a silent state, thereby improving the simulation level of the digital human and enhancing the user's interactive experience. Simultaneously, since the cloud management platform can provide the user with pre-prepared default audio and / or default video of the first digital human, there is no need to continuously occupy the computing resources corresponding to the first digital human in a silent state, thus reducing resource waste.

[0017] In some implementations, when the second output information is provided to the user before the first output information is provided, the last frame of the default video in the second output information is an adjacent frame to the first frame of the video in the first output information.

[0018] In this embodiment of the application, by connecting the last frame of the default video in the second output information with the first frame of the video in the first output information, the user can avoid problems such as discontinuous video or frame skipping, thereby improving the user's interactive experience.

[0019] In conjunction with the first aspect, in some implementations, the first configuration information of the tenant is received. The first configuration information is used to configure at least one of the following: whether to store the default audio set and / or default video set of the first digital human in the user's device, whether to enable the function of sending the first instruction information to the first node; and according to the first configuration information, to provide the user with a first access link, which is used to provide interactive services to the user.

[0020] In some embodiments, the tenant is the provider of the interaction service for the first digital human, and the user is the user of the interaction service for the first digital human.

[0021] In this embodiment, the cloud management platform can provide users with corresponding access links based on the tenant's configuration. When a user uses the digital human's interactive service through the access link, the platform can determine whether to provide the user with the digital human's default audio and / or default video, and / or whether to use the method provided in this embodiment, thereby avoiding the need to continuously occupy computing resources in the digital human's silent state and reducing resource waste.

[0022] In conjunction with the first aspect, in some implementations, when the first output information includes a video corresponding to the first text information obtained using the image of the first digital human, the first output information is determined based on the third output information and the first text information. The third output information includes a base video obtained using the image of the first digital human, or the third output information includes the last frame of a default video obtained using the image of the first digital human and the base video obtained using the image of the first digital human.

[0023] In conjunction with the first aspect, in some implementations, the third output information includes: the last frame of the default video obtained using the image of the first digital human, the first frame of the base video obtained using the image of the first digital human, the reference frame of the base video obtained using the image of the first digital human, and the base video obtained using the image of the first digital human, wherein the reference frame of the base video belongs to the base video.

[0024] In some embodiments, the base video is used to indicate the image of the first digital human, the image of the first digital human including at least one of the following: appearance, expression, posture, movement, etc.

[0025] In this embodiment of the application, when providing a user with a default video obtained using the image of the first digital human, the last frame of the default video can be used to determine the first output information, thereby ensuring that the default video and the first output information are connected smoothly, avoiding problems such as frame skipping or discontinuous video footage perceived by the user, and thus improving the user's interactive experience.

[0026] In conjunction with the first aspect, in some implementations, the third output information is determined based on the first identification information, which includes the identification information of the first digital human and / or the identification information of the second output information, which includes at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human.

[0027] In this embodiment of the application, the second node determines the voice and / or image of the first digital human based on the first identification information, thereby determining the third output information, which in turn facilitates the determination of the first output information.

[0028] In conjunction with the first aspect, in some implementations, when the first database includes first output information, the index of the first output information in the first database is a first index, which is determined based on first identification information and first text information. The first identification information includes the identification information of the first digital human and / or the identification information of the second output information. The second output information includes at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human.

[0029] In this embodiment, determining the first index using the first identification information and the first text information ensures that the first output information retrieved from the first database includes the relevant audio and / or video of the first digital human. Furthermore, when the first identification information includes the identification information of the second output information, querying the first database using the first index ensures a smooth transition between the first and second output information retrieved from the first database, thereby avoiding issues such as frame skipping or discontinuous video footage perceived by the user.

[0030] In conjunction with the first aspect, in some implementations, the first output information is sent to the first node when the first database does not include the first output information.

[0031] In this embodiment of the application, when there is no audio and / or video corresponding to the first text information in the first database, the cloud management platform can also send the first output information to the first node after obtaining the first output information, so that when the first input information is received again in the future, the corresponding audio and / or video can be directly provided to the user through the first node, thereby reducing the waste of resources and improving the efficiency of providing answers to users.

[0032] Secondly, an interaction method is provided. This method includes a cloud management platform, a first node, and a second node. The cloud management platform receives first input information from a user, which requests a response corresponding to the user's input. In response to the first input information, the cloud management platform determines first text information, which indicates the text of the response corresponding to the user's input. The cloud management platform sends first instruction information to the first node, which includes the first text information and instructs the first node to provide first output information to the user. The first node stores a first database, which includes at least one audio segment and / or at least one video segment of a first digital human. The first output information includes at least one of the following: audio corresponding to the first text information obtained using the voice of the first digital human, or video corresponding to the first text information obtained using the image of the first digital human. The first node determines whether the first database includes the first output information based on the first instruction information. If the first database includes the first output information, the first node provides the first output information to the user. If the first database does not include the first output information, the first node sends second instruction information to the cloud management platform; the cloud management platform then sends third instruction information to the second node and provides the first output information to the user. The second indication information is used to indicate that the first database does not include the first output information. The third indication information is used to instruct the second node to determine the first output information based on the first text information.

[0033] In this embodiment, during user interaction with the digital human, after the user inputs a question, the cloud management platform generates a corresponding answer text based on the question and sends an instruction to the first node based on the answer text. This allows the first node to query a first database to obtain the corresponding audio and / or video by the first node, and then provide it to the user. Since the cloud management platform can instruct the first node to provide the user with the corresponding audio and / or video without needing to use the computing resources of the second node to infer the answer text each time, it reduces the waste of computing resources in the second node and improves the efficiency of providing answers to the user, thereby enhancing the user experience.

[0034] Thirdly, a cloud management platform is provided. This cloud management platform includes modules for implementing the first aspect or any possible implementation thereof.

[0035] Fourthly, this application provides a computing device cluster, including at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method in the first aspect or any possible implementation of the first aspect.

[0036] Fifthly, this application provides a computer program product containing instructions that, when executed by a cluster of computer devices, cause the cluster of computer devices to perform the method described in the first aspect or any possible implementation thereof.

[0037] In a sixth aspect, this application provides a computer-readable storage medium including computer program instructions that, when executed by a cluster of computing devices, perform the method described in the first aspect or any possible implementation thereof.

[0038] In a seventh aspect, a chip system is provided, the chip system including logic circuitry for coupling with an input / output interface, through which data is transmitted to perform the method described in the first aspect or any possible implementation thereof. Attached Figure Description

[0039] Figure 1 This is a schematic structural diagram of an interactive system according to an embodiment of this application.

[0040] Figure 2 This is a schematic structural diagram of an interaction method according to another embodiment of this application.

[0041] Figure 3This is a schematic flowchart of an interaction method according to an embodiment of this application.

[0042] Figure 4 This is a schematic flowchart of an interaction method according to another embodiment of this application.

[0043] Figure 5 This is a schematic diagram of a first graphical interface according to an embodiment of this application.

[0044] Figure 6 This is a schematic diagram of a first graphical interface according to another embodiment of this application.

[0045] Figure 7 This is a schematic structural diagram of a cloud management platform according to an embodiment of this application.

[0046] Figure 8 This is a schematic structural diagram of a computing device according to an embodiment of this application.

[0047] Figure 9 This is a schematic structural diagram of a computing device cluster according to an embodiment of this application.

[0048] Figure 10 This is a schematic diagram illustrating the connection between computing devices 800A and 800B via a network according to an embodiment of this application. Detailed Implementation

[0049] The technical solutions in this application will now be described with reference to the accompanying drawings.

[0050] This application will present various aspects, embodiments, or features relating to a system comprising multiple devices, components, modules, etc. It should be understood and appreciated that individual systems may include additional devices, components, modules, etc., and / or may not include all the devices, components, modules, etc. discussed in conjunction with the accompanying drawings. Furthermore, combinations of these approaches are also possible.

[0051] Furthermore, in the embodiments of this application, the words "exemplary," "for example," etc., are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" in the embodiments of this application should not be construed as being better or more advantageous than other embodiments or design schemes. Specifically, the use of the term "exemplary" is intended to present the concept in a concrete manner.

[0052] The business scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0053] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0054] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0055] To facilitate understanding of the technical solutions in the embodiments of this application, some terms involved in the embodiments of this application are introduced below.

[0056] 1. Cloud Management Platform

[0057] A cloud management platform is used to manage the infrastructure that provides cloud services. This infrastructure includes at least one cloud data center, and each cloud data center includes at least one compute node. Each compute node is, for example, a container, a virtual machine (VM), a server, a computing device, etc. Each compute node includes cloud service resources, thereby providing corresponding cloud services to tenants.

[0058] The cloud management platform can be located in a cloud data center and can provide access interfaces (such as user interfaces or application program interfaces, APIs). Tenants can remotely access these interfaces to register a cloud account and password on the cloud management platform and log in. After successful authentication of the cloud account and password, the tenant can further select and purchase compute nodes (such as containers, virtual machines, servers, computing devices, etc.) with specific specifications (e.g., processors, memory, disks) on the cloud management platform. After successful purchase, the cloud management platform provides a remote login account and password for the purchased compute node, allowing the tenant to remotely log in to the compute node and install and run their applications on it. Therefore, tenants can create, manage, log in to, and operate compute nodes in the cloud data center through the cloud management platform. These compute nodes can also be referred to as Elastic Compute Service (ECS) or Elastic Instances (different cloud service providers may use different names).

[0059] It should be understood that cloud service tenants can be individuals, businesses, schools, hospitals, government agencies, etc.

[0060] The cloud management platform's functions include, but are not limited to, a user console, compute management services, network management services, storage management services, authentication services, and image management services. The user console provides an interface or API for interaction with tenants. The compute management service manages servers running virtual machines and containers, as well as bare metal servers. The network management service manages network services (such as gateways and firewalls). The storage management service manages storage services (such as data bucket services). The authentication service manages tenant account passwords. The image management service manages virtual machine images.

[0061] 2. Digital Human

[0062] A digital human is a digitized humanoid created using digital technology that closely resembles a human. Digital humans can interact with users in real time, such as through voice, images, and text, thereby providing personalized and intelligent services. The interaction process involves the user inputting a question (e.g., text or voice), and the computing node using the digital human's voice and / or image to provide the user with real-time audio and / or video responses.

[0063] 3. Content Delivery Network (CDN)

[0064] CDN is a distributed network architecture that caches content on the storage node closest to the user's device by deploying storage nodes around the world, thereby achieving fast and efficient content distribution.

[0065] Figure 1 This is a schematic structural diagram of the cloud scenario 100 provided in the embodiments of this application. Figure 1 The cloud scenario 100 includes a cloud management platform 110. This cloud management platform 110 manages the infrastructure providing cloud services, including at least one data center (e.g., data center 120). The data center 120 includes at least one compute node cluster, each cluster containing at least one compute node, such as a container, virtual machine, server, or computing device. For example, the data center 120 includes compute node clusters 130 and / or 140, where cluster 130 includes compute nodes 131 and / or 132, and cluster 140 includes compute nodes 141 and / or 142. Tenants can apply for access to resources in the data center 120 through the cloud management platform 110, thereby utilizing these resources to run digital human interaction services and provide digital human interaction services to users. The tenant is a public cloud tenant who has registered a public cloud account and purchased public cloud resources. The user is a user who interacts with a digital human using the digital human interaction services provided by the tenant. The cloud management platform 110 is used to execute the methods provided in the embodiments of this application, for example... Figure 3 or Figure 4 The method in the middle.

[0066] In some embodiments, multiple computing nodes in a computing node cluster are directly connected or connected via a network, such as a wide area network or a local area network.

[0067] In some embodiments, the data center 120 may further include at least one storage node cluster, each storage node cluster including at least one storage node. This application embodiment does not limit the type of storage node; for example, the storage node may be a centralized storage node or a distributed storage node. Storage nodes are used to store tenant data, or to store data required and / or generated when performing the methods in this application embodiment. Exemplarily, multiple storage nodes in a storage node cluster are directly connected or connected via a network, such as a wide area network (WAN) or a local area network (LAN).

[0068] Figure 2 This is a schematic structural diagram of the interactive system 200 provided in the embodiments of this application. Figure 2 The interactive system 200 includes a user module 210, a task management module 220, a large model module 230, a storage module 240, and an inference module 250.

[0069] In some embodiments, the various modules in the interactive system 200 are managed by a cloud management platform, such as... Figure 1The cloud management platform 110 is included. The user module 210 can be deployed on a user's computing device, such as a laptop, server, desktop computer, or smart device. The task management module 220, large model module 230, storage module 240, and inference module 250 can be deployed on a server, which may include, for example, […]. Figure 1 The computing nodes in the process.

[0070] In some embodiments, the storage module 240 is deployed in a first node, and the inference module 250 is deployed in a second node. The first node and the second node are different nodes.

[0071] For example, the first node belongs to the CDN. That is, the first node is the CDN storage node corresponding to the user's device. In other words, the first node is a CDN storage node that is physically close to the user's device.

[0072] For example, the physical distance between the first node and the user's device is less than or equal to the physical distance between the second node and the user's device.

[0073] For example, the second node is a computing node managed by the cloud management platform. The second node runs a digital human interaction service.

[0074] For example, the first node and the second node may belong to the same or different cloud vendors. If the first node and the second node belong to the same cloud vendor, they can be managed by the same cloud management platform.

[0075] User module 210 is used to receive input information from users or tenants, such as first request information, first input information, or first configuration information. User module 210 is also used to transmit the input information to task management module 220 and receive response information from task management module 220. This response information may include, for example, first text information, a first access link, or first output information. User module 210 is also used to provide output information to users or tenants, such as first text information, a first access link, first output information, and second output information. The first request information, first input information, first configuration information, first text information, first access link, first output information, and second output information are described in [reference needed]. Figure 3 or Figure 4 The description in the text.

[0076] In some embodiments, the user module 210 is used to provide a graphical interface for users or tenants, enabling users or tenants to perform selection, input, or upload operations in the graphical interface, thereby transmitting input information to the user module 210. Alternatively, the user module 210 is used to provide a graphical interface for users or tenants, the graphical interface including output information, enabling users or tenants to obtain the output information through the graphical interface.

[0077] In some embodiments, the user module 210 is further configured to send first instruction information to the storage module 240. The first instruction information includes first text information, which instructs the storage module 240 to provide first output information to the user. The user module 210 is also configured to receive the first output information from the storage module 240.

[0078] The task management module 220 receives first input information from the user module 210 and determines first text information based on the first output information. The task management module 220 then transmits the first text information to the user module 210.

[0079] In some embodiments, the task management module 220 queries a second database based on the first output information to determine the first text information. The second database includes at least one piece of text information. Alternatively, the task management module 220 transmits the first input information to the large model module 230, thereby receiving the first text information from the large model module 230.

[0080] In some embodiments, when the first input information is text information, the task management module 220 directly transmits the first input information to the large model module 230. When the first input information is voice information, the task management module 220 converts the voice information into corresponding text information and transmits the text information to the large model module 230.

[0081] For example, the task management module 220 converts voice information into corresponding text information using automatic speech recognition (ASR) technology.

[0082] In some embodiments, the task management module 220 is further configured to receive second indication information from the storage module 240, the second indication information including first text information, the second indication information indicating that the storage module 240 does not contain first output information. The task management module 220 sends third indication information to the inference module 250, the third indication information indicating that the inference module 250 determines the first output information based on the first text information. The task management module 220 is further configured to receive the first output information from the inference module 250 and send the first output information to the storage module 240 and / or the user module 210.

[0083] In some embodiments, the task management module 220 is further configured to receive first configuration information from the user module 210. The task management module 220 is further configured to determine a first access link based on the first configuration information and send it to the user module 210.

[0084] In some embodiments, the task management module 220 is further configured to receive a first request message from the user module 210, the first request message being a request to interact with the digital human. In response to the first request message, the task management module 220 is further configured to send second output information to the user module 210.

[0085] The large model module 230 is used to receive first input information from the task management module 220, wherein the first input information is text information. The large model module 230 is also used to obtain first text information based on the first input information and send the first text information to the task management module 220. Alternatively, the large model module 230 is used to receive text information corresponding to the first input information from the task management module 220, wherein the first input information is voice information. The large model module 230 is also used to obtain first text information based on the text information corresponding to the first input information and send the first text information to the task management module 220.

[0086] In some embodiments, the large model module 230 includes a first large model, which is used to determine the corresponding answer text (e.g., first text information) based on user input information (e.g., first input information). The first large model is a large language model (LLM), or the first large model is obtained by training an LLM.

[0087] Storage module 240 is used to receive first instruction information from user module 210. Storage module 240 is also used to determine, based on the first text information, whether the first database includes first output information. See [link to first database]. Figure 3 or Figure 4 The description in the document is as follows: When the first output information is included in the first database, the storage module 240 sends the first output information to the user module 210. When the first output information is not included in the first database, the storage module 240 sends second instruction information to the task management module 220.

[0088] In some embodiments, the storage module 240 receives first output information from the task management module 220.

[0089] In some embodiments, the storage module 240 stores a first database.

[0090] In some embodiments, the storage module 240 determines the first index based on the first text information, or based on the first identification information and the first text information. The first identification information includes the identification information of the first digital human and / or the identification information of the second output information. See [link to documentation]. Figure 3 or Figure 4The storage module 240 determines whether the first database contains the first output information based on the first index. In other words, the storage module 240 determines whether the first database contains the information corresponding to the first index. If the first database contains the information corresponding to the first index, the information corresponding to the first index is used as the first output information. If the first database does not contain the information corresponding to the first index, the first text information and the first identification information are sent to the task management module 220.

[0091] For example, the storage module 240 is further configured to receive first identification information from the user module 210. That is, the user module 210 is further configured to send the first identification information to the storage module 240.

[0092] In some embodiments, after receiving the first output information from the task management module 220, the storage module 240 stores the first output information in the first database. The index of the first output information is the first index.

[0093] The inference module 250 is used to receive third instruction information from the task management module 220. The inference module 250 is also used to determine first output information based on the first text information and send it to the task management module 220.

[0094] In some embodiments, the inference module 250 includes at least one computing node. The at least one computing node includes, for example, […]. Figure 1 The computing nodes in the system. Each computing node is used to perform a reasoning task, which includes generating audio and / or video of the digital human using its voice and / or image. That is, each computing node runs an interactive service for the digital human. In other words, each computing node is used to infer, based on the answer text, the voice and / or image of the digital human to obtain audio and / or video responses in the voice and / or image of the digital human. At any given time, each computing node can perform one or more reasoning tasks. In other words, each computing node includes one or more sets of resources, each set of resources being used to perform one reasoning task. Each set of resources includes computing resources and storage resources.

[0095] For example, each computing node in the inference module 250 sends first statistics to the task management module 220, which indicates the number of available resource groups in each computing node. The available resource group is a set of resources in the computing node for which no inference tasks are being performed. The task management module 220 assigns inference tasks to each computing node based on the first statistics. For example, the number of inference tasks assigned by the task management module 220 to each computing node is less than or equal to the number of available resource groups in each computing node.

[0096] In some embodiments, the computing node in the inference module 250 determines third output information based on the first identification information. The third output information includes at least one of the following: audio obtained using the voice of the first digital human, and video obtained using the image of the first digital human. The computing node is also configured to determine the first output information based on the first text information and the third output information.

[0097] For example, the computing node in the inference module 250 inputs the first text information into the first inference model and obtains the output of the first inference model, i.e., obtains the first output information. Alternatively, the computing node in the inference module 250 inputs the third output information and the first text information into the second inference model and obtains the output of the first inference model, i.e., obtains the first output information. The third output information, the first inference model, and the second inference model are described in [reference needed]. Figure 3 or Figure 4 The description in the text.

[0098] In some embodiments, the computing node in the inference module 250 determines first output information based on text-to-speech (TTS) technology, the first output information including audio corresponding to the first text information obtained using the voice of a digital human.

[0099] During the interaction between users and digital humans Figure 2 The interactive system 200 can generate corresponding answer text based on the user's input question, and query a database based on the answer text to obtain the corresponding audio and / or video, which is then provided to the user. Since the interactive system 200 can provide the corresponding audio and / or video to the user by querying the database, without needing to use computing resources to reason from the answer text each time, it reduces the waste of computing resources and improves the efficiency of providing answers to the user, thereby enhancing the user experience. Furthermore, if the database does not contain the corresponding audio and / or video for the answer text, the interactive system 200 can store the corresponding audio and / or video after obtaining it, so that when the first input information is received again subsequently, the corresponding audio and / or video answer can be directly provided to the user, further reducing resource waste and improving the efficiency of providing answers to the user.

[0100] Figure 3 This is a schematic flowchart of the interaction method provided in the embodiments of this application. Figure 3 The methods described can be executed by the cloud management platform, for example... Figure 1 The cloud management platform 110, or the computing nodes managed by the cloud management platform, can execute it, for example... Figure 1 The computing nodes in the process. Figure 3The method includes the following steps.

[0101] 310, Receive the user's first input information.

[0102] The cloud management platform receives the user's initial input information, which is used to request the response corresponding to the user's input. This user is the one using the digital human interaction service to interact with the digital human.

[0103] In some embodiments, the first input information includes a question entered by the user. Exemplarily, the first input information is text information or voice information.

[0104] In some embodiments, the cloud management platform provides a first graphical interface to the user, allowing the user to select, input, or upload first input information within the first graphical interface. The specific form of this first graphical interface is not limited in the embodiments of this application.

[0105] 320, in response to the first input information, determine the first text information.

[0106] The cloud management platform determines the first text information based on the first input information. This first text information indicates the answer text corresponding to the user's input. In other words, the first text information includes the answer text corresponding to the first input information.

[0107] In some embodiments, after obtaining the first input information, the cloud management platform queries a second database based on the first input information to determine whether the second database includes the first text information. The second database includes at least one piece of text information. If the second database includes the first text information, the cloud management platform directly determines the first text information. Alternatively, if the second database does not include the first text information, the cloud management platform inputs the first input information into a first large model and obtains the output of the first large model, i.e., obtains the first text information. Alternatively, after obtaining the first input information, the cloud management platform inputs the first input information into the first large model and obtains the output of the first large model, i.e., obtains the first text information.

[0108] For example, the first major model is used to determine the corresponding answer text based on the user's input information (i.e., the question entered by the user). The first major model is an LLM, or the first major model is obtained by training an LLM.

[0109] For example, before step 320, the cloud management platform obtains the first large model that has already been trained. Alternatively, before step 320, the cloud management platform trains the LLM to obtain the first large model.

[0110] In some embodiments, when the first input information is voice information, the first input information includes audio of a text message read aloud by the user's voice. The cloud management platform determines the text information corresponding to the first input information. That is, the cloud management platform determines the text information included in the audio of the first input information. The cloud management platform inputs the text information corresponding to the first input information into the first large model to obtain the first text information. When the first input information is text information, the cloud management platform directly inputs the first input information into the first large model to obtain the first text information.

[0111] For example, the cloud management platform determines the second index based on the first input information. For example, the second index is either the hash value of the first input information or a feature vector of the first input information.

[0112] For example, the cloud management platform inputs the first input information into the first feature extraction model to obtain the feature vector of the first input information. Specifically, the cloud management platform inputs the first input information into the first feature extraction model to obtain the output of the first feature extraction model, i.e., obtains the feature vector of the first input information. The first feature extraction model can be a model trained using deep learning based on a first training dataset. The first training dataset may include at least one piece of text information, at least one feature vector, and a mapping relationship between at least one piece of text information and at least one feature vector.

[0113] For example, before determining the feature vector of the first input information, the cloud management platform can obtain a pre-trained first feature extraction model. Alternatively, before determining the feature vector of the first input information, the cloud management platform can obtain a first training dataset and train the model based on the first training dataset to obtain a pre-trained first feature extraction model.

[0114] For example, the cloud management platform queries the second database based on the second index to determine whether the second database contains the information corresponding to the second index. The information corresponding to the second index is the first text information. The second database includes at least one text information and an index for each text information. The method for determining the index of each text information is similar to the method for determining the second index, and will not be repeated here.

[0115] For example, the cloud management platform determines the text information corresponding to the first input information based on ASR technology.

[0116] 330, Send the first instruction information to the first node.

[0117] The cloud management platform sends a first instruction message to the first node, and correspondingly, the first node receives the first instruction message from the cloud management platform. This first instruction message includes first text information, which instructs the first node to provide first output information to the user. The first node stores a first database, which includes at least one audio segment and / or at least one video segment of the first digital human. The first output information includes at least one of the following: audio corresponding to the first text information obtained using the voice of the first digital human, or video corresponding to the first text information obtained using the image of the first digital human. The first digital human is a digital human that the user requests to interact with.

[0118] In some embodiments, the audio of the first digital human is audio of the first digital human reading a text message aloud using its voice. The video of the first digital human includes at least one image frame, or the video of the first digital human includes audio and at least one image frame. In other words, the video of the first digital human may or may not include audio. When the video of the first digital human includes at least one image frame, the video of the first digital human includes at least one image frame of the first digital human making facial expressions and / or gestures using its image. The image of the first digital human includes at least one of the following: the shape of the first digital human, the facial expressions of the first digital human, the gestures of the first digital human, the actions of the first digital human, etc. When the video of the first digital human includes audio and at least one image frame, the video of the first digital human includes: audio of the first digital human reading a text message aloud using its voice, and at least one image frame of the first digital human making some facial expressions and / or gestures using its image.

[0119] In some embodiments, the audio corresponding to the first text information obtained using the voice of the first digital human includes: audio of the first text information being read aloud using the voice of the first digital human. The video corresponding to the first text information obtained using the image of the first digital human includes the image of the first digital human, and the lip movements of the first digital human in the video match the lip movements used to pronounce the first text information. Alternatively, the video corresponding to the first text information obtained using the image of the first digital human includes the image of the first digital human, and the lip movements of the first digital human in the video match the lip movements used to pronounce the first text information, and the video includes audio corresponding to the first text information obtained using the voice of the first digital human.

[0120] In a video containing first text information obtained using the image of a first digital human, and also including audio information containing the first text information obtained using the voice of the first digital human, the lip movements of the first digital human in the video correspond to the audio. In other words, when the video is played, the lip movements of the first digital human in the video match the lip movements corresponding to the sounds in the audio.

[0121] In some embodiments, the first node belongs to a CDN. That is, the first node is the CDN storage node corresponding to the user's device. In other words, the first node is a CDN storage node that is physically close to the user's device.

[0122] Optionally, after receiving the first instruction information, the first node determines whether the first output information is included in the first database based on the first text information. If the first output information is included in the first database, the first node provides the first output information to the user. If the first output information is not included in the first database, the first node sends second instruction information to the cloud management platform. This second instruction information is used to indicate that the first output information is not included in the first database.

[0123] In some embodiments, the cloud management platform provides a second graphical interface to the user, which includes first output information provided by the first node, allowing the user to view the first output information in the second graphical interface. The specific form of this second graphical interface is not limited in the embodiments of this application.

[0124] For example, the second graphical interface may be the same as or different from the first graphical interface.

[0125] In some embodiments, the first node determines the first index based on the first text information. Alternatively, the first node determines the first index based on the first text information and the first identification information. The first identification information includes the identification information of the first digital human and / or the identification information of the second output information. The second output information includes at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human. See [link to documentation] for the second output information and its identification information. Figure 4 The description in the text.

[0126] For example, the first instruction information includes first text information. Alternatively, the first instruction information includes first text information and first identification information.

[0127] For example, the first database includes an index for each audio or video segment. The method for determining the index in the first database is not limited in this application embodiment. For example, the index in the first database can be any of the following: a hash value of the text information corresponding to each audio or video segment, a feature vector of the text information corresponding to each audio or video segment, a hash value of the first character sequence corresponding to each audio or video segment, or a feature vector of the first character sequence corresponding to each audio or video segment. The first character sequence corresponding to each audio segment includes: the digital human identification information corresponding to each audio segment and the text information corresponding to each audio segment. The first character sequence corresponding to each video segment includes: the digital human identification information corresponding to each video segment and the text information corresponding to each video segment. Alternatively, the first character sequence corresponding to each video segment includes: the digital human identification information corresponding to each video segment, the identification information of the second output information corresponding to each video segment, and the text information corresponding to each video segment. Alternatively, the first character sequence corresponding to each video segment includes: the digital human identification information corresponding to each video segment, the identification information of the second output information corresponding to each video segment, and the text information corresponding to each video segment.

[0128] In some embodiments, the cloud management platform provides the user with second output information before and / or after providing the user with the first output information. In other words, when the digital human does not interact with the user (i.e., does not answer the user's questions), the cloud management platform provides the user with default audio obtained using the voice of the first digital human and / or default video obtained using the image of the first digital human, thereby enhancing the realism of the digital human and thus improving the user's interactive experience.

[0129] In some embodiments, the first node queries the first database based on the first index to determine whether the first database includes information indexed as the first index. If the first database includes audio or video with the first index, the first node provides the user with that audio or video, and the cloud management platform does not need to execute step 340. If the first database does not include audio or video with the first index, the first node sends second instruction information to the cloud management platform, causing the cloud management platform to execute step 340.

[0130] 340, if the first output information is not included in the first database, receive the second instruction information from the first node and send the third instruction information to the second node to provide the first output information to the user using the interactive service.

[0131] If the first output information is not included in the first database, the cloud management platform receives a second instruction from the first node. This second instruction indicates that the first output information is not included in the first database. After receiving the second instruction, the cloud management platform sends a third instruction to the second node to provide the first output information to the user using an interactive service. This third instruction instructs the second node to determine the first output information based on the first text information. The second node runs the interactive service, which determines the first output information based on the first text information.

[0132] For example, after the cloud management platform sends the third instruction information to the second node, it receives the first output information from the second node and provides the first output information to the tenant.

[0133] In some embodiments, the second node is a computing node managed by the cloud management platform.

[0134] For example, the physical distance between the first node and the user's device is less than or equal to the physical distance between the second node and the user's device.

[0135] For example, the first node and the second node may belong to the same or different cloud vendors. If the first node and the second node belong to the same cloud vendor, they can be managed by the same cloud management platform.

[0136] Optionally, after obtaining the third instruction information, the second node inputs the first text information into the first inference model to obtain the output of the first inference model, i.e., obtains the first output information. See [link to first inference model]. Figure 4 The description in the text.

[0137] In some embodiments, the third instruction information includes the first text information. Alternatively, the third instruction information includes the first text information and the first identification information.

[0138] Optionally, the first node determines the third output information based on the first identification information. The third output information includes at least one of the following: audio obtained using the voice of the first digital human, and video obtained using the image of the first digital human. The cloud management platform determines the first output information based on the first text information and the third output information.

[0139] In some embodiments, where the first output information includes audio corresponding to the first text information obtained using the voice of the first digital human, the third output information includes base audio obtained using the voice of the first digital human. This base audio is used to indicate the voice of the first digital human, and the text information corresponding to this base audio may be the same as or different from the first text information.

[0140] For example, the base audio included in the third output information belongs to the base audio set of the first digital human, which includes at least one base audio. Each base audio is used to indicate the voice of the first digital human. When the base audio set includes multiple base audios, the content of at least two base audios in the base audio set is different.

[0141] In some embodiments, when the first output information includes a video corresponding to the first text information obtained using the image of the first digital human, the third output information includes a base video obtained using the image of the first digital human, or the third output information includes the last frame of a default video obtained using the image of the first digital human and the base video obtained using the image of the first digital human. The base video is used to indicate the image of the first digital human, and the text information corresponding to the base video may be the same as or different from the first text information. The default video obtained using the image of the first digital human is the default video in the second output information. Including the last frame of the default video in the second output information in the third output information allows for a smoother transition between the video in the first output information and the default video in the second output information, thereby avoiding problems such as frame skipping or video discontinuity perceived by the user, and thus improving the user's visual experience.

[0142] For example, the base video included in the third output information belongs to a base video set of the first digital human, which includes at least one base video. Each base video is used to indicate the image of the first digital human. When the base video set includes multiple base videos, the content of at least two base videos in the base video set is different.

[0143] In some embodiments, when the first output information includes a video corresponding to the first text information obtained using the image of the first digital human, the third output information includes: the last frame of the default video obtained using the image of the first digital human, the first frame of the base video obtained using the image of the first digital human, a reference frame of the base video obtained using the image of the first digital human, and the base video obtained using the image of the first digital human. The reference frame of the base video belongs to the base video, that is, the reference frame of the base video is any frame in the base video.

[0144] In some embodiments, the first node queries a third database based on first identification information to determine the basic audio set and / or basic video set of the first digital human. The third database includes at least one basic audio set and / or basic video set of a digital human. The first digital human belongs to at least one of these at least one digital human. If the third output information includes basic audio obtained using the voice of the first digital human, the cloud management platform selects one or more basic audio samples from the basic audio set of the first digital human to determine the third output information. If the third output information includes basic video obtained using the image of the first digital human, the cloud management platform selects one or more basic videos from the basic video set of the first digital human to determine the third output information.

[0145] In some embodiments, the first node inputs the first text information and the third output information into the second inference model to obtain the output of the second inference model, i.e., obtains the first output information. See [link to second inference model]. Figure 4 The description in the text.

[0146] Optionally, if the first output information includes audio corresponding to the first text information obtained using the voice of the first digital human, the first node converts the first text information into audio corresponding to the first text information obtained using the voice of the first digital human according to TTS technology, thereby determining the first output information.

[0147] Optionally, the cloud management platform provides a second graphical interface to the user, which includes first output information determined by the second node, allowing the user to view the first output information in the second graphical interface. The specific form of this second graphical interface is not limited in this embodiment.

[0148] Optionally, after determining the first output information, the cloud management platform sends the first output information to the first node, causing the first node to store the first output information in the first database. The index of the first output information in the first database is the first index. The method for determining this first index is described in [link to documentation]. Figure 4 The description in the text.

[0149] In some embodiments, the first node performs hierarchical caching of the first output information in the first database. For example, the more times the first output information is provided to the user, the longer the first output information is stored in the first database.

[0150] Optionally, while providing the user with the first output information, the user is also provided with the first text information.

[0151] Optionally, before step 310, the cloud management platform receives the user's first configuration information. This first configuration information is used to configure at least one of the following: whether to store the first digital human's default audio set and / or default video set on the user's device, and whether to enable the function of sending first instruction information to the first node. The first digital human's default audio set includes at least one default audio track. The first digital human's default video set includes at least one default video track.

[0152] Optionally, the cloud management platform provides a first access link to the user based on the first configuration information. This first access link is used to provide interactive services to the user.

[0153] Figure 3 The method described herein allows for the generation of corresponding answer text based on user input of a question during interaction with a digital human. This answer text is then sent to a first node, which in turn queries a first database to retrieve the corresponding audio and / or video, providing it to the user. Since the cloud management platform can instruct the first node to provide the user with the corresponding audio and / or video without repeatedly utilizing the computing resources of a second node to infer the answer text, it reduces wasted computing resources in the second node and improves the efficiency of providing answers to the user, thus enhancing the user experience. Furthermore, if the first database does not contain the corresponding audio and / or video, the cloud management platform can still send it to the first node after obtaining it. This allows the first node to store the audio and / or video, facilitating direct provision of the corresponding audio and / or video to the user upon receiving the same input again, further reducing wasted computing resources in the second node and improving the efficiency of providing answers to the user.

[0154] In some embodiments, Figure 3 The methods in [the document] can be applied to, for example... Figure 4 The interaction method shown. Figure 4 This is a schematic flowchart of the interaction method provided in the embodiments of this application. Figure 4 The method described can be executed by the cloud management platform, the first node, and the second node. The cloud management platform is, for example, a... Figure 1 The cloud management platform 110 is described above. The first node and the second node are as described in the above embodiments. Figure 4 The method includes the following steps.

[0155] 401, Receive the tenant's first configuration information.

[0156] The cloud management platform receives first configuration information from the tenant. This first configuration information is used to configure at least one of the following: whether to store the default audio set and / or default video set of the first digital human on the user's device, and whether to enable the function of sending first instruction information to the first node. The default audio set of the first digital human includes at least one default audio of the first digital human. The default video set of the first digital human includes at least one default video of the first digital human.

[0157] In some embodiments, the cloud management platform provides a third graphical interface to the tenant, allowing the tenant to select, input, or upload first configuration information, thereby transmitting the first configuration information to the cloud management platform. The specific form of this third graphical interface is not limited in the embodiments of this application.

[0158] 402, Based on the first configuration information, determine the first access link.

[0159] The cloud management platform determines the first access link based on the tenant's initial configuration information. This first access link is used to provide interactive services to the user, including services for interacting with the first digital human. The tenant is the provider of the interactive services for the first digital human, and the user is the consumer of the interactive services for the first digital human.

[0160] In some embodiments, the first access link may be accessed by one or more users, each of whom can use the first digital human's interactive services through the first access link.

[0161] In some embodiments, when a user uses the interactive service of the first digital human through a first access link, if the first configuration information is used to configure storing the default audio set and / or default video set of the first digital human in the user's device, then the cloud management platform transmits the default audio set and / or default video set of the first digital human to the user's device, and the user's device stores the default audio set and / or default video set of the first digital human. The default video in the default audio set and / or default video set is determined by the cloud management platform (e.g., pre-configured) or determined according to the first configuration information. If the first configuration information is used to configure not storing the default audio set or default video set of the first digital human in the user's device, then the cloud management platform does not transmit the default audio set or default video set of the first digital human to the user's device, and the user's device does not need to store the default audio set or default video set of the first digital human. If the first configuration information is used to configure enabling the function of sending first indication information to the first node, then the cloud management platform executes the method provided in the embodiments of this application after receiving the user's first input information. If the first configuration information is used to configure the function of disabling (i.e. not enabling) sending the first instruction information to the first node, then the cloud management platform will not execute the method provided in this embodiment after receiving the user's first input information.

[0162] 403 indicates that the user's first request information has been received.

[0163] The cloud management platform receives the user's first request information, which is used to request interaction with the first digital human.

[0164] In some embodiments, a user sends a first request to the cloud management platform by accessing a first access link. For example, by clicking the first access link, the user requests to open a page for interacting with the first digital human, such as a first graphical interface. The embodiments of this application do not limit the specific form of the first graphical interface.

[0165] For example, the first graphical interface is as follows Figure 5 or Figure 6 As shown. Figure 5 This is a schematic diagram of the first graphical interface 500. (Example) Figure 5 As shown, Figure 5The first graphical interface includes a first area 510 and a second area 520. The first area 510 displays the image of the first digital human, which includes at least one of the following: appearance, expression, posture, and actions. The second area 520 receives user input, such as initial input information. The second area 520 includes an input control 521, a confirmation control, and a cancellation control. The input control 521 is used for the user to input text information about interacting with the first digital human. The confirmation control sends the user-inputted text information to the cloud management platform. The cancellation control cancels the current interaction. Figure 6 This is a schematic diagram of the first graphical interface 600. (For example...) Figure 6 As shown, Figure 6 The first graphical interface 600 includes a first area 610. This first area 610 is used to display the image of the first digital human, and it also includes an input control 611. The input control 611 is used by the user to input voice information for interacting with the first digital human. The input control 611 is also used to send the user-inputted voice information to the cloud management platform.

[0166] 404, in response to the first request information, provides the user with second output information.

[0167] After receiving the first request information, the cloud management platform responds by providing the user with second output information. This second output information includes at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human. In other words, after the user opens a page for interacting with the first digital human through the first access link, the page provides the user with the default audio and / or default video of the first digital human, thereby enhancing the realism of the first digital human and improving the user's interactive experience.

[0168] For example, the default audio included in the second output information belongs to the default audio set of the first digital human, which includes at least one default audio of the first digital human. Each default audio includes audio of the first digital human when it is not interacting with the user, such as the voice of the first digital human in a silent state (i.e., not speaking). When the default audio set includes multiple default audios, the content of at least two default audios in the default audio set is different. Similarly, the default video included in the second output information belongs to the default video set of the first digital human, which includes at least one default video of the first digital human. Each default video includes at least one image frame of the first digital human when it is not interacting with the user, or each default video includes audio and at least one image frame of the first digital human when it is not interacting with the user. The at least one image frame includes, for example, the facial expression and / or posture of the first digital human in a silent state (i.e., not speaking). When the default video set includes multiple default videos, the content of at least two default videos in the default video set is different.

[0169] For example, the default audio set and / or default video set may be configured by the cloud management platform (e.g., pre-configured) or may be configured by the first configuration information.

[0170] For example, the first request information includes the identification information of the first digital human. After receiving the first request information, the cloud management platform determines the default audio set and / or default video set of the first digital human based on the identification information. The cloud management platform determines a default audio from the default audio set and / or a default video from the default video set of the first digital human, thereby determining the second output information and providing it to the user. Alternatively, the cloud management platform transmits the default audio set and / or default video set of the first digital human to the user's device, and the user's device stores the default audio set and / or default video set of the first digital human. The user's device may also determine a default audio from the default audio set and / or a default video from the default video set of the first digital human, thereby determining the second output information and providing it to the user.

[0171] In some embodiments, when the second output information includes a default video of the first digital human, the default video in the second output information is played in a forward-reverse loop when the second output information is provided to the user. That is, the default video in the second output information is played once in forward order, then once in reverse order, then once in forward order again, and then once in reverse order, thereby achieving loop playback. This forward-reverse loop playback order avoids issues such as frame skipping or video discontinuity perceived by the user when providing the second output information, thus improving the user's visual experience.

[0172] Step 405: Receive the user's first input information. The implementation of step 405 is similar to that of step 310, and will not be described again here.

[0173] 406. In response to the first input information, determine the first text information. The implementation of step 406 is similar to that of step 320, and will not be described in detail here.

[0174] 407. Send the first instruction information to the first node. The implementation of step 407 is similar to that of step 330, and will not be described again here.

[0175] 408. Determine the first index based on the first text information.

[0176] Optionally, the first node determines the first index based on the first text information. For example, the first index is the hash value of the first text information, or the first index is the feature vector of the first text information.

[0177] For example, the first node determines the feature vector of the first text information based on the first feature extraction model. Specifically, the first node inputs the first text information into the first feature extraction model and obtains the output of the first feature extraction model, i.e., obtains the feature vector of the first text information. The first feature extraction model may be a model trained by deep learning based on a first training dataset. The first training dataset may include at least one piece of text information, at least one feature vector, and a mapping relationship between at least one piece of text information and at least one feature vector.

[0178] For example, before determining the feature vector of the first text information, the first node can obtain a pre-trained first feature extraction model. Alternatively, before determining the feature vector of the first text information, the first node can obtain a first training dataset and train the model based on the first training dataset to obtain a pre-trained first feature extraction model.

[0179] Optionally, the first node determines a first index based on the first text information and the first identification information. The first identification information includes the identification information of the first digital human and / or the identification information of the second output information. For example, the first index is the hash value of the second character sequence, or the first index is the feature vector of the second character sequence. The second character sequence includes the first text information and the first identification information.

[0180] In some embodiments, the identification information of the second output information includes at least one of the following: identification information of the default audio included in the second output information, and identification information of the default video included in the second output information.

[0181] In some embodiments, when the default audio set of the first digital human includes only one default audio, the first node determines the identifier information of the default audio included in the second output information based on the identifier information of the first digital human, or the first node determines the identifier information of the first digital human based on the identifier information of the default audio included in the second output information. In other words, when the default audio set of the first digital human includes only one default audio, the first identifier information includes the identifier information of the first digital human, or the first identifier information includes the identifier information of the default audio in the second output information, or the first identifier information includes the identifier information of the first digital human and the identifier information of the default audio in the second output information. Similarly, when the default video set of the first digital human includes only one default video, the first node determines the identifier information of the default video included in the second output information based on the identifier information of the first digital human, or the first node determines the identifier information of the first digital human based on the identifier information of the default video included in the second output information. In other words, when the default video set of the first digital human includes only one default video, the first identification information includes the identification information of the first digital human, or the first identification information includes the identification information of the default video in the second output information, or the first identification information includes the identification information of the first digital human and the identification information of the default video in the second output information.

[0182] For example, the first node inputs the second character sequence into the first feature extraction model and obtains the output of the first feature extraction model, that is, obtains the first index.

[0183] 409. Determine whether the first database includes the information corresponding to the first index.

[0184] The first node queries the first database based on the first index to determine whether the first database contains the information corresponding to the first index. If the first database contains the information corresponding to the first index, step 410 is executed, and steps 411-417 are not required. If the first database does not contain the information corresponding to the first index, steps 411-415 are executed.

[0185] In some embodiments, the first database is stored in a first node, which refers to... Figure 2 or Figure 3 The description in the text.

[0186] 410 provides the user with the first output information.

[0187] If the first node determines that the first database contains information corresponding to the first index, the first node provides the first output information to the user.

[0188] In some embodiments, the first node sends first output information to the cloud management platform, causing the second graphical interface to include the first output information. This second graphical interface is a graphical interface provided by the cloud management platform to the user. That is, the user can view the first output information through this second graphical interface. The specific form of this second graphical interface is not limited in the embodiments of this application. The second graphical interface is described in step 415.

[0189] 411, Send the second instruction message to the cloud management platform.

[0190] If the first node determines that the first database does not contain the information corresponding to the first index, the first node sends a second indication message to the cloud management platform. This second indication message indicates that the first database does not contain the first output information. See the description in step 340 for the second indication message.

[0191] 412, send the third instruction information to the second node.

[0192] After receiving the second instruction information, the cloud management platform determines that the first output information is not included in the first database. The cloud management platform then sends a third instruction information to the second node, which instructs the second node to run an interactive service to determine the first output information based on the first text information. This third instruction information is described in step 340.

[0193] 413. Determine the first output information based on the first text information.

[0194] Optionally, the first node inputs the first text information into the first inference model and obtains the output of the first inference model, i.e., obtains the first output information. The first inference model can be a model trained using deep learning based on a second training dataset.

[0195] For example, the second training dataset may include at least one piece of text information, at least one piece of output audio, and a mapping relationship between the at least one piece of text information and the at least one piece of output audio. Each of the at least one piece of output audio corresponds to a piece of text information. In other words, each piece of output audio is an audio corresponding to a piece of text information obtained using the voice of a digital human. That is, the voice in each piece of output audio is the voice of a digital human, and each piece of output audio is audio of the corresponding text information read aloud using the voice of a digital human.

[0196] For example, the second training dataset may include at least one piece of text information, at least one output video, and a mapping relationship between the at least one piece of text information and the at least one output video. Each of the at least one output video corresponds to a piece of text information. In other words, each output video is a video corresponding to a piece of text information obtained using the image of a digital human. For example, the lip movements of the digital human in each output video are matched with the pronunciation lip movements of the corresponding text information.

[0197] For example, the second training dataset may include at least one piece of text information, at least one piece of output audio, at least one piece of output video, and a mapping relationship between the at least one piece of text information, at least one piece of output audio, and at least one piece of output video. The text information corresponding to one piece of output audio and one piece of output video is the same. That is, the lip movements of the first digital human in one piece of output video match the lip movements of the sounds in the output audio corresponding to that piece of output video.

[0198] For example, before determining the first output information, the second node can obtain a pre-trained first inference model. Alternatively, before determining the first output information, the second node can obtain a second training dataset and train the model based on the second training dataset to obtain a pre-trained first inference model.

[0199] Optionally, the second node determines the third output information based on the first identification information. The third output information includes at least one of the following: audio obtained using the voice of the first digital human, and video obtained using the image of the first digital human. This third output information is described in step 350. The second node determines the first output information based on the first text information and the third output information.

[0200] In some embodiments, the second node inputs the first text information and the third output information into the second inference model to obtain the output of the second inference model, i.e., obtains the first output information. The second inference model can be a model trained using deep learning based on a third training dataset.

[0201] For example, the third training dataset may include at least one piece of text information, at least one piece of base audio, at least one piece of output audio, and a mapping relationship between at least one piece of text information, at least one piece of base audio, and at least one piece of output audio. The output audio is similar to the output audio in the second training dataset. Each of the at least one piece of base audio is used to indicate the voice of the digital human; that is, each piece of base audio is audio obtained using the voice of the digital human. The text information corresponding to each piece of base audio may be the same as or different from the text information corresponding to one or more of the at least one piece of output audio.

[0202] For example, the third training dataset may include at least one piece of text information, at least one base video, at least one output video, and a mapping relationship between at least one piece of text information, at least one base video, and at least one output video. The output video is similar to the output video in the second training dataset. Each base video in the at least one base video is used to indicate the image of the digital human; that is, each base video is a video obtained using the image of the digital human. The text information corresponding to each base video may be the same as or different from the text information corresponding to one or more output videos in the at least one output video.

[0203] For example, the third training dataset may include at least one piece of text information, at least one piece of base audio, at least one piece of base video, at least one piece of output audio, at least one piece of output video, and mapping relationships between at least one piece of text information, at least one piece of base audio, at least one piece of base video, at least one piece of output audio, and at least one piece of output video. The text information corresponding to one piece of output audio and one piece of output video is the same. That is, the lip movements of the first digital human in one piece of output video match the lip movements of the sounds in the output audio corresponding to that piece of output video.

[0204] For example, before determining the first output information, the second node can obtain a pre-trained second inference model. Alternatively, before determining the first output information, the second node can obtain a third training dataset and train the model based on the third training dataset to obtain a pre-trained second inference model.

[0205] Optionally, if the first output information includes audio corresponding to the first text information obtained using the voice of the first digital human, the second node converts the first text information into audio corresponding to the first text information obtained using the voice of the first digital human according to TTS technology, thereby determining the first output information.

[0206] 414, Send the first output information to the cloud management platform.

[0207] After determining the first output information, the second node sends the first output information to the cloud management platform.

[0208] 415 provides the first output information to the user.

[0209] After obtaining the first output information, the cloud management platform provides the first output information to the user. Exemplarily, the cloud management platform provides the user with a second graphical interface, allowing the user to view the first output information in the second graphical interface. The specific form of this second graphical interface is not limited in this embodiment.

[0210] In some embodiments, the second graphical interface is the same as the first graphical interface, for example, the second graphical interface and the first graphical interface have the same uniform resource locator (URL). Alternatively, the second graphical interface and the first graphical interface are different graphical interfaces, for example, the second graphical interface and the first graphical interface have different uniform resource locators (URLs).

[0211] For example, when the second graphical interface is the first graphical interface 500, the cloud management platform provides the user with the first output information through the first area 510 in the first graphical interface 500. Alternatively, when the second graphical interface is the first graphical interface 600, the cloud management platform provides the user with the first output information through the first area 610 in the first graphical interface 600.

[0212] Optionally, when the cloud management platform or the first node provides the user with the first output information, the cloud management platform also provides the user with the first text information. For example, the cloud management platform provides the user with the first output information and the first text information through the first area 510 in the first graphical interface 500 or the first area 610 in the first graphical interface 600, and the first text information is provided to the user in the form of subtitles.

[0213] In some embodiments, the last frame of the default video in the second output information is an adjacent frame to the first frame of the video in the first output information. In other words, when the cloud management platform provides the first output information to the user, the first frame of the video in the first output information is played or displayed after the last frame of the default video in the second output information, thereby achieving a smooth transition when the default video and the answer video (i.e., the video in the first output information) are connected, improving the user's visual experience.

[0214] Optionally, after the cloud management platform provides the user with the first output information, it executes step 416 or 417.

[0215] 416 provides the user with a second output message.

[0216] In some embodiments, the first frame of the default video in the second output information and the last frame of the video in the first output information are adjacent frames. In other words, after the cloud management platform provides the first output information to the user, the first frame of the default video in the second output information is played or displayed after the last frame of the video in the first output information, thereby achieving a smooth transition when the default video and the answer video (i.e., the video in the first output information) are connected, improving the user's visual experience.

[0217] In some embodiments, the default video in the second output information in step 416 may be the same as or different from the default video in the second output information in step 404. In other words, after providing the first output information to the user, the cloud management platform may provide the user with a default audio that is different from the default audio provided before providing the first output information, and / or provide the user with a default video that is different from the default video provided before providing the first output information.

[0218] 417, Send the first output information to the first node.

[0219] After determining the first output information, the cloud management platform sends the first output information to the first node. Upon receiving the first output information, the first node stores it in the first database. The index of this first output information in the first database is the first index.

[0220] Figure 4 The method described can provide users with corresponding access links based on tenant configurations. This allows the system to determine whether to provide the digital human's default audio and / or video when the user uses the interactive service through that link, thus avoiding the continuous consumption of computing resources when the digital human doesn't answer questions and reducing resource waste. Furthermore, Figure 4 The method described above can further reduce the waste of computing resources in the second node by instructing the first node to provide the user with the corresponding audio and / or video, without having to use the computing resources in the second node to reason from the answer text to obtain the corresponding audio and / or video each time.

[0221] Figure 7 This is a schematic structural diagram of the cloud management platform provided in the embodiments of this application. Figure 7 The cloud management platform 700 includes a transceiver module 710 and a processing module 720. Figure 7 The cloud management platform 700 can be used to execute Figure 3 or Figure 4 The method in the middle. Figure 7 The cloud management platform 700 can be applied to Figure 1 In the cloud management platform or computing nodes.

[0222] The cloud management platform 700 is used for execution. Figure 3 In the method described above, the transceiver module 710 is used to: receive first input information from the user; send first instruction information to the first node; and, if the first output information is not included in the first database, receive second instruction information from the first node and send third instruction information to the second node to provide the first output information to the user using the interactive service. The transceiver module 710 is used to perform... Figure 3Steps 310, 330, and 340 in the above steps. Processing module 720 is used to: determine the first text information in response to the first input information; processing module 720 is used to execute... Figure 3 Step 320 in the process.

[0223] The cloud management platform 700 is used for execution. Figure 4 When the cloud management platform executes the method, the transceiver module 710 is used to: receive the tenant's first configuration information; receive the user's first request information; receive the user's first input information; and send the first instruction information to the first node. The transceiver module 710 is used to perform... Figure 4 Steps 401, 403, 405, and 407 are described in the text. Processing module 720 is used to: determine a first access link based on first configuration information; provide second output information to the user in response to a first request; and determine first text information in response to the user's first input information. Processing module 720 is used to execute... Figure 4 Steps 402, 404, and 406 in the process.

[0224] In some embodiments, the cloud management platform 700 is used to perform Figure 4 When the cloud management platform executes the method, the transceiver module 710 is also used to: receive second instruction information; send third instruction information to the second node; receive first output information; provide first output information to the user; provide second output information to the user; and send first output information to the first node. The transceiver module 710 is used to execute... Figure 4 Steps 412, 415, 416, and 417 in the text.

[0225] The cloud management platform 700 is used for execution. Figure 4 When the first node in the process executes a method, the transceiver module 710 is used to: receive first instruction information; provide first output information to the user, or send second instruction information to the cloud management platform. The transceiver module 710 is used to execute... Figure 4 Steps 410 or 411 in the first text are described. Processing module 720 is used to: determine a first index based on the first text information; and determine whether the first database includes information corresponding to the first index. Processing module 720 is used to execute... Figure 4 Steps 408 and 409 in the process.

[0226] In some embodiments, the cloud management platform 700 is used to perform Figure 4 When the method executed by the first node in the process is executed, the transceiver module 710 is also used to: receive the first output information.

[0227] The cloud management platform 700 is used for execution. Figure 4 When the second node in the process executes the method, the transceiver module 710 is used to: receive third instruction information; and send first output information to the cloud management platform. The transceiver module 710 is used to execute... Figure 4 Step 414. Processing module 720 is used to: determine first output information based on the first text information. Processing module 720 is used to execute... Figure 4 Step 413 in the process.

[0228] Both the transceiver module 710 and the processing module 720 can be implemented in software or in hardware. For example, the implementation of the processing module 720 will be described below. Similarly, the implementation of the transceiver module 710 can be referenced from the implementation of the processing module 720.

[0229] As an example of a software functional unit, the processing module 720 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the processing module 720 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0230] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0231] As an example of a hardware functional unit, the processing module 720 may include at least one computing device, such as a server. Alternatively, the processing module 720 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0232] The processing module 720 includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the processing module 720 can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the processing module 720 can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0233] Therefore, the modules of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0234] It should be noted that the above embodiments of the device, when executing the above methods, are only illustrative examples of the division of functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. For example, the transceiver module 710 can be used to execute any step in the above methods, and the processing module 720 can be used to execute any step in the above methods. The steps implemented by the transceiver module 710 and the processing module 720 can be specified as needed, and the device can achieve all its functions by implementing different steps in the above methods through the transceiver module 710 and the processing module 720 respectively.

[0235] Furthermore, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments above, which will not be repeated here.

[0236] The method provided in this application can be executed by a computing device, which can also be referred to as a computer system. It includes a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as processing units, memory, and memory control units; the functions and structure of this hardware will be described in detail later. The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software. Optionally, the computer system can be a handheld device such as a smartphone, or a terminal device such as a personal computer; this application does not particularly limit this, as long as the method provided in this application can be used. The executing entity of the method provided in this application can be a computing device, or a functional module within the computing device capable of calling and executing programs.

[0237] Figure 8 This is a schematic structural block diagram of a computing device 800 provided in an embodiment of this application. The computing device 800 may be a server, a computer, or other device with computing capabilities. Figure 8 The computing device 800 shown includes at least one processor 810 and a memory 820.

[0238] It should be understood that this application does not limit the number of processors and memories in the computing device 800.

[0239] The processor 810 executes instructions in the memory 820, causing the computing device 800 to implement the method provided in this application. Alternatively, the processor 810 executes instructions in the memory 820, causing the computing device 800 to implement the various functional modules provided in this application, thereby implementing the method provided in this application.

[0240] Optionally, the computing device 800 also includes a communication interface 830. The communication interface 830 uses a transceiver module, such as, but not limited to, a network interface card or a transceiver, to enable communication between the computing device 800 and other devices or communication networks.

[0241] Optionally, the computing device 800 also includes a system bus 840, wherein the processor 810, memory 820, and communication interface 830 are respectively connected to the system bus 840. The processor 810 can access the memory 820 through the system bus 840; for example, the processor 810 can perform data read / write or code execution in the memory 820 through the system bus 840. The system bus 840 is a peripheral component interconnect express (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus 840 is divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0242] In one possible implementation, the processor 810 primarily functions to interpret the instructions (or code) of a computer program and process data within the computer software. The instructions of the computer program and the data within the computer software can be stored in the memory 820 or the cache of the processor 810.

[0243] Optionally, the processor 810 may be an integrated circuit chip with signal processing capabilities. By way of example and not limitation, the processor 810 may be a general-purpose processor, a digital signal processor (DSP), an ASIC, an FPGA, or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Among these, a general-purpose processor is a microprocessor, etc. For example, the processor 810 may be a central processing unit (CPU).

[0244] The memory 820 provides runtime space for processes in the computing device 800. For example, the memory 820 stores the computer program (specifically, the program code) used to generate the process. After the computer program is run by the processor to generate a process, the processor allocates corresponding storage space for the process in the memory 820. Furthermore, the aforementioned storage space further includes text segments, initialized data segments, bit initialized data segments, stack segments, heap segments, etc. The memory 820 stores data generated during the process's execution, such as intermediate data or process data, in the aforementioned process-specific storage space.

[0245] Optionally, the memory, also known as RAM, is used to temporarily store the data processed by the processor 810, as well as data exchanged with external storage devices such as hard disks. As long as the computer is running, the processor 810 will load the data required for processing into RAM for computation, and then transfer the result back out after the computation is complete.

[0246] By way of example and not limitation, memory 820 may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile storage medium may be, for example, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory is random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus DRAM (DRDRAM). It should be noted that the memory 820 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0247] The structures of the computing device 800 listed above are merely illustrative and are not limited thereto. The computing device 800 in this application includes various hardware components in existing computer systems. For example, the computing device 800 also includes other memories besides the memory 820, such as disk storage. Those skilled in the art should understand that the computing device 800 may also include other devices necessary for normal operation. Furthermore, depending on specific needs, those skilled in the art should understand that the computing device 800 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the computing device 800 may only include the devices necessary for implementing the embodiments of this application, and may not necessarily include... Figure 8All the devices shown.

[0248] This application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server. In some embodiments, the computing device may also be a desktop computer, a laptop computer, or a smartphone, or other terminal device.

[0249] like Figure 9 As shown, the computing device cluster includes at least one computing device 800. The memory 820 of one or more computing devices 800 in the computing device cluster may store the same instructions for performing the methods described above.

[0250] In some possible implementations, the memory 820 of one or more computing devices 800 in the computing device cluster may also each store a portion of the instructions for executing the above-described method. In other words, a combination of one or more computing devices 800 can jointly execute the instructions of the above-described method.

[0251] It should be noted that the memories 820 in different computing devices 800 within the computing device cluster can store different instructions, each used to execute a portion of the functions of the aforementioned device. That is, the instructions stored in the memories 820 of different computing devices 800 can implement the functions of one or more modules within the aforementioned device.

[0252] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 10 One possible implementation is shown. For example... Figure 10 As shown, the two computing devices 800A and 800B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device.

[0253] It should be understood that Figure 10 The functions of the computing device 800A shown can also be performed by multiple computing devices 800. Similarly, the functions of the computing device 800B can also be performed by multiple computing devices 800.

[0254] In this embodiment of the application, a computer program product containing instructions is also provided. The computer program product may be software or a program product containing instructions that can run on a computing device cluster or be stored on any available medium. When run by the computing device cluster, it causes the computing device cluster to perform the methods provided above, or causes the computing device cluster to perform the functions of the apparatus provided above.

[0255] This application embodiment also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., high-density digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that, when executed by a cluster of computing devices, cause the cluster of computing devices to perform the method provided above.

[0256] In this application embodiment, a chip system is also provided. The chip system includes logic circuitry for coupling with an input / output interface to transmit data via the input / output interface, thereby executing the methods provided above.

[0257] This application embodiment also provides an interactive system. The interactive system includes... Figure 3 or Figure 4 The cloud management platform, the first node, and the second node described herein.

[0258] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0259] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0260] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0261] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0262] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0263] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0264] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An interaction method, characterized in that, The method is applied to a cloud management platform, which manages infrastructure providing cloud services, the infrastructure including at least one computing node, and the method includes: Receive first input information from the user, the first input information being used to request the answer corresponding to the user's input; In response to the first input information, first text information is determined, which is used to indicate the answer text corresponding to the user's input; A first instruction message is sent to a first node, the first instruction message including the first text information. The first instruction message is used to instruct the first node to provide first output information to the user. The first node stores a first database, the first database including at least one audio segment and / or at least one video segment of the first digital human. The first output information includes at least one of the following: audio corresponding to the first text information obtained using the voice of the first digital human; or video corresponding to the first text information obtained using the image of the first digital human; or... If the first output information is not included in the first database, a second indication information is received from the first node, and a third indication information is sent to the second node to provide the first output information to the user using the interactive service. The second node runs the interactive service. The second indication information is used to indicate that the first output information is not included in the first database, and the third indication information is used to instruct the second node to determine the first output information based on the first text information.

2. The method according to claim 1, characterized in that, The first node belongs to a Content Delivery Network (CDN).

3. The method according to claim 1 or 2, characterized in that, If the first output information is not included in the first database, the step of sending third indication information to the second node to provide the first output information to the user includes: Send a third indication message to the second node, wherein the second node belongs to the at least one computing node; Receive the first output information from the second node; Provide the first output information to the user.

4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Before and / or after providing the first output information to the user, second output information is provided to the user, the second output information including at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Receive the first configuration information of the tenant, the first configuration information is used to configure at least one of the following: whether to store the default audio set and / or default video set of the first digital human in the user's device, whether to enable the function of sending the first indication information to the first node; Based on the first configuration information, a first access link is provided to the user, and the first access link is used to provide the interactive service to the user.

6. The method according to any one of claims 1 to 5, characterized in that, In the case where the first output information includes a video corresponding to the first text information obtained using the image of the first digital human, the first output information is determined based on the third output information and the first text information. The third output information includes a base video obtained using the image of the first digital human, or the third output information includes the last frame of a default video obtained using the image of the first digital human and the base video obtained using the image of the first digital human.

7. The method according to claim 6, characterized in that, The third output information includes: the last frame of the default video obtained using the image of the first digital human, the first frame of the base video obtained using the image of the first digital human, a reference frame of the base video obtained using the image of the first digital human, and the base video obtained using the image of the first digital human, wherein the reference frame of the base video belongs to the base video.

8. The method according to claim 6 or 7, characterized in that, The third output information is determined based on the first identification information, which includes the identification information of the first digital human and / or the identification information of the second output information. The second output information includes at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human.

9. The method according to any one of claims 1 to 8, characterized in that, When the first output information is included in the first database, the index of the first output information in the first database is a first index. The first index is determined based on the first identification information and the first text information. The first identification information includes the identification information of the first digital human and / or the identification information of the second output information. The second output information includes at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human.

10. The method according to any one of claims 1 to 9, characterized in that, If the first output information is not included in the first database, the method further includes: Send the first output information to the first node.

11. An interaction method, characterized in that, The method includes: The cloud management platform receives the user's first input information, which is used to request the answer corresponding to the user's input. In response to the first input information, the cloud management platform determines the first text information, which is used to indicate the answer text corresponding to the user's input. The cloud management platform sends a first instruction information to the first node. The first instruction information includes the first text information. The first instruction information is used to instruct the first node to provide the user with first output information. The first node stores a first database. The first database includes at least one audio segment and / or at least one video segment of the first digital human. The first output information includes at least one of the following: audio corresponding to the first text information obtained using the voice of the first digital human, and video corresponding to the first text information obtained using the image of the first digital human. The first node determines whether the first database includes the first output information based on the first indication information; If the first database includes the first output information, the first node provides the first output information to the user; or... If the first output information is not included in the first database: The first node sends a second instruction to the cloud management platform, the second instruction indicating that the first database does not include the first output information; The cloud management platform sends a third instruction message to the second node, the third instruction message being used to instruct the second node to determine the first output message based on the first text message; The cloud management platform provides the user with the first output information.

12. A cloud management platform, characterized in that, The cloud management platform is used to manage the infrastructure that provides cloud services, the infrastructure including at least one computing node, and the cloud management platform includes: The transceiver module is used to receive the user's first input information, which is used to request the answer corresponding to the user's input. The processing module is configured to, in response to the first input information, determine first text information, wherein the first text information is used to indicate the answer text corresponding to the user's input; The transceiver module is configured to send first instruction information to a first node. The first instruction information includes the first text information and is used to instruct the first node to provide first output information to the user. The first node stores a first database, which includes at least one audio segment and / or at least one video segment of a first digital human. The first output information includes at least one of the following: audio corresponding to the first text information obtained using the voice of the first digital human; or video corresponding to the first text information obtained using the image of the first digital human. In the case that the first output information is not included in the first database, the transceiver module is configured to receive second indication information from the first node and send third indication information to the second node to provide the first output information to the user using the interactive service. The second node runs the interactive service. The second indication information is used to indicate that the first output information is not included in the first database, and the third indication information is used to instruct the second node to determine the first output information based on the first text information.

13. The cloud management platform according to claim 12, characterized in that, The first node belongs to a Content Delivery Network (CDN).

14. The cloud management platform according to claim 12 or 13, characterized in that, In the case that the first output information is not included in the first database, the transceiver module is specifically used for: Send a third indication message to the second node, wherein the second node belongs to the at least one computing node; Receive the first output information from the second node; Provide the first output information to the user.

15. The cloud management platform according to any one of claims 12 to 14, characterized in that, The transceiver module is further configured to provide the user with second output information before and / or after providing the first output information, the second output information including at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human.

16. The cloud management platform according to any one of claims 12 to 15, characterized in that, The transceiver module is further configured to receive first configuration information from the tenant, wherein the first configuration information is configured to configure at least one of the following: whether to store the default audio set and / or default video set of the first digital human in the user's device, whether to enable the function of sending the first indication information to the first node; The processing module is further configured to provide the user with a first access link based on the first configuration information, wherein the first access link is used to provide the user with the interactive service.

17. The cloud management platform according to any one of claims 12 to 16, characterized in that, In the case where the first output information includes a video corresponding to the first text information obtained using the image of the first digital human, the first output information is determined based on the third output information and the first text information. The third output information includes a base video obtained using the image of the first digital human, or the third output information includes the last frame of a default video obtained using the image of the first digital human and the base video obtained using the image of the first digital human.

18. The cloud management platform according to claim 17, characterized in that, The third output information includes: the last frame of the default video obtained using the image of the first digital human, the first frame of the base video obtained using the image of the first digital human, a reference frame of the base video obtained using the image of the first digital human, and the base video obtained using the image of the first digital human, wherein the reference frame of the base video belongs to the base video.

19. The cloud management platform according to claim 17 or 18, characterized in that, The third output information is determined based on the first identification information, which includes the identification information of the first digital human and / or the identification information of the second output information. The second output information includes at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human.

20. The cloud management platform according to any one of claims 12 to 19, characterized in that, When the first output information is included in the first database, the index of the first output information in the first database is a first index. The first index is determined based on the first identification information and the first text information. The first identification information includes the identification information of the first digital human and / or the identification information of the second output information. The second output information includes at least one of the following: default audio obtained using the voice of the first digital human, and default video obtained using the image of the first digital human.

21. The cloud management platform according to any one of claims 12 to 20, characterized in that, If the first output information is not included in the first database, the transceiver module is further configured to send the first output information to the first node.

22. A computing device cluster, characterized in that, It includes at least one computing device, each computing device including a processor and memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1 to 10.

23. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device cluster, the computing device cluster causes the computing device cluster to perform the method as described in any one of claims 1 to 10.

24. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a cluster of computing devices, perform the method as described in any one of claims 1 to 10.