Communication method and apparatus

By generating digital human videos with natural expressions through the interaction between the first and third network elements, the problem of high resource consumption for expression processing in call assistant services is solved, thus improving the user experience.

WO2026153011A1PCT designated stage Publication Date: 2026-07-23HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-12-17
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

In current call assistant services, the processing of facial expressions for digital humans consumes a lot of resources, and the expressions are unrelated to the output content of the call assistant, resulting in low fun and playability and an inadequate user experience.

Method used

By interacting with the first network element and the third network element, digital human videos based on facial expression description information and emotional information are generated, achieving natural facial expressions with minimal computational cost and avoiding increasing model complexity.

Benefits of technology

It improves the user's interactive experience, making the digital human's expressions more natural and reducing processing complexity and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025143306_23072026_PF_FP_ABST
    Figure CN2025143306_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of communications, and provides a communication method and apparatus. The method comprises: receiving a first notification, the first notification being used for indicating that a first call assistant service has been started, and the first notification comprising an identifier of a digital human corresponding to the first call assistant service; on the basis of the first notification, requesting first information, the first information being used for indicating first expression description information of the digital human corresponding to the first call assistant service; acquiring second information, the second information comprising prompt information replied by a call assistant and emotion information matching the prompt information; on the basis of the first information and the second information, generating first video information, the first video information being used for indicating a video of the digital human corresponding to the first call assistant service; and sending the first video information. Since the video of the digital human is generated on the basis of the first expression description information and the emotion information matching the prompt information, expressions of the digital human in the video of the digital human are more natural. On this basis, user service experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

A communication method and apparatus

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202510082066.8, filed on January 17, 2025, entitled "A Communication Method and Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of communication technology, and in particular to a communication method and apparatus. Background Technology

[0004] Digital humans (or meta-humans) are digitized human figures created using digital technology, closely resembling human forms. Breaking physical boundaries, digital humans offer anthropomorphic services, characterized by hyper-realism and strong interactivity, and represent a future development trend.

[0005] When users engage with the call assistant service, the voice-driven digital avatar of the call assistant can enhance the user's interactive experience. However, in the current call assistant service, driving the digital avatar's facial expressions requires a complex processing model, resulting in high processing resource consumption. Summary of the Invention

[0006] This application provides a communication method and apparatus to make the expressions of the digital human more natural when using a call assistant service, thereby improving the user experience without increasing the complexity of model processing.

[0007] Firstly, this application provides a communication method applicable to a first network element, which provides media resource management functions. The first network element can be the network element itself, a component within the network element (e.g., a processor, chip, or chip system), or a logical module or software implementing all or part of the functions of the first network element. This application does not specifically limit the scope here. The execution is as follows:

[0008] Receive a first notification, which indicates that a first call assistant service has been activated. The first notification includes: the identifier of the digital human corresponding to the first call assistant service; request first information according to the first notification, which indicates the first expression description information of the digital human corresponding to the first call assistant service; obtain second information, which includes: prompt information from the call assistant and emotion information matching the prompt information; generate first video information according to the first and second information, which indicates the video of the digital human corresponding to the first call assistant service; and send the first video information.

[0009] The prompt message can be audio or voice text. This is merely an example and not a specific limitation.

[0010] In this application, after confirming that the first call assistant service has been activated, the system requests the first facial expression information of the digital human corresponding to the first call assistant service. Then, based on the prompt information from the call assistant's response, the emotional information matching the prompt information, and the first facial expression description information, a digital human video is generated. Because the digital human video is generated based on the first facial expression information and the emotional information matching the prompt information, the digital human's expressions in the video are more natural. Furthermore, the digital human's expression generation does not involve complex processing models, requiring only minimal computational cost and enabling high-concurrency operation. Therefore, when using the call assistant service, the voice-driven digital human's expressions are more natural, which can improve the user's service experience.

[0011] In one alternative approach, the first network element determines second expression description information based on the second information and the first information. The second expression description information is associated with the emotion information matched by the prompt information, and the second expression description information belongs to the first expression description information. The third information is obtained based on the second expression description information. The third information is used to indicate the video frame or image frame of the digital human corresponding to the first call assistant service that matches the second expression description information. The first video information is generated based on the second information and the third information.

[0012] After determining the facial expression description information that matches the second information, the corresponding digital human video frame or image frame is requested in order to generate a video of the digital human. Based on this, the digital human's facial expressions in the video are more natural.

[0013] In one alternative approach, the first network element determines the second expression description information based on the second information and the first information. The second expression description information is associated with the emotion information matched by the prompt information, and the second expression description information belongs to the first expression description information. The first video information is generated based on the second expression description information and the third information.

[0014] After determining the second expression description information that matches the second information, when the second expression description information is a specific emotion file, the first network element can generate the first video information based on the second expression description information and the second information.

[0015] In one alternative approach, the first network element generates fourth information based on the second information, the fourth information being used to indicate a label for emotional information that matches the prompt information; and determines the second expression description information based on the fourth information and the first information.

[0016] In this application, after the first network element generates fourth information based on the second information, it searches for the first information based on the fourth information to determine the matching second facial expression description information, thereby generating a video of the digital human. The facial expressions of the digital human in the video generated in this way are more natural.

[0017] In one alternative approach, the first network element also receives fourth information from the second network element, the fourth information being used to indicate a label for emotion information matching the prompt information; and determines second expression description information based on the fourth information and the first information.

[0018] In this application, after the first network element obtains the fourth information from the second network element, it searches for the first information based on the fourth information to determine the matching second facial expression description information, and then generates a video of the digital human. The facial expressions of the digital human in the video generated in this way are more natural.

[0019] In one optional manner, the first network element sends a first request message to the third network element based on the second expression description information. The first request message includes: the second expression description information. The first request message is used to request video frames or image frames of the digital human corresponding to the first call assistant service that matches the second expression description information; and receives third information from the third network element.

[0020] The first network element obtains third information through signaling interaction with the third network element, without increasing the processing complexity of the first network element, and can improve data processing efficiency.

[0021] In one alternative approach, the fourth information is associated with the following: the duration of the prompt message and the frame rate of the video of the digital human corresponding to the first call assistant service.

[0022] When the duration of the fourth information and the prompt information are related to the frame rate of the digital human's video, it can be ensured that the digital video generated by the first network element is compatible with the prompt information and meets the requirements of the video configuration.

[0023] In one alternative approach, the fourth piece of information is a sequence used to indicate emotional information.

[0024] When the fourth piece of information is a sequence, it facilitates the identification and processing by the first network element, thereby improving data processing efficiency.

[0025] In one alternative approach, the first network element sends a second request message to the third network element based on the first notification. The second request message includes: the identifier of the digital human corresponding to the first call assistant service, and the second request message is used to request the first expression description information; and receives the first information from the third network element.

[0026] The first network element obtains the first information through signaling interaction with the third network element, without increasing the processing complexity of the first network element, and can improve data processing efficiency.

[0027] In one alternative approach, the first network element receives second information from the second network element.

[0028] Secondly, this application provides a communication method applicable to a second network element for providing call assistant services. The second network element can be the second network element itself, a component within the second network element (e.g., a processor, chip, or chip system), or a logic module or software implementing all or part of the functions of the second network element. This application does not specifically limit the scope here. The execution is as follows:

[0029] Receive multimedia information; obtain second information based on the multimedia information, the second information including: prompt information from the call assistant and emotional information matching the prompt information; send the second information to the first network element.

[0030] In one alternative approach, the second network element further generates fourth information based on the second information, the fourth information being used to indicate a label for the emotional information that matches the prompt information; and sends the fourth information to the first network element.

[0031] In one alternative approach, the fourth information is associated with the following: the duration of the prompt message and the frame rate of the video of the digital human corresponding to the first call assistant service.

[0032] In one alternative approach, the fourth piece of information is a sequence used to indicate emotional information.

[0033] Thirdly, this application provides a communication method applicable to a third network element used to provide digital human facial expression description files. This third network element can be the third network element itself, a component within the third network element (e.g., a processor, chip, or chip system), or a logic module or software implementing all or part of the functions of the third network element. This application does not specifically limit the scope of this application.

[0034] Execute as follows:

[0035] Receive a second request message from a first network element, the second request message including: the identifier of the digital human corresponding to the first call assistant service, the second request message being used to request first expression description information; send first information to the first network element, the first information being used to indicate the first expression description information of the digital human corresponding to the first call assistant service.

[0036] In one alternative approach, the third network element also receives a first request message from the first network element, the first request message including: second expression description information, the first request message being used to request video frames or image frames of the digital human corresponding to the first call assistant service that match the second expression description information; and sending third information to the first network element, the third information being used to indicate video frames or image frames of the digital human corresponding to the first call assistant service that match the second expression description information.

[0037] Fourthly, embodiments of this application provide a communication device, which can be a first network element, a second network element, or a third network element. The communication device has the functions to implement the first to third aspects described above. For example, the communication device includes modules, units, or means corresponding to the steps involved in the first to third aspects. These functions, units, or means can be implemented by software, hardware, or hardware executing corresponding software.

[0038] In one possible design, the communication device includes a processing unit and a transceiver unit. The transceiver unit can be used to send and receive signals to enable communication between the communication device and other devices. The processing unit can be used to perform some internal operations of the communication device. The transceiver unit can be called an input / output unit, a communication unit, etc., and can be a transceiver; the processing unit can be a processor. When the communication device is a module (e.g., a chip) in a communication device, the transceiver unit can be an input / output interface, input / output circuit, or input / output pins, etc., and can also be called an interface, communication interface, or interface circuit, etc.; the processing unit can be a processor, processing circuit, or logic circuit, etc.

[0039] In another possible design, the communication device includes a processor and may further include a transceiver for transmitting and receiving signals. The processor executes program instructions to perform the methods in any of the possible designs or implementations of the first to third aspects described above. The communication device may also include one or more memories coupled to the processor, which may store necessary computer programs or instructions for implementing the functions involved in the first to third aspects described above. The processor can execute the computer programs or instructions stored in the memory, and when the computer programs or instructions are executed, the communication device implements the methods in any of the possible designs or implementations of the first to third aspects described above.

[0040] In another possible design, the communication device includes a processor that can be coupled to a memory. The memory can store necessary computer programs or instructions for implementing the functions described in the first to third aspects above. The processor can execute the computer programs or instructions stored in the memory, causing the communication device to implement the methods in any possible design or implementation of the first to third aspects above when the computer programs or instructions are executed.

[0041] In another possible design, the communication device includes a processor and an interface circuit, wherein the processor is used to communicate with other devices through the interface circuit and to perform the methods in any possible design or implementation of the first to third aspects described above.

[0042] Understandably, in the fourth aspect mentioned above, the processor can be implemented in hardware or software. When implemented in hardware, the processor can be a logic circuit, integrated circuit, etc.; when implemented in software, the processor can be a general-purpose processor that reads software code stored in memory. Furthermore, there can be one or more processors, and one or more memories. The memory can be integrated with the processor, or the memory and processor can be separate. In specific implementations, the memory can be integrated with the processor on the same chip, or it can be set on different chips. This application does not limit the type of memory or the arrangement of the memory and processor.

[0043] Fifthly, embodiments of this application provide a communication system including the aforementioned first network element, second network element, and third network element. The first network element can be used to execute the method in the first aspect, the second network element can be used to execute the method in the second aspect, and the third network element can be used to execute the method in the third aspect. Furthermore, it should be noted that in each aspect, there may be processes executed interactively by multiple devices or network elements; the corresponding processes cannot be executed by a single device or network element. Instead, they are mainly executed through the interaction of corresponding devices or network elements, which will not be elaborated upon here.

[0044] Sixthly, this application provides a chip system including a processor and potentially a memory for implementing the methods described in the first to third aspects above. The chip system may be composed of chips or may include chips and other discrete devices.

[0045] In a seventh aspect, this application also provides a computer-readable storage medium storing computer-readable instructions that, when executed on a computer, cause the computer to perform the methods described in the first to third aspects.

[0046] Eighthly, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the methods of the embodiments of the first to third aspects described above.

[0047] The technical effects that can be achieved by the second to eighth aspects mentioned above can be referred to the description of the technical effects that can be achieved by the corresponding possible design schemes in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0048] Figure 1A shows a schematic diagram of a communication scenario;

[0049] Figure 1B shows a schematic diagram of another communication scenario;

[0050] Figure 2A shows a schematic diagram of a communication architecture;

[0051] Figure 2B shows a schematic diagram of another communication architecture;

[0052] Figure 3 shows a flowchart of a communication method provided in an embodiment of this application;

[0053] Figure 4 shows a schematic diagram of determining a second facial expression description file according to an embodiment of this application;

[0054] Figure 5 shows a schematic diagram of a data processing flow provided in an embodiment of this application;

[0055] Figures 6 to 10 show schematic flowcharts of a communication method provided in an embodiment of this application;

[0056] Figure 11 shows a schematic diagram of the structure of a communication device provided in an embodiment of this application;

[0057] Figure 12 shows a schematic diagram of the structure of a communication device provided in an embodiment of this application;

[0058] Figure 13 shows a schematic diagram of the structure of a communication device provided in an embodiment of this application. Detailed Implementation

[0059] In the embodiments of this application, the solutions in each embodiment can be used in a reasonable combination, and the explanations or descriptions of various terms, similar operations, or steps appearing in the embodiments can be referenced or explained to each other in the embodiments, without limitation.

[0060] In the embodiments of this application, "transmission" includes "sending" and / or "receiving." "Sending" and "receiving" indicate the direction of signal transmission. For example, "sending information to XX" can be understood as the destination of the information being XX, which can include direct transmission via the air interface or indirect transmission by other units or modules via the air interface. "Receiving information from YY" can be understood as the source of the information being YY, which can include direct reception from YY via the air interface or indirect reception from YY by other units or modules via the air interface. "Sending" can also be understood as the "output" of a chip interface, and "receiving" can also be understood as the "input" of a chip interface. In other words, sending and receiving can occur between devices, such as between access network devices and terminal devices, or within a device, such as between components, modules, chips, software modules, or hardware modules within the device via a bus, wiring, or interface.

[0061] In this application embodiment, the number of nouns, unless otherwise specified, refers to "singular nouns or plural nouns," that is, "one or more." "At least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone, where A / B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. For example, A / B means: A or B. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and / or c means the following combinations: a exists alone, b exists alone, c exists alone, a and b exist simultaneously, a and c exist simultaneously, b and c exist simultaneously, or a, b, and c exist simultaneously, where a, b, and c can be single or multiple.

[0062] In the embodiments of this application, "when," "if," and "if" all refer to the device taking corresponding actions under certain objective circumstances, and are not time-limited, nor do they require the device to perform a judgment action, nor do they imply any other limitations. Unless otherwise specified, "if" and "if" are interchangeable, and "when" and "in the case of" are interchangeable. "When" and "if" / "if" are interchangeable. "Associated" and "corresponding" are interchangeable.

[0063] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0064] In this application, the ordinal numbers such as "first" and "second" are used to distinguish multiple objects, and are not used to limit the size, content, order, timing, priority, or importance of the multiple objects. For example, the first parameter and the second parameter refer to two different parameters, and do not indicate a difference in the priority or importance of these two parameters.

[0065] The technical solutions provided in the embodiments of this application can be applied to various communication systems, such as cellular systems related to the 3rd Generation Partnership Project (3GPP), such as Long Term Evolution (LTE) communication systems, 5th Generation (5G) mobile communication systems / New Radio (NR) communication systems, or future-oriented evolution systems, or other similar communication systems. Other similar communication systems include Wireless Fidelity (WiFi), Vehicle-to-Everything (V2X), Spark Link systems, Bluetooth systems, Near Field Communication (NFC) systems, and Internet of Things (IoT) systems, such as Ambient IoT (A-IoT / A-IoT) and Narrow Band Internet of Things (NB-IoT). Alternatively, the solutions provided in the embodiments of this application can also be applied to communication systems that integrate two or more of the above systems. It should be understood that IoT technology is widely used in various industries; for example, IoT technology can be applied to scenarios such as logistics, warehousing, industrial manufacturing, identity recognition, or environmental monitoring.

[0066] Figure 1A illustrates an application scenario applicable to this application. A user calls a call assistant via an IP multimedia subsystem (IMS) network, and the call assistant provides the user with the required services. For example, as shown in Figure 1A, a user requests the call assistant to "check the weather," "check tourist attractions," "order food," or "remind them of the basketball game on Saturday morning," etc. The call assistant can provide the user with feedback on "weather conditions," "tourist attraction information," "food order status," and "basketball game reminder," etc. Here, the call assistant is a service subscribed to by the user, and the call assistant is implemented based on artificial intelligence (AI). Optionally, the call assistant is implemented based on a large language model (LLM). Figure 1A illustrates a customer-to-manufacturer (C2M) scenario, that is, a scenario where the user directly calls the call assistant.

[0067] Figure 1B illustrates another application scenario applicable to this application. User A calls User B via the IMS network, and the service is provided by a call assistant subscribed to by User A. For example, as shown in Figure 1B, User A calls User B via the IMS network, asking B, "Are you free tonight? Let's have dinner together." After User B agrees to the dinner request, User A requests the call assistant to "help order food," and the call assistant then orders food for both User A and User B. Figure 1A illustrates a customer-to-consumer (C2C) scenario, where users call each other, and during the call, the user subscribed to the call assistant wakes up the call assistant, which then helps the user complete the task.

[0068] The following describes the network architecture applicable to the communication method proposed in this application. Referring to Figure 2A, a schematic diagram of a network architecture provided in an embodiment of this application is shown. The network architecture shown in Figure 2A includes a proxy call session control function (P-CSCF) network element, a serving call session control function (S-CSCF) network element, an IMS access media gateway (AGW), a home subscriber server (HSS), an IMS application server (AS), a media function (MF) network element, a data channel signaling function (DCSF) network element, a network exposure function (NEF) network element, a terminal device, a data channel application server (DCAS), and a mobile IMS. The MF also includes an AI assistant logic unit (subscriber AI agent function, SAAF). Optionally, this network architecture also includes a data channel application repository (DCAR). In the architecture shown in Figure 2B, the SAAF and MF are set up separately.

[0069] This section provides a brief introduction to the network elements (or functional network elements, functional entities, nodes, devices, etc.) shown in Figures 2A and 2B:

[0070] 1. Terminal Equipment: Terminal equipment, also known as user equipment (UE), terminal, terminal device, access terminal, user unit, user station, mobile station, mobile station (MS), mobile terminal (MT), remote station, remote terminal, mobile device, user terminal, wireless communication equipment, user agent, or user device, is a device with wireless or wired communication capabilities. For example, terminal equipment can connect to wireless access equipment via an air interface, or it can connect to wired access equipment via a wired interface. In terms of product form, terminal equipment can include handheld devices, vehicle-mounted devices, wearable devices, or computing devices with communication functions. For example, terminal devices can be mobile phones, tablets, computers with wireless communication capabilities (such as laptops, PDAs, etc.), mobile internet devices (MIDs), virtual reality (VR) terminals, augmented reality (AR) terminals, smartwatches, smart bracelets, smart glasses, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical care, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, SIP phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, other processing devices connected to wireless modems, terminals in Internet of Things (IoT) systems, desktop computers, or third-generation partnerships. The project, UE, etc., as defined by the 3GPP standard specifications are not restricted.

[0071] 2. P-CSCF network element: The entry node for users to access the IMS network, mainly responsible for forwarding SIP signaling between IMS users and the home network.

[0072] 3. S-CSCF Network Element: The S-CSCF is the unified entry point of the IMS user's home network, responsible for allocating or querying the S-CSCF that provides services to the user.

[0073] 4. IMS AGW: IMS AGW can provide IMS network access gateway and media gateway functions.

[0074] 5. DCSF: Used to provide signaling control functions for data channels.

[0075] 6. DCAR: DCAR is a repository used to store data channel applications (DC Apps). After a developer completes a DC App, it is uploaded to the operator's DCS, which then saves it to the DCAR. When needed, the DCS downloads the DC App from the DCAR to its local machine for further processing.

[0076] The DC control network element in this application embodiment includes DCSF, and optionally also includes DCAR.

[0077] 7. HSS: In IMS, HSS serves as a database for storing user information, where users store their data. As an example and not a limitation, user data may include information about the data channel services that the user has contracted with the operator.

[0078] 8. IMS AS: The IMS AS is the top-level application layer device in the IMS system, providing basic and supplementary services such as multimedia conferencing, unified communications, SMS gateways, and standard operator desks. The IMS network is an open system based on IP that provides various multimedia services to users. The IMS AS interacts with the CSCF (CSCF) to trigger and execute various network services. Furthermore, the terminal device communicates with the IMS AS, and the IMS AS can establish data channels for the terminal device. For example, the IMS AS can be a multimedia telephony application server (MMTEL AS) or a telephone application server (TAS).

[0079] 9. DC AS: DC AS is used to provide data channel service logic. It can be understood that DC AS can be deployed within the IMS network. In this case, DC AS can be directly connected to other network elements in the IMS network through interfaces (such as service interfaces); or, DC AS can be deployed outside the IMS network. In this case, DC AS can be connected to other network elements in the IMS network through network exposure function (NEF) network elements (or IMS border network management).

[0080] 10. NEF Network Element: The NEF network element is used to securely expose various services of the 5GC network to third parties.

[0081] 11. MF Network Element: Also known as a media server, it is used to provide media resource management functions, including media resource management of media channels and media resource management of data channels.

[0082] 12. SAAF Network Element: Responsible for the logical functions of the user AI Agent, used to implement call assistant functions. This SAAF network element can be set in the MF or deployed independently; this application does not specifically limit its deployment.

[0083] Mb, Iq, Gm, Mw, N70, N71 / Sh, N72 / Sc, DC1, DC2, DC3, DC4, DC5, ISC, MDC1, MDC2, MDC3, and N33 provide service interfaces for the network elements in Figure 2A or Figure 2B, which are used to call the corresponding service operations.

[0084] To enhance user interaction, the call assistant service has introduced a voice-driven digital human feature. In call assistant scenarios, voice commands control the digital human avatar, increasing the interactive experience between the user and the call assistant. A digital human is a virtual character with a digital appearance, displayed through devices such as mobile phones, televisions, augmented reality (AR) glasses, or virtual reality (VR) glasses. Digital humans possess a human-like or near-realistic appearance, with intuitive anthropomorphic features such as facial features, gender, and personality. Driven by digital technology, digital humans can exhibit human-like behaviors, including language, facial expressions, and gestures. Furthermore, artificial intelligence can enable digital humans to possess simple thoughts, recognize their environment, and interact with people. Digital humans are widely used in film and television production, virtual broadcasters, virtual education, gaming, virtual customer service, virtual tour guides, and real-time communication.

[0085] Currently, the voice assistant only drives the digital avatar's mouth movements, not its facial expressions (the digital avatar's expressions are unrelated to the call assistant's output). The digital avatar's expressions remain neutral throughout, resulting in low fun and playability, and an unsatisfactory call experience. To enable the digital avatar to support richer expressions, the call assistant would need to employ a more powerful model, consuming significant processing resources.

[0086] Based on this, this application provides a communication method to make the expressions of the digital human more natural when using a call assistant service, thereby improving the user experience without increasing the complexity of model processing. This method is applicable to the communication architecture shown in Figure 2A or Figure 2B above, and is implemented through the interaction of a first network element and a third network element. The first network element provides media resource management functions, such as the aforementioned MF. This first network element can be itself, a component within the first network element (e.g., a processor, chip, or chip system), or a logical module or software implementing all or part of the first network element. The third network element provides the time-frequency frames or image frames of the digital human corresponding to the call assistant service, which can be understood as a video material warehouse. This third network element is typically deployed in an operator's facility or a cloud server. This third network element can be itself, a component within the third network element (e.g., a processor, chip, or chip system), or a logical module or software implementing all or part of the third network element. This application does not specifically limit the scope of this method. Furthermore, when this communication method is applied to the communication architecture shown in Figure 2B, it also involves a second network element, which can be understood as a call assistant, such as the SAAF mentioned above. This is only illustrative and not specifically limited. Additionally, this communication method also involves terminal devices, data channel application servers, etc. Referring to Figure 3, the following is executed. The specific number of the first network element, second network element, third network element, terminal device, and data channel application server is not limited here; Figure 3 illustrates one as an example:

[0087] Step 301: The first network element receives a first notification, which indicates that the first call assistant service has been started. The first notification includes the identifier of the digital human corresponding to the first call assistant service.

[0088] Different users may subscribe to the same or different call assistant services. Each user can be associated with one or more terminal devices, and can use any of these devices to make a call and access the IMS network to activate the call assistant service. For example, user A's call assistant service subscription only includes voice services and does not include a digital persona. User B's call assistant service subscription corresponds to digital persona B. User C's call assistant service subscription corresponds to digital persona C. Users D and E's call assistant service subscriptions correspond to the same digital persona W. The image of digital persona W can be a public figure, such as actor W, artist W, or cartoon character W, etc. This is only an example and not a specific limitation.

[0089] Furthermore, subscribers to the First Call Assistant service have signed up for the corresponding digital persona. The First Call Assistant service may be a single call assistant service or multiple call assistant services. For example, the First Call Assistant service may include call assistant services for user 1 and user 2. This is merely an example and not a specific limitation.

[0090] In one possible implementation, the user makes a phone call through a terminal device, accesses the IMS network, and thus triggers the first notification.

[0091] In another possible implementation, after the data channel application server obtains the information to initiate the first call assistant service, it sends a first notification to the first network element. For example, the data channel application server is a DC AS (Distributed Data Assistant Server). The DC AS stores the user's call assistant subscription information and can determine when the call assistant service needs to be initiated. Upon determining that the agreed-upon time has arrived, the call assistant service is initiated, and a first notification is sent to the first network element (e.g., MF).

[0092] For example, the first notification may also include: the start time of the first call assistant service, the duration of the first call assistant service, etc.

[0093] Step 302: The first network element requests the first information according to the first notification. The first information is used to indicate the first expression description information of the digital human corresponding to the first call assistant service.

[0094] In one possible implementation, the first network element sends a second request message to the third network element based on a first notification. The second request message includes an identifier for the digital human corresponding to the first call assistant service and is used to request first expression description information. The third network element retrieves the first information based on the identifier of the digital human corresponding to the first call assistant service, and then sends the first information back to the first network element. The first network element obtains the first information through signaling interaction with the third network element, which does not increase the processing complexity of the first network element and can improve data processing efficiency.

[0095] In another possible implementation, the first network element forwards the first notification to the third network element to obtain the first information.

[0096] The first information can be either the description information of the first expression or the index information of the description information of the first expression. No specific limitation is made here.

[0097] The first expression description information can be offline-generated expression materials, which can be obtained based on an AI model and are not specifically limited here. This first expression description information may include an index of a certain emotion file of the digital human. For example, the first expression description information includes the start frame number, end frame number, or duration of a certain emotion of the digital human*. For example, the start and end frame numbers of the digital human X's happiness level 1, the start and end frame numbers of the digital human X's happiness level 2, the start and duration of the digital human X's sadness level 1, the start and end frame numbers of the digital human X's surprise level 1, and the start and end frame numbers of the digital human X's neutral level 1. Optionally, the first expression description information may also include the start and end frame numbers of different emotion transitions. For example, the first expression description information includes the start and end frame numbers of the digital human X's transition from happiness level 1 to neutral level 1. This first expression description information may also include specific emotion files. For example, the first expression description information includes all video frames or image frames of the digital human X at happiness level 1. This is merely an example to illustrate the description of the first emotional expression; its specific application is not limited. Furthermore, the degree of emotion can be predefined; for example, happiness can be set to three levels: slightly happy, moderately happy, and extremely happy. This is not specifically limited here.

[0098] Step 303: The first network element obtains the second information, which includes: the prompt information from the call assistant and the emotion information matching the prompt information.

[0099] In one possible implementation, the first network element is the MF in Figure 2A. This first network element is co-located with the second network element (e.g., SAAF). The first network element can directly implement the call assistant service, and the first network element can obtain the second information through internal data processing.

[0100] In another possible implementation, the first network element is the MF in Figure 2B. This first network element is separate from the second network element (e.g., SAAF). After receiving the user's multimedia information (audio, video, or text, etc.), the second network element can send second information to the first network element.

[0101] In the second message, the call assistant's response can be audio or voice text. For example, a user makes a phone call, accesses the IMS network, and activates the call assistant service, saying, "I completed an important task today." The call assistant's response could be the audio message, "Congratulations! I'm glad you made such progress." Alternatively, the response could be the text message, "Congratulations! I'm glad you made such progress." The matching emotion for the response is "Happy." This is merely an example and not a specific limitation.

[0102] Step 304: The first network element generates first video information based on the first information and the second information. The first video information is used to indicate the video of the digital human corresponding to the first call assistant service.

[0103] In one possible implementation, the first network element processes the first information and the second information to obtain the first video information. For example, the first information is a specific emotion file. The first network element can retrieve the first information based on the emotion information in the second information, obtain the emotion file matching that emotion information, and then perform video editing to integrate the prompt information into the video, thus obtaining the digital human's video.

[0104] In another possible implementation, the first network element determines second facial expression description information based on second information and first information. This second facial expression description information is associated with the emotion information matching the prompt information, and it belongs to the first facial expression description information. Third information is obtained based on the second facial expression description information. This third information is used to indicate the video frame or image frame of the digital human corresponding to the first call assistant service that matches the second facial expression description information. First video information is generated based on the second and third information (e.g., video and audio arrangement based on AI). After determining the facial expression description information that matches the second information, the corresponding digital human video frame or image frame is requested to generate a video of the digital human. The digital human's expressions in the video generated in this way are more natural.

[0105] In another possible implementation, the first network element determines second facial expression description information based on the second information and the first information. The second facial expression description information is associated with the emotion information matching the prompt information, and the second facial expression description information belongs to the first facial expression description information. First video information is generated based on the second facial expression description information and the second information. After determining the second facial expression description information that matches the second information, if the second facial expression description information is a specific emotion file, the first network element can generate the first video information based on the second facial expression description information and the second information.

[0106] In one specific implementation, the first network element can refer to the second information to filter out the second expression description information from the first information. For example, the first network element can search for the second expression description information corresponding to the emotion information in the first expression description information based on the emotion information in the second information. Alternatively, the first network element can input the second information and the first information into the AI ​​processing model and derive the second expression description information based on the AI ​​processing model. This is only an example and does not specifically limit how the second expression description information is determined.

[0107] In another specific implementation, the first network element can generate fourth information based on the second information. The fourth information is used to indicate the label of the emotional information that matches the prompt information. The first network element determines the second expression description information based on the fourth information and the first information. In this application, after the first network element generates the fourth information based on the second information, it searches for the first information based on the fourth information to determine the matching second expression description information, and then generates a video of the digital human. The expressions of the digital human in the video generated in this way are more natural. The fourth information is associated with the following information: the duration of the prompt information and the frame rate (e.g., 25 FPS) of the video of the digital human corresponding to the first call assistant service. When the fourth information is related to the duration of the prompt information and the frame rate of the video of the digital human, it can be guaranteed that the video of the digital human generated by the first network element is compatible with the prompt information and meets the requirements of the video configuration. For example, if the duration of the prompt information is 5 seconds and the frame rate of the video of the digital human is 25 FPS, then the length of the fourth information is 5*25. For example, the fourth information is a sequence used to indicate emotional information, such as h,h,h,h,h,h,n,n,n,n,n.

[0108] In another specific implementation, the first network element also receives fourth information from the second network element; the first network element determines second expression description information based on the fourth information and the first information. After obtaining the fourth information from the second network element, the first network element searches for the first information based on the fourth information to determine the matching second expression description information, thereby generating a video of the digital human. The expressions of the digital human in the video generated in this way are more natural.

[0109] As shown in Figure 4, the sequence indicating emotional information is h,h,h,h,h,h,h,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n. Based on this sequence, a second expression description file corresponding to the first information is selected. For example, the frame numbers of the video frames are [001,101,102,103,203,204,205,206,…], where the duration of the prompt information is 40 / 25 = 1.6 seconds, the video frame rate is 25 fps, and one video frame corresponds to one sequence element in the sequence.

[0110] The first network element can send a first request message to the third network element based on the second expression description information. The first request message includes the second expression description information and is used to request video frames or image frames of the digital human corresponding to the first call assistant service that matches the second expression description information. The third network element retrieves third information based on the second expression description information and then sends the third information to the first network element. The first network element obtains the third information through signaling interaction with the third network element, which does not increase the processing complexity of the first network element and can improve data processing efficiency.

[0111] Step 305: The first network element sends the first video information.

[0112] For example, the first network element may send the first video information to the terminal device, or the first network element may send the first video information to other applications (APPs). No specific limitations are specified here.

[0113] As shown in Figure 5, the offline-generated digital human facial expression materials are stored in the third network element. After the call assistant service is activated, the call assistant can provide feedback text information and corresponding emotional information (i.e., second information) based on the large language model. For example, the second information is "<Happy> The weather is great today, you can go on a picnic. <Normal> It will rain tomorrow, <Sad> You can only stay at home." The first network element generates a sequence indicating emotional information based on the second information, for example, h,h,h,h,h,h,h,h,h,n,n,n,n,n,n,n,s,s,s,s,s. The first network element determines the time-frequency frame or image frame corresponding to the sequence through facial expression arrangement, and then constructs the digital human's time-frequency based on the lip-sync model and the text information provided by the call assistant based on the large language model. For example, if different levels of happiness are represented by h1 (slight happiness), h2 (moderate happiness), and h3 (extreme happiness), then the sequence indicating emotional information could also be h1,h1,h2,h2,h3,h3, etc., which is only illustrated here as an example.

[0114] In this application, after confirming that the first call assistant service has been activated, the system requests the first facial expression information of the digital human corresponding to the first call assistant service. Then, based on the prompt information from the call assistant's response, the emotional information matching the prompt information, and the first facial expression description information, a digital human video is generated. Because the digital human video is generated based on the first facial expression information and the emotional information matching the prompt information, the digital human's expressions in the video are more natural. Furthermore, the digital human's expression generation does not involve complex processing models, requiring only minimal computational cost and enabling high-concurrency operation. Therefore, when using the call assistant service, the voice-driven digital human's expressions are more natural, which can improve the user's service experience.

[0115] To better illustrate the scheme of this application, a specific example is provided below. The example uses MF as the first network element, SAAF as the second network element, the material warehouse as the third network element, DC AS as the data channel application server, and UE1 and UE2 as terminal devices. In Figures 6-9 below, MF and SAAF are separate network elements, while in Figure 10, MF and SAAF are combined network elements. Optionally, in specific applications, interactions between other network elements may also be involved, which are not limited here.

[0116] Referring to Figure 6, it is assumed that the digital personalities corresponding to the call assistant services subscribed to by UE1 and UE2 are different, and are digital personalities specifically customized for UE1 and UE2 respectively. The following explanation uses UE1 as an example; UE2 can refer to it for understanding. In Figure 6, MF automatically generates the fourth information. The execution is as follows:

[0117] Step 601: UE1 responds to the user's call, accesses the IMS network, triggers the user's subscribed call assistant service 1, and sends the first notification to MF.

[0118] Step 602, MF sends a second request message to the material warehouse. The second request message includes: the identifier of digital human 1 corresponding to call assistant service 1.

[0119] Step 603: The material repository sends the expression description file corresponding to Digital Human 1 back to MF.

[0120] The expression description file corresponding to the digital human 1 is an index of a certain emotion file of the digital human 1, which can be understood by referring to the description in step 302 above, and will not be explained in detail here.

[0121] In step 604, after UE1 responds to the user's multimedia information, it sends multimedia information to SAAF.

[0122] For example, the multimedia message is the audio message "An important task was completed today!".

[0123] Step 605: SAAF generates second information based on the multimedia information using the LLM service.

[0124] The second message includes the text "Congratulations! We're so glad you've made such progress," along with the emotional message "Happy."

[0125] Step 606: SAAF sends the second message to MF.

[0126] Step 607: MF generates fourth information based on the second information.

[0127] Step 608: MF determines the second expression description information based on the fourth information and the first expression description information.

[0128] This can be understood by referring to the description in step 304 above, and will not be repeated here.

[0129] Step 609: MF sends a first request message to the material repository, which includes the second emoticon description information.

[0130] Step 610: The material warehouse sends a third message to MF.

[0131] Steps 609 and 610 can be understood by referring to the description in step 304 above, and will not be repeated here.

[0132] Step 611: MF generates first video information based on the third information and the second information.

[0133] Step 612: MF sends the first video information to UE1.

[0134] In this example, emoji materials are created offline and stored in an emoji repository. During a call, MF performs real-time emoji arrangement to achieve the goal of lightweight emoji-driven communication, thus associating emojis with LLM's response content.

[0135] Referring to Figure 7, it is assumed that the digital personalities corresponding to the call assistant services subscribed to by UE1 and UE2 are different, and are digital personalities specifically customized for UE1 and UE2 respectively. The following explanation uses UE1 as an example; UE2 can refer to it for understanding. In Figure 7, SAAF generates the fourth information. The execution is as follows:

[0136] Step 701: UE1 responds to the user's call, accesses the IMS network, triggers the user's subscribed call assistant service 1, and sends the first notification to MF.

[0137] In addition, the DC AS can also send a first notification to the MF when it determines the scheduled start time of the call assistant service 1 for UE1.

[0138] Step 702, MF sends a second request message to the material warehouse. The second request message includes: the identifier of digital human 1 corresponding to call assistant service 1.

[0139] Step 703: The material repository sends the expression description file corresponding to Digital Human 1 back to MF.

[0140] The expression description file corresponding to the digital human 1 is an index of a certain emotion file of the digital human 1, which can be understood by referring to the description in step 302 above, and will not be explained in detail here.

[0141] Step 704: After responding to the user's multimedia information, UE1 sends the multimedia information to SAAF.

[0142] For example, the multimedia message is the audio message "An important task was completed today!".

[0143] Step 705: SAAF generates second information based on the multimedia information using the LLM service.

[0144] The second message includes the text "Congratulations! We're so glad you've made such progress," along with the emotional message "Happy."

[0145] Step 706: SAAF generates fourth information based on the second information.

[0146] Step 707: SAAF sends the second and fourth messages to MF.

[0147] For example, SAAF sends the prompt message and the fourth message from the second message to MF.

[0148] Step 708: MF determines the second expression description information based on the fourth information and the first expression description information.

[0149] This can be understood by referring to the description in step 304 above, and will not be repeated here.

[0150] Step 709: MF sends a first request message to the material repository, which includes the second emoticon description information.

[0151] Step 710: The material warehouse sends a third message to MF.

[0152] Steps 709 and 710 can be understood by referring to the description in step 304 above, and will not be repeated here.

[0153] Step 711: MF generates first video information based on the third information and the second information.

[0154] Step 712: MF sends the first video information to UE1.

[0155] In this example, emoji materials are created offline and stored in an emoji repository. During a call, MF arranges emojis in real time based on the fourth information fed back by SAAF, achieving the goal of lightweight emoji-driven communication and associating emojis with the LLM's response content.

[0156] Referring to Figure 8, it is assumed that UE1 and UE2 have the same digital persona associated with their call assistant service, for example, a certain artist. The following explanation uses UE1 as an example; UE2 can refer to it for understanding. In Figure 8, MF automatically generates the fourth piece of information. The execution is as follows:

[0157] Step 801: UE1 responds to the user's call, accesses the IMS network, triggers the user's subscribed call assistant service 1, and sends the first notification to MF.

[0158] Step 802, MF sends a second request message to the material warehouse. The second request message includes: the identifier of digital human 1 corresponding to call assistant service 1.

[0159] Step 803: The material repository sends the expression description file corresponding to Digital Human 1 back to MF.

[0160] Among them, the artist's image data is relatively small, so all emotion files can be downloaded at once when the digital human service is launched. Therefore, the expression description file corresponding to digital human 1 is all the emotion files of digital human 1, which can be understood by referring to the description in step 302 above, and will not be elaborated here.

[0161] Step 804: After responding to the user's multimedia information, UE1 sends the multimedia information to SAAF.

[0162] For example, the multimedia message is the audio message "An important task was completed today!".

[0163] Step 805: SAAF generates second information based on the multimedia information using the LLM service.

[0164] The second message includes the text "Congratulations! We're so glad you've made such progress," along with the emotional message "Happy."

[0165] Step 806, SAAF sends the second message to MF.

[0166] Step 807: MF generates fourth information based on the second information.

[0167] Step 808: MF determines the second expression description information based on the fourth information and the first expression description information.

[0168] This can be understood by referring to the description in step 304 above, and will not be repeated here.

[0169] Step 809: MF generates first video information based on the second expression description information and the second information.

[0170] Step 810: MF sends the first video information to UE1.

[0171] In this example, emoji materials are created offline and stored in an emoji repository. During a call, MF performs real-time emoji arrangement to achieve the goal of lightweight emoji-driven communication, thus associating emojis with LLM's response content.

[0172] Referring to Figure 9, it is assumed that UE1 and UE2 have the same digital persona associated with their subscribed call assistant service, for example, a certain artist. The following explanation uses UE1 as an example; UE2 can refer to it for understanding. In Figure 9, SAAF generates the fourth piece of information. The execution is as follows:

[0173] Step 901: UE1 responds to the user's call, accesses the IMS network, triggers the user's subscribed call assistant service 1, and sends the first notification to MF.

[0174] Step 902, MF sends a second request message to the material warehouse. The second request message includes: the identifier of digital human 1 corresponding to call assistant service 1.

[0175] Step 903: The material repository sends the expression description file corresponding to Digital Human 1 back to MF.

[0176] Among them, the artist's image data is relatively small, so all emotion files can be downloaded at once when the digital human service is launched. Therefore, the expression description file corresponding to digital human 1 is all the emotion files of digital human 1, which can be understood by referring to the description in step 302 above, and will not be elaborated here.

[0177] Step 904: After responding to the user's multimedia information, UE1 sends the multimedia information to SAAF.

[0178] For example, the multimedia message is the audio message "An important task was completed today!".

[0179] Step 905: SAAF generates second information based on the multimedia information using the LLM service.

[0180] The second message includes the text "Congratulations! We're so glad you've made such progress," along with the emotional message "Happy."

[0181] Step 906: SAAF generates fourth information based on the second information.

[0182] Step 907, SAAF sends the second and fourth messages to MF.

[0183] Step 908: MF determines the second expression description information based on the fourth information and the first expression description information.

[0184] This can be understood by referring to the description in step 304 above, and will not be repeated here.

[0185] Step 909: MF generates first video information based on the second expression description information and the second information.

[0186] Step 910: MF sends the first video information to UE1.

[0187] In this example, emoji materials are created offline and stored in an emoji repository. During a call, MF arranges emojis in real time based on the fourth information fed back by SAAF, achieving the goal of lightweight emoji-driven communication and associating emojis with the LLM's response content.

[0188] Referring to Figure 10, it is assumed that the digital personalities corresponding to the call assistant services subscribed to by UE1 and UE2 are different, and are digital personalities specifically customized for UE1 and UE2 respectively. The following explanation uses UE1 as an example; UE2 can refer to it for understanding. MF and SAAF are jointly set up and executed as follows:

[0189] Step 1001: UE1 responds to the user's call, accesses the IMS network, triggers the user's subscribed call assistant service 1, and sends the first notification to MF.

[0190] Step 1002, MF sends a second request message to the material warehouse. The second request message includes: the identifier of digital human 1 corresponding to call assistant service 1.

[0191] Step 1003: The material warehouse sends the expression description file corresponding to Digital Human 1 back to MF.

[0192] The expression description file corresponding to the digital human 1 is an index of a certain emotion file of the digital human 1, which can be understood by referring to the description in step 302 above, and will not be explained in detail here.

[0193] Step 1004: After responding to the user's multimedia information, UE1 sends the multimedia information to MF.

[0194] For example, the multimedia message is the audio message "An important task was completed today!".

[0195] Step 1005: MF generates second information based on the multimedia information using the LLM service.

[0196] The second message includes the text "Congratulations! We're so glad you've made such progress," along with the emotional message "Happy."

[0197] Step 1006: MF generates fourth information based on the second information.

[0198] Step 1007: MF determines the second expression description information based on the fourth information and the first expression description information.

[0199] This can be understood by referring to the description in step 304 above, and will not be repeated here.

[0200] Step 1008: MF sends a first request message to the material repository, which includes second emoticon description information.

[0201] Step 1009: The material warehouse sends a third message to MF.

[0202] Steps 609 and 610 can be understood by referring to the description in step 304 above, and will not be repeated here.

[0203] Step 1010: MF generates first video information based on the third information and the second information.

[0204] Step 1011: MF sends the first video information to UE1.

[0205] In this example, emoji materials are created offline and stored in an emoji repository. During a call, MF performs real-time emoji arrangement to achieve the goal of lightweight emoji-driven communication, thus associating emojis with LLM's response content.

[0206] The foregoing primarily describes the solutions provided by the embodiments of this application from the perspective of device interaction. It is understood that, in order to achieve the above functions, each device may include corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0207] The embodiments of this application can divide the device into functional units according to the above method examples. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0208] In the case of using integrated units, FIG11 shows a possible exemplary block diagram of the communication device involved in the embodiments of this application. As shown in FIG11, the communication device 1100 may include a processing unit 1101 and a transceiver unit 1102. The processing unit 1101 is used to control and manage the operation of the communication device 1100. The transceiver unit 1102 is used to support communication between the communication device 1100 and other devices. Optionally, the transceiver unit 1102 may include a receiving unit and / or a transmitting unit, respectively used to perform receiving and transmitting operations. Optionally, the communication device 1100 may also include a storage unit for storing the program code and / or data of the communication device 1100. The transceiver unit may be called an input / output unit, a communication unit, etc., and the transceiver unit may be a transceiver; the processing unit may be a processor. When the communication device is a module (e.g., a chip) in a communication device, the transceiver unit may be an input / output interface, an input / output circuit, or an input / output pin, etc., and may also be called an interface, a communication interface, or an interface circuit, etc.; the processing unit may be a processor, a processing circuit, or a logic circuit, etc. For example, the device can be the first network element or the second network element described above.

[0209] More detailed descriptions of the above-mentioned processing unit 1101 and transceiver unit 1102 can be obtained directly from the relevant descriptions in the above-mentioned method embodiments, and will not be repeated here.

[0210] Figure 12 shows a communication device 1200 provided in this application. The communication device 1200 can be a chip or a chip system. This communication device can be located in the device involved in any of the above method embodiments, such as a terminal or network device, to perform the actions corresponding to that device.

[0211] Optionally, a chip system can consist of chips or include chips and other discrete components.

[0212] The communication device 1200 includes a processor 1210.

[0213] The processor 1210 is configured to execute a computer program stored in the memory 1220 to implement the operation of each device in any of the above method embodiments.

[0214] The communication device 1200 may also include a memory 1220 for storing computer programs.

[0215] Optionally, the memory 1220 and the processor 1210 are coupled. Coupling is an indirect coupling or communication connection between devices, units, or modules, which can be electrical, mechanical, or other forms, for information exchange between devices, units, or modules. Optionally, the memory 1220 and the processor 1210 are integrated together.

[0216] There can be one or more processors 1210 and memory 1220, and there is no limitation.

[0217] Optionally, in practical applications, the communication device 1200 may or may not include a transceiver 1230, as illustrated by the dashed box in the figure. The communication device 1200 can exchange information with other devices through the transceiver 1230. The transceiver 1230 can be a circuit, a bus, a transceiver, or any other device that can be used for information exchange.

[0218] In one possible implementation, the communication device 1200 can be a terminal or network device as described in the above-described methods.

[0219] This application embodiment does not limit the specific connection medium between the transceiver 1230, processor 1210, and memory 1220. In Figure 12, the memory 1220, processor 1210, and transceiver 1230 are connected via a bus, indicated by a thick line. The connection methods between other components are merely illustrative and not intended to be limiting. The bus can be an address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used in Figure 12, but this does not indicate that there is only one bus or one type of bus. In this application embodiment, the processor can be a general-purpose processor, digital signal processor, application-specific integrated circuit, field-programmable gate array, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, and can implement or execute the methods, steps, and logic block diagrams disclosed in this application embodiment. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in this application embodiment can be directly reflected as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0220] In the embodiments of this application, the memory can be non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or it can be volatile memory, such as random-access memory (RAM). The memory can also be any other medium capable of carrying or storing desired program code in the form of instructions or data structures, and accessible by a computer, but is not limited thereto. The memory in the embodiments of this application can also be a circuit or any other device capable of implementing storage functions, used to store computer programs, program instructions, and / or data.

[0221] Based on the above embodiments, referring to Figure 13, this application embodiment also provides another communication device 1300, including: an interface circuit 1310 and a logic circuit 1320; the interface circuit 1310 can be understood as an input / output interface, which can be used to execute the transmission and reception steps of each device in any of the above method embodiments, and the logic circuit 1320 can be used to run code or instructions to execute the methods executed by each device in any of the above embodiments, which will not be described in detail again.

[0222] Based on the above embodiments, this application also provides a computer-readable storage medium storing instructions that, when executed, cause the methods executed by the devices in any of the above method embodiments to be implemented. The computer-readable storage medium may include various media capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory, random access memory, magnetic disk, or optical disk.

[0223] Based on the above embodiments, this application provides a communication system, which includes the terminal and network device mentioned in any of the above method embodiments, and can be used to execute the methods executed by each device in any of the above method embodiments.

[0224] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0225] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.

[0226] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0227] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

Claims

1. A communication method characterized by comprising: Applied to the first network element, including: Receive a first notification, the first notification being used to indicate that the first call assistant service has been started, the first notification including: the identifier of the digital human corresponding to the first call assistant service; According to the first notification request for first information, the first information is used to indicate the first facial expression description information of the digital human corresponding to the first call assistant service; Obtain second information, which includes: prompt information from the call assistant's response and emotion information matching the prompt information; Based on the first information and the second information, first video information is generated, which is used to indicate the video of the digital human corresponding to the first call assistant service. Send the first video information.

2. The method of claim 1, wherein, The step of generating the first video information based on the first information and the second information includes: The second expression description information is determined based on the second information and the first information. The second expression description information is associated with the emotion information matched by the prompt information. The second expression description information belongs to the first expression description information. The third information is obtained based on the second expression description information. The third information is used to indicate the video frame or image frame of the digital human corresponding to the first call assistant service that matches the second expression description information. The first video information is generated based on the second information and the third information.

3. The method of claim 1, wherein, The step of generating the first video information based on the first information and the second information includes: The second expression description information is determined based on the second information and the first information. The second expression description information is associated with the emotion information matched by the prompt information. The second expression description information belongs to the first expression description information. The first video information is generated based on the second facial expression description information and the third information.

4. The method according to claim 2 or 3, characterized in that, Determining the second facial expression description information based on the second information and the first information includes: The fourth information is generated based on the second information, and the fourth information is used to indicate the label of the emotion information that matches the prompt information; The second expression description information is determined based on the fourth information and the first information.

5. The method according to claim 2 or 3, characterized in that, The method further includes: Receive fourth information from the second network element, the fourth information being used to indicate a tag of emotion information that matches the prompt information; Determining the second facial expression description information based on the second information and the first information includes: The second expression description information is determined based on the fourth information and the first information.

6. The method according to any one of claims 2-5, characterized in that, The step of obtaining the third information based on the second facial expression description information includes: A first request message is sent to a third network element according to the second expression description information. The first request message includes the second expression description information. The first request message is used to request a video frame or image frame of the digital human corresponding to the first call assistant service that matches the second expression description information. Receive the third information from the third network element.

7. The method according to any one of claims 4-6, characterized by, The fourth piece of information is associated with the following: The duration of the prompt message is related to the frame rate of the video of the digital human corresponding to the first call assistant service.

8. The method according to any one of claims 4-7, characterized in that, The fourth piece of information is a sequence used to indicate the emotional information.

9. The method according to any one of claims 1-8, characterized in that, The request for the first information based on the first notification includes: According to the first notification, a second request message is sent to the third network element. The second request message includes: the identifier of the digital human corresponding to the first call assistant service. The second request message is used to request the first expression description information. Receive the first information from the third network element.

10. The method according to any one of claims 1-9, characterized in that, The acquisition of the second information includes: Receive the second information from the second network element.

11. A communication method, characterized in that, Applied to the second network element, including: Receive multimedia information; The second information is obtained based on the multimedia information, and the second information includes: prompt information from the call assistant's reply and emotion information matching the prompt information; Send the second information to the first network element.

12. The method according to claim 11, characterized in that, The method further includes: The fourth information is generated based on the second information, and the fourth information is used to indicate the label of the emotion information that matches the prompt information; The fourth information is sent to the first network element.

13. The method according to claim 11 or 12, characterized in that, The fourth piece of information is associated with the following: The duration of the prompt message is related to the frame rate of the video of the digital human corresponding to the first call assistant service.

14. The method according to any one of claims 11-13, characterized in that, The fourth piece of information is a sequence used to indicate the emotional information.

15. A communication method, characterized in that, Applied to third-party network elements, including: Receive a second request message from the first network element, the second request message including: the identifier of the digital human corresponding to the first call assistant service, the second request message being used to request the first expression description information; Send first information to the first network element, the first information being used to indicate the first expression description information of the digital human corresponding to the first call assistant service.

16. The method according to claim 15, characterized in that, The method further includes: Receive a first request message from a first network element, the first request message including: second expression description information, the first request message is used to request a video frame or image frame of the digital human corresponding to the first call assistant service that matches the second expression description information; Send a third message to the first network element, the third message being used to indicate the video frame or image frame of the digital human corresponding to the first call assistant service that matches the second expression description information.

17. A communication device, characterized in that, include: A first network element for performing the method as described in any one of claims 1 to 10, or a second network element for performing the method as described in any one of claims 11 to 14, or a third network element for performing the method as described in any one of claims 15 to 16.

18. A communication device, characterized in that, It includes at least one processor; and a communication interface communicatively connected to said at least one processor; said at least one processor causes the method as described in any one of claims 1 to 16 to be executed by executing instructions stored in memory.

19. A computer-readable storage medium, characterized in that, The computer contains a computer program or instructions that, when executed on a computer, cause the computer to perform the method as described in any one of claims 1 to 16.

20. A computer program product, characterized in that, When the computer reads and executes the computer program product, the method described in any one of claims 1 to 16 is performed.