Communication method and device

By using the first network element to generate natural digital human expression videos based on prompt information and emotional information in the call assistant service, the problems of large resource consumption and unnatural expressions caused by complex models in the prior art are solved, and the user experience is improved.

CN120075533APending Publication Date: 2025-05-30XIAN RUIXIN TECH CO LTD

Patent Information

Application Number
CN202510082066.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the current call assistant business, driving digital human expressions requires complex processing models, resulting in large resource consumption and poor expression performance, which affects the user experience.

Method used

After receiving the notification through the first network element, a digital expression video information is requested and generated. Based on the call assistant's prompt information and emotional information, a natural and uncomplicated expression expression is generated.

Benefits of technology

It realizes that without increasing the complexity of model processing, it improves the naturalness of digital human expressions and improves the user's interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075533A_ABST
    Figure CN120075533A_ABST
Patent Text Reader

Abstract

The invention provides a communication method and device, and relates to the technical field of communication. A first notification is received, the first notification is used for indicating that the first call assistant service is started, and the first notification comprises an identifier of a digital person corresponding to the first call assistant service; requesting first information according to the first notification, wherein the first information is used for indicating first expression description information of a digital person corresponding to the first call assistant service; obtaining second information, wherein the second information comprises prompt information replied by the call assistant and emotion information matched with the prompt information; according to the first information and the second information, first video information is generated, and the first video information is used for indicating a video of a digital person corresponding to the first call assistant service; and sending the first video information. The video of the digital person is generated based on the first expression expression information and the emotion information matched with the prompt information, so that the expression of the digital person in the video of the digital person is more natural. Therefore, the service experience of the user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a communication method and apparatus. Background Art

[0002] A Digital Human is a digital human image created by using digital technologies and similar to the human image. Digital Humans break physical boundaries to provide anthropomorphic services, featuring hyper-realism and strong interactivity, which is the future development trend.

[0003] When a user uses the call assistant service, the digital human image of the call assistant can be driven by voice, which can improve the user's interaction experience. However, in the current call assistant service, driving the expression of the digital human requires a complex processing model to achieve, consuming a large amount of processing resources. Summary of the Invention

[0004] This application provides a communication method and apparatus, so that when the call assistant service is adopted, the expression of the digital human can be driven more naturally, the user experience can be improved, and the complexity of model processing is not increased.

[0005] In a first aspect, this application provides a communication method, which can be applied to a first network element for providing a media resource management function. The first network element can be the first network element itself, or a component in the first network element (such as a processor, a chip, or a chip system, etc.), or can also be a logical module or software that implements all or part of the functions of the first network element. This application does not specifically limit this here. The method is executed as follows:

[0006] Receive a first notification, where the first notification is used to indicate that a first call assistant service has been started, and the first notification includes: an identifier of the digital human corresponding to the first call assistant service; request first information according to the first notification, where the first information is used to indicate first expression description information of the digital human corresponding to the first call assistant service; obtain second information, where the second information includes: a prompt message replied by the call assistant and emotion information matching the prompt message; generate first video information according to the first information and the second information, where the first video information is used to indicate a video of the digital human corresponding to the first call assistant service; and send the first video information.

[0007] Among them, the prompt message can be audio or voice text. This is only an exemplary illustration here and is not specifically limited.

[0008] In this application, after determining that the first call assistant service has been started, request the first expression description information of the digital human corresponding to the first call assistant service, and then generate a digital human video based on the prompt information replied by the call assistant, the emotion information matching the prompt information, and the first expression description information. Since the video of this digital human is generated based on the first expression description information and the emotion information matching the prompt information, the expression of the digital human in the video of this digital human is more natural. In addition, the generation of the expression of this digital human does not introduce a model with complex processing, only requires a small amount of computing cost, and can run in high concurrency. Based on this, when using the call assistant service, the expression of the voice-driven digital human is more natural, which can improve the user's service experience.

[0009] In an optional manner, the first network element determines the second expression description information according to the second information and the first information. The second expression description information is associated with the emotion information matching the prompt information, and the second expression description information belongs to the first expression description information; obtain the third information according to the second expression description information, and the third information is used to indicate the video frame or picture frame of the digital human corresponding to the first call assistant service that matches the second expression description information; generate the first video information according to the second information and the third information.

[0010] After determining the expression description information matching the second information, request the corresponding digital human video frame or picture frame to generate the video of the digital human. Based on this, the expression of the digital human in the generated video of the digital human is more natural.

[0011] In an optional manner, the first network element determines the second expression description information according to the second information and the first information. The second expression description information is associated with the emotion information matching the prompt information, and the second expression description information belongs to the first expression description information; generate the first video information according to the second expression description information and the third information.

[0012] After determining the second expression description information matching the second information, when the second expression description information is a specific emotion file, the first network element can generate the first video information based on the second expression description information and the second information.

[0013] In an optional manner, the first network element generates the fourth information according to the second information, and the fourth information is used to indicate the label of the emotion information matching the prompt information; determine the second expression description information according to the fourth information and the first information.

[0014] In this application, after the first network element generates the fourth information based on the second information, it searches for the first information based on the fourth information and then determines the second expression description information matching it, and then generates the video of the digital human. Based on this, the expression of the digital human in the generated video of the digital human is more natural.

[0015] In an alternative manner, the first network element also receives fourth information from the second network element, where the fourth information is used to indicate a tag of the emotion information that matches the prompt information; the second expression description information is determined according to the fourth information and the first information.

[0016] In this application, after the first network element obtains the fourth information from the second network element, it searches for the first information based on the fourth information, and then determines the second expression description information that matches it, and then generates a video of the digital human. The expression of the digital human in the video of the digital human generated based on this is more natural.

[0017] In an alternative manner, the first network element sends a first request message to the third network element according to the second expression description information, where the first request message includes: the second expression description information, and the first request message is used to request a video frame or a picture frame of the digital human corresponding to the first call assistant service that matches the second expression description information; the third information is received from the third network element.

[0018] The first network element obtains the third information through signaling interaction with the third network element, which does not increase the processing complexity of the first network element and can improve the data processing efficiency.

[0019] In an alternative manner, the fourth information is associated with the following information: the duration of the prompt information and the frame rate of the video of the digital human corresponding to the first call assistant service.

[0020] When the fourth information is related to the duration of the prompt information and the frame rate of the video of the digital human, it can ensure that the video of the digital human generated by the first network element is adapted to the prompt information and meets the requirements of the adapted video configuration.

[0021] In an alternative manner, the fourth information is a sequence used to indicate emotion information.

[0022] When the fourth information is a sequence, it is convenient for the first network element to perform recognition processing and improves the data processing efficiency.

[0023] In an alternative manner, the first network element sends a second request message to the third network element according to the first notification, where the second request message includes: the identifier of the digital human corresponding to the first call assistant service, and the second request message is used to request the first expression description information; the first information is received from the third network element.

[0024] The first network element obtains the first information through signaling interaction with the third network element, which does not increase the processing complexity of the first network element and can improve the data processing efficiency.

[0025] In an alternative manner, the first network element receives second information from the second network element.

[0026] Second aspect, the present application provides a communication method, which can be applied to a second network element for providing a call assistant service. The second network element can be the second network element itself, or a component in the second network element (such as a processor, a chip, or a chip system, etc.), or can also be a logic module or software that implements all or part of the functions of the second network element. The present application does not specifically limit this here. The execution is as follows:

[0027] Receive multimedia information; obtain second information according to the multimedia information, where the second information includes: the prompt information replied by the call assistant and the emotion information matching the prompt information; send the second information to the first network element.

[0028] In an optional manner, the second network element also generates fourth information according to the second information, where the fourth information is used to indicate the label of the emotion information matching the prompt information; send the fourth information to the first network element.

[0029] In an optional manner, the fourth information is associated with the following information: the duration of the prompt information and the frame rate of the video of the digital human corresponding to the first call assistant service.

[0030] In an optional manner, the fourth information is a sequence used to indicate emotion information.

[0031] Third aspect, the present application provides a communication method, which can be applied to a third network element for providing a digital human expression description file. The third network element can be the third network element itself, or a component in the third network element (such as a processor, a chip, or a chip system, etc.), or can also be a logic module or software that implements all or part of the functions of the third network element. The present application does not specifically limit this here. The execution is as follows:

[0032] Receive a second request message from the first network element, where the second request message includes: the identifier of the digital human corresponding to the first call assistant service, and the second request message is used to request the first expression description information; send the first information to the first network element, where the first information is used to indicate the first expression description information of the digital human corresponding to the first call assistant service.

[0033] In an optional manner, the third network element also receives a first request message from the first network element, where the first request message includes: the second expression description information, and the first request message is used to request the video frame or picture frame of the digital human corresponding to the first call assistant service that matches the second expression description information; send the third information to the first network element, where the third information is used to indicate the video frame or picture frame of the digital human corresponding to the first call assistant service that matches the second expression description information.

[0034] Fourthly, an embodiment of the present application provides a communication device, which may be a first network element, a second network element, or a third network element. The communication device is capable of implementing the functions of the above first to third aspects. For example, the communication device includes modules or units or means corresponding to the steps involved in the above first to third aspects. The functions or units or means may be implemented by software, or by hardware, or by hardware executing corresponding software.

[0035] In a possible design, the communication device includes a processing unit and a transceiver unit. The transceiver unit may be used to transmit and receive signals to achieve communication between the communication device and other devices. The processing unit may be used to perform some internal operations of the communication device. The transceiver unit may be referred to as an input / output unit, a communication unit, etc. The transceiver unit may be a transceiver. The processing unit may be a processor. When the communication device is a module (such as a chip) in a communication device, the transceiver unit may be an input / output interface, an input / output circuit, or input / output pins, etc., and may also be referred to as an interface, a communication interface, or an interface circuit, etc. The processing unit may be a processor, a processing circuit, or a logic circuit, etc.

[0036] In another possible design, the communication device includes a processor and may further include a transceiver. The transceiver is used to transmit and receive signals. The processor executes program instructions to complete the methods in any possible design or implementation manner of the above first to third aspects. The communication device may further include one or more memories. The memories are used to be coupled with the processor. The memories may store necessary computer programs or instructions for implementing the functions involved in the above first to third aspects. The processor may execute the computer programs or instructions stored in the memories. When the computer programs or instructions are executed, the communication device implements the methods in any possible design or implementation manner of the above first to third aspects.

[0037] In another possible design, the communication device includes a processor. The processor may be used to be coupled with a memory. The memory may store necessary computer programs or instructions for implementing the functions involved in the above first to third aspects. The processor may execute the computer programs or instructions stored in the memory. When the computer programs or instructions are executed, the communication device implements the methods in any possible design or implementation manner of the above first to third aspects.

[0038] In another possible design, the communication device includes a processor and an interface circuit. The processor is used to communicate with other devices through the interface circuit and execute the methods in any possible design or implementation manner of the above first to third aspects.

[0039] Understandably, in the above fourth aspect, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor that implements by reading software code stored in a memory. In addition, the above processors can be one or more, and the memories can be one or more. The memory can be integrated with the processor, or the memory and the processor can be separately arranged. In the specific implementation process, the memory can be integrated with the processor on the same chip, or can be separately arranged on different chips. The embodiments of the present application do not limit the type of the memory and the setting manner of the memory and the processor.

[0040] In a fifth aspect, an embodiment of the present application provides a communication system, which includes the above first network element, second network element, and third network element. Among them, the first network element can be used to execute the method in the first aspect, the second network element can be used to execute the method in the second aspect, and the third network element can be used to execute the method in the third aspect. In addition, it should be noted that there may be processes in which multiple devices or network elements interact and execute in each aspect, and the corresponding processes cannot be executed by a single device or network element. Mainly, the corresponding processes are executed by the corresponding devices or network elements interacting with each other, which will not be elaborated here.

[0041] In a sixth aspect, the present application provides a chip system, which includes a processor and may further include a memory for implementing the methods described in the above first aspect to third aspect. The chip system can be composed of chips, or can include chips and other discrete devices.

[0042] In a seventh aspect, the present application further provides a computer-readable storage medium, in which computer-readable instructions are stored. When the computer-readable instructions run on a computer, the computer is enabled to execute the methods in the first aspect to third aspect.

[0043] In an eighth aspect, the present application provides a computer program product containing instructions, which when running on a computer, enables the computer to execute the methods of the embodiments in the above first aspect to third aspect.

[0044] The technical effects that can be achieved in the above second aspect to eighth aspect can be described with reference to the technical effects that can be achieved by the corresponding possible design solutions in the above first aspect. The present application will not repeat them here. Description of the Drawings

[0045] Figure 1A Shows a schematic diagram of a communication scenario;

[0046] Figure 1B Shows a schematic diagram of another communication scenario;

[0047] Figure 2A shows a schematic diagram of a communication architecture;

[0048] Figure 2B shows another schematic diagram of a communication architecture;

[0049] Figure 3 shows a schematic flowchart of a communication method provided by an embodiment of the present application;

[0050] Figure 4 shows a schematic diagram of determining a second expression description file provided by an embodiment of the present application;

[0051] Figure 5 shows a schematic flowchart of a data processing process provided by an embodiment of the present application;

[0052] Figures 6 to 10 shows a schematic flowchart of a communication method provided by an embodiment of the present application;

[0053] Figure 11 shows a schematic structural diagram of a communication device provided by an embodiment of the present application;

[0054] Figure 12 shows a schematic structural diagram of a communication device provided by an embodiment of the present application;

[0055] Figure 13 shows a schematic structural diagram of a communication device provided by an embodiment of the present application. Detailed implementation manners

[0056] In the embodiments of the present application, the solutions in the embodiments can be combined and used reasonably, and the explanations or descriptions of each term, similar operations, or steps appearing in the embodiments can be referred to or explained mutually in each embodiment, which is not limited herein.

[0057] In the embodiments of the present application, "transmission" includes "sending" and / or "receiving". Among them, "sending" and "receiving" indicate the direction of signal transmission. For example, "sending information to XX" can be understood as the destination of the information being XX, which can include directly sending through the air interface, and also include indirectly sending through the air interface by other units or modules. "Receiving information from YY" can be understood as the source of the information being YY, which can include directly receiving from YY through the air interface, and can also include indirectly receiving from YY through the air interface from other units or modules. "Sending" can also be understood as the "output" of the chip interface, and "receiving" can also be understood as the "input" of the chip interface. In other words, sending and receiving can be carried out between devices. For example, between an access network device and a terminal device, or can also be carried out within a device. For example, sending or receiving between components, modules, chips, software modules or hardware modules within a device through a bus, trace or interface.

[0058] In the embodiments of the present application, for the number of nouns, unless otherwise specified, it means "singular noun or plural noun", that is, "one or more". "At least one" means one or more, and "multiple" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A / B can be singular or plural. The character " / " generally indicates that the front and back associated objects are in an "or" relationship. For example, A / B means: A or B. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b and (or) c means the following combinations: a exists alone, b exists alone, c exists alone, a and b exist simultaneously, a and c exist simultaneously, b and c exist simultaneously, or a, b and c exist simultaneously, where a, b, c can be single or multiple.

[0059] In the embodiments of the present application, "when...", "if" and "in case" all mean that the device will perform corresponding processing under certain objective circumstances, not limited to time, and it is not required that the device must have a judgment action when implemented, nor does it mean that there are other limitations. Unless otherwise specified, "if" and "in case" can be replaced, and "when..." and "in the case of..." can be replaced. "When..." and "if" / "in case" can be replaced. "Associate" and "correspond" can be replaced.

[0060] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0061] The ordinal numbers such as "first" and "second" mentioned in the embodiments of the present application are used to distinguish multiple objects and are not used to limit the size, content, order, timing, priority, or importance of multiple objects. For example, the first parameter and the second parameter refer to two different parameters, and do not indicate differences in the priority or importance of these two parameters.

[0062] The technical solutions provided in the embodiments of the present application can be applied to various communication systems, such as cellular systems related to the 3rd generation partnership project (3GPP), for example, long term evolution (LTE) communication systems, the sixth generation (5G) mobile communication systems / new radio (NR) communication systems, or future evolution systems, or other similar communication systems. Other similar communication systems include, for example, wireless fidelity (WiFi), vehicle to everything (V2X), spark link systems, Bluetooth systems, near field communication systems, internet of things (IoT) systems, such as ambient IoT (A-IoT / A-IoT), narrowband internet of things (NB-IoT), etc. Alternatively, the solutions provided in the embodiments of the present application can also be applied to communication systems that integrate two or more of the above systems. It should be understood that IoT technology is widely used in various industrial fields. For example, IoT technology can be applied to scenarios such as logistics, warehousing, industrial manufacturing, identity recognition, or environmental monitoring.

[0063] Figure 1A An application scenario applicable to the present application is shown. The user calls the call assistant based on the IP multimedia subsystem (IMS) network, and the call assistant provides the services required by the user. Exemplarily, as Figure 1AAs shown, the user requests the call assistant to "help query the weather", "help query scenic spots", "help order food", "remind of the basketball game on Saturday morning", etc., and the call assistant can feedback to the user "weather conditions", "scenic spot information", "food ordering situation", and "remind of the basketball game", etc. Among them, the call assistant is a service subscribed by the user, and the call assistant is implemented based on artificial intelligence (AI). Optionally, the call assistant is implemented based on a large language model (LLM). Among them, Figure 1A What is shown is a customer to manufacturer (C2M) scenario, that is, a scenario where the user directly calls the call assistant.

[0064] Figure 1B Another application scenario applicable to this application is shown. User A calls User B based on the IMS network, and services are provided based on the call assistant subscribed by User A. Exemplarily, as Figure 1B shown, User A communicates with User B based on the IMS network and asks B "Are you free tonight? Let's have dinner together". After User B agrees to the dinner request, User A requests the call assistant to "help order food", and the call assistant then orders food for User A and User B. Among them, Figure 1A What is shown is a customer (consumer) to customer (consumer) (C2C) scenario, that is, a scenario where the user calls the user. During the call, the user who subscribes to the call gravity wakes up the call assistant, and the call assistant helps the user complete the matter.

[0065] The following introduces the network architecture applicable to the communication method proposed in this application. Refer to Figure 2A which is a schematic diagram of a network architecture provided by an embodiment of this application. Figure 2AThe network architecture shown includes a proxy call session control function (P-CSCF) network element, a serving call session control function (S-CSCF) network element, an IMS access media gateway (AGW), a home subscriber server (HSS), an IMS application server (AS), a media function (MF) network element, a data channel signaling function (DCSF) network element, a network exposure function (NEF) network element, a terminal device, a data channel application server (DC AS), and mobile IMS. Among them, an AI assistant logic unit (subscriber AI agent function, SAAF) is also set in the MF. Optionally, the network architecture further includes a data channel application repository (DCAR). In Figure 2B 's architecture, the SAAF and the MF are separately set.

[0066] Here, a brief introduction is given to Figures 2A to 2B each network element (or functional network element, functional entity, node, device, etc.) shown in

[0067] 1. Terminal device: A terminal device can be referred to as a user equipment (UE), terminal, terminal device, access terminal, user unit, user station, mobile station, mobile station (MS), mobile terminal (MT), remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent, or user device, etc., and is a device with wireless or wired communication capabilities. For example, a terminal device can be connected to a wireless access device through an air interface, or a terminal device can be connected to a wired access device through a wired interface. In terms of product form, a terminal device can include a handheld device with communication functions, a vehicle-mounted device, a wearable device, or a computing device, etc. Exemplarily, a terminal device can be a mobile phone, a tablet (pad), a computer with wireless communication functions (such as a laptop computer, a handheld computer, etc.), a mobile internet device (MID), a virtual reality (VR) terminal, an augmented reality (AR) terminal, a smart watch, a smart bracelet, smart glasses, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, a cellular phone, a cordless phone, a SIP phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication functions, other processing devices connected to a wireless modem, a terminal in an internet of things (IoT) system, a desktop computer, or a UE defined by the 3rd generation partnership project (3GPP) standard specifications, etc., without limitation.

[0068] 2. P-CSCF network element: The entry node for a user to access the IMS network, mainly responsible for forwarding SIP signaling between an IMS user and the home network.

[0069] 3. S-CSCF Network Element: S-CSCF is the unified entry point of the IMS user home network, responsible for allocating or querying the S-CSCF that serves the user.

[0070] 4. IMS AGW: IMS AGW can provide the functions of an IMS network access gateway and a media gateway.

[0071] 5. DCSF: Used to provide data channel signaling control functions.

[0072] 6. DCAR: DCAR is a repository for storing data channel applications (DC Apps). After a developer completes the development of a DC App, it is uploaded to the operator's DCS, and then saved in DCAR by DCS. DCS downloads the DC App from DCAR to the local when needed for subsequent processing.

[0073] The DC control network element in the embodiments of this application includes DCSF, and optionally also includes DCAR.

[0074] 7. HSS: HSS serves as a database for storing user information in IMS, saving user data. By way of example and not limitation, user data may include information on the data channel services subscribed to by the user with the operator.

[0075] 8. IMS AS: IMS AS is the application layer device at the top layer in the IMS system, providing basic services and supplementary services, such as multimedia conferencing, converged communication, SMS gateway, standard operator console and other services. The IMS network is an open system based on IP bearer and providing various multimedia services to users. IMS AS interacts with CSCF to trigger and execute various network services. Moreover, there is a communication connection between the terminal device and IMS AS, and IMS AS can establish a data channel for the terminal device. Exemplarily, IMS AS can be a multimedia telephony application server (MMTEL AS) or a telephony application server (TAS).

[0076] 9. DC AS: DC AS is used to provide data channel service logic. It can be understood that DC AS can be deployed in the IMS network. In this case, DC AS can be directly connected to other network elements in the IMS network through an interface (such as a service-based interface); alternatively, DC AS can also be deployed outside the IMS network. In this case, DC AS can be connected to other network elements in the IMS network through a network exposure function (NEF) network element (or the IMS border network management).

[0077] 10. NEF network element: The NEF network element is used to securely expose various services of the 5GC network to third parties.

[0078] 11. MF network element: It can also be called a media server and is used to provide media resource management functions, including media resource management for media channels and media resource management for data channels.

[0079] 12. SAAF network element: It is responsible for the logical functions of the user AI Agent and is used to implement the call assistant function. The SAAF network element can be set in the MF or can also be deployed independently. This application does not specifically limit this here.

[0080] Mb, Iq, Gm, Mw, N70, N71 / Sh, N72 / Sc, DC1, DC2, DC3, DC4, DC5, ISC, MDC1, MDC2, MDC3, N33 are Figure 2A or Figure 2B network elements in [] provide service-based interfaces for calling corresponding service-based operations.

[0081] Among them, for the call assistant service to enhance the user interaction experience, the voice-driven digital human function is introduced. In the call assistant service scenario, the voice drives the digital human image, increasing the interaction experience between the user and the call assistant. Among them, the digital human is a virtual human image with a digitalized appearance, which is displayed relying on a display device. For example, the display effect can be presented through devices such as mobile phones, TVs, augmented reality (AR) technology, or virtual reality (VR) glasses. The digital human has an appearance similar to or lifelike to that of a human, with anthropomorphic characteristics such as an intuitive appearance, gender, and personality. Driven by digital technology, the digital human can have behaviors similar to humans, including expression capabilities such as language, expressions, and actions. Even through artificial intelligence, the digital human can have simple thoughts, recognize the external environment, and interact with people. Digital humans are widely used in scenarios such as film and television production, virtual anchors, virtual teaching, game entertainment, virtual customer service, virtual tour guides, and real-time communication.

[0082] The current voice only drives the mouth movements of the digital avatar and does not drive the expressions of the digital human (the expressions of the current digital human are independent of the output content of the call assistant). The expression of the digital human remains neutral throughout, with low interest and playability, and insufficient call experience. If it is desired that the digital human supports more diverse expressions, the call assistant needs to use a more powerful computing model to achieve this, resulting in high processing resource consumption.

[0083] Based on this, the present application provides a communication method to make the expression of the digital human more natural when using the call assistant service, improve the user experience, and not increase the complexity of model processing. This method is applicable to the communication architectures described above Figure 2A or Figure 2B and is implemented through the interaction between the first network element and the third network element. Among them, the first network element is used to provide media resource management functions, such as the MF described above. The first network element can be the first network element itself, or a component within the first network element (such as a processor, chip, or chip system, etc.), or it can also be a logical module or software that implements all or part of the first network element. The third network element is used to provide the time-frequency frames or picture frames of the digital human corresponding to the call assistant service, which can be understood as a video material warehouse. The third network element is usually deployed in an operator or a cloud server. The third network element can be the third network element itself, or a component within the third network element (such as a processor, chip, or chip system, etc.), or it can also be a logical module or software that implements all or part of the third network element. The present application does not specifically limit this here. In addition, when this communication method is applied to the communication architecture described in Figure 2B it also involves a second network element, which can be understood as a call assistant, such as the SAAF described above. This is only for illustrative purposes and not specifically limited. Additionally, this communication method also involves terminal devices, data channel application servers, etc. Referring to Figure 3 the following steps are executed, and the number of the first network element, the second network element, the third network element, the terminal device, and the data channel application server is not specifically limited here Figure 3 Taking one as an example for illustration:

[0084] Step 301, the first network element receives a first notification, where the first notification is used to indicate that the first call assistant service has been started, and the first notification includes: the identifier of the digital human corresponding to the first call assistant service.

[0085] Among them, the call assistant services subscribed by different users may be the same or different. Among them, a user can correspond to one or more terminal devices, and the user can use any one of the one or more terminal devices to make a call and access the IMS network to initiate the call assistant service. Exemplarily, the call assistant service subscribed by user A is only a voice service and does not include a digital human. The call assistant service subscribed by user B corresponds to digital human B. The call assistant service subscribed by user C corresponds to digital human C. The call assistant services subscribed by user D and user E correspond to the same digital human W. Among them, the image of digital human W can be a certain public figure, for example, actor W, artist W, or cartoon character W, etc. This is only an exemplary description here and is not specifically limited.

[0086] In addition, the subscribing user of the first call assistant service subscribes to the digital human corresponding to the call assistant service. The first call assistant service may be one call assistant service or multiple call assistant services. For example, the first call assistant service includes the call assistant service of user 1 and the call assistant service of user 2. This is only an exemplary description here and is not specifically limited.

[0087] In one possible implementation, the user makes a call through the terminal device and accesses the IMS network, thereby triggering the first notification.

[0088] In another possible implementation, after the data channel application server obtains the information to start the first call assistant service, it sends a first notification to the first network element. Exemplarily, the data channel application server is DC AS. DC AS stores the call assistant subscription information of the user, can clarify when the call assistant service needs to be started, and starts the call assistant service when the agreed time arrives, and sends a first notification to the first network element (such as MF).

[0089] Exemplarily, the first notification further includes: the start time of the first call assistant service, the duration of the first call assistant service, etc.

[0090] Step 302, the first network element requests the first information according to the first notification, and the first information is used to indicate the first expression description information of the digital human corresponding to the first call assistant service.

[0091] In one possible implementation, the first network element sends a second request message to the third network element according to the first notification. The second request message includes: the identifier of the digital human corresponding to the first call assistant service. The second request message is used to request the first expression description information; the third network element retrieves the first information according to the identifier of the digital human corresponding to the first call assistant service, and then the third network element sends the first information to the first network element. The first network element obtains the first information through the signaling interaction with the third network element, which does not increase the processing complexity of the first network element and can improve the data processing efficiency.

[0092] In another possible implementation, the first network element forwards a first notification to the third network element to obtain first information.

[0093] The first information may be first expression description information or index information of the first expression description information. It is not specifically limited herein.

[0094] The first expression description information may be an expression material produced offline. The expression material produced offline can be obtained based on an AI model, which is not specifically limited herein. The first expression description information may include an index of a certain emotion file of the digital human. Exemplarily, the first expression description information includes the start frame number, end frame number, or duration of a video of a certain emotion of the digital human *. For example, the start frame number and end frame number of the digital human X with a happiness level of 1, the start frame number and end frame number of the digital human X with a happiness level of 2, the start frame number and duration of the digital human X with a sadness level of 1, the start frame number and end frame number of the digital human X with a surprise level of 1, and the start frame number and end frame number of the digital human X with a neutral level of 1. Optionally, the first expression description information may further include the start frame number and end frame number of the switching between different emotions. For example, the first expression description information includes the switching start frame number and end frame number of the digital human X switching from a happiness level of 1 to a neutral level of 1. The first expression description information may further include a specific emotion file. Exemplarily, the first expression description information includes all video frames or picture frames of the digital human X with a happiness level of 1. Only exemplary descriptions of the first expression description information are provided herein, and it is not limited in specific applications. Additionally, the degree of emotion can be predefined. For example, three levels of happiness are set, namely slightly happy, moderately happy, and extremely happy, which are not specifically limited herein.

[0095] Step 303, the first network element obtains second information, where the second information includes: the prompt information replied by the call assistant and the emotion information matching the prompt information.

[0096] In one possible implementation, the first network element is the Figure 2A MF in, and the first network element is co-located with the second network element (such as SAAF). The first network element can directly implement the call assistant service, so the first network element can obtain the second information through internal data processing.

[0097] In another possible implementation, the first network element is the Figure 2B MF in, and the first network element is separated from the second network element (such as SAAF). After the second network element receives the multimedia information (audio, video, or text, etc.) of the user, it can send the second information to the first network element.

[0098] Among them, the call assistant reply prompt information in the second information can be audio or voice text. Exemplarily, the user dials a call, accesses the IMS network, and activates the call assistant service. The user says, "I completed an important task today." The prompt information replied by the call assistant can be the audio "Congratulations! I'm glad you've made such progress." Or, the prompt information is the text "Congratulations! I'm glad you've made such progress." Among them, the emotion information matching the prompt information is <happy>. This is only an exemplary illustration here and is not specifically limited.

[0099] Step 304, the first network element generates first video information according to the first information and the second information, and the first video information is used to indicate the video of the digital human corresponding to the first call assistant service.

[0100] In a possible implementation manner, the first network element processes the first information and the second information to obtain the first video information. Exemplarily, the first information is a specific emotion file. The first network element can retrieve the first information based on the emotion information in the second information, obtain the emotion file matching the emotion information, and then perform video editing to fuse the prompt information into the video to obtain the video of the digital human.

[0101] In another possible implementation manner, the first network element determines second expression description information according to the second information and the first information. The second expression description information is associated with the emotion information matching the prompt information, and the second expression description information belongs to the first expression description information; obtains third information according to the second expression description information, and the third information is used to indicate the video frame or picture frame of the digital human corresponding to the first call assistant service that matches the second expression description information; generates first video information according to the second information and the third information (for example, performs video and audio arrangement based on AI). After determining the expression description information matching the second information, requests the corresponding digital human video frame or picture frame to generate the video of the digital human, and the expression of the digital human in the generated video of the digital human is more natural.

[0102] In still another possible implementation manner, the first network element determines second expression description information according to the second information and the first information. The second expression description information is associated with the emotion information matching the prompt information, and the second expression description information belongs to the first expression description information; generates first video information according to the second expression description information and the second information. After determining the second expression description information matching the second information, when the second expression description information is a specific emotion file, the first network element can generate the first video information based on the second expression description information and the second information.

[0103] Among them, in a specific implementation, the first network element may screen out the second expression description information from the first information with reference to the second information. For example, the first network element searches for the second expression description information corresponding to the emotion information in the first expression description information based on the emotion information in the second information, or the first network element inputs the second information and the first information into an AI processing model, and obtains the second expression description information based on the AI processing model. Here, it is only an exemplary illustration, and the specific method of determining the second expression description information is not limited.

[0104] In another specific implementation, the first network element may generate the fourth information according to the second information, and the fourth information is used to indicate the label of the emotion information matching the prompt information; the first network element determines the second expression description information according to the fourth information and the first information. In this application, after the first network element generates the fourth information based on the second information, it searches for the first information based on the fourth information and then determines the second expression description information matching it, and then generates the video of the digital human. The expression of the digital human in the generated video of the digital human is more natural. Among them, the fourth information is associated with the following information: the duration of the prompt information and the frame rate of the video of the digital human corresponding to the first call assistant service (such as 25 FPS). When the fourth information is related to the duration of the prompt information and the frame rate of the video of the digital human, it can ensure that the video of the digital human generated by the first network element is adapted to the prompt information and meets the requirements of the adapted video configuration. For example, if the duration of the prompt information is 5 s and the frame rate of the video of the digital human is 25 FPS, then the length of the fourth information is 5 * 25. Exemplarily, the fourth information is a sequence used to indicate emotion information, such as h, h, h, h, h, h, n, n, n, n, n.

[0105] In still another specific implementation, the first network element also receives the fourth information from the second network element; the first network element determines the second expression description information according to the fourth information and the first information. After the first network element obtains the fourth information from the second network element, it searches for the first information based on the fourth information and then determines the second expression description information matching it, and then generates the video of the digital human. The expression of the digital human in the generated video of the digital human is more natural.

[0106] Such as Figure 4As shown, the sequence indicating the emotional information is h,h,h,h,h,h,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n,n. A second expression description file corresponding to the first information is selected based on the sequence. For example, the frame numbers of the video frames are [001, 101, 102, 103, 203, 204, 205, 206, …], where the duration of the prompt information is 40 / 25=1.6 seconds, the frame rate of the video is 25fps, and one video frame corresponds to one sequence element in the sequence.

[0107] Among them, the first network element can send a first request message to the third network element according to the second expression description information, and the first request message includes: the second expression description information, and the first request message is used to request the video frame or picture frame of the digital person corresponding to the first call assistant service matching the second expression description information; the third network element retrieves the third information according to the second expression description information, and then sends the third information to the first network element. The first network element obtains the third information through signaling interaction with the third network element, which does not increase the processing complexity of the first network element and can improve data processing efficiency.

[0108] Step 305: The first network element sends first video information.

[0109] Exemplarily, the first network element may send the first video information to the terminal device, or the first network element may send the first video information to other applications (APPs), which is not specifically limited here.

[0110] like Figure 5 As shown, the offline generated digital human expression material is stored in the third network element. After the call assistant service is started, the call assistant can feedback text information and emotional information corresponding to the text information (i.e., the second information) based on the large language model. For example, the second information is "<Happy> Today's weather is very good, you can go out for an outing. <Normal> It will rain tomorrow, <Sad> You can only stay at home." The first network element generates a sequence indicating emotional information based on the second information, for example, h,h,h,h,h,h,h,h,n,n,n,n,n,n,n,s,s,s,s. The first network element determines the time-frequency frame or picture frame corresponding to the sequence through expression arrangement, and then constructs the time-frequency of the digital human based on the lip-driven model and the call assistant feedback text information based on the large language model. Exemplarily, if the degree of happiness is different, h1 corresponds to slight happiness, h2 corresponds to moderate happiness, and h3 corresponds to extreme happiness, then the sequence indicating emotional information can also be h1, h1, h2, h2.h3, h3, etc., which is only exemplified here.

[0111] In this application, after determining that the first call assistant service has been started, a request is sent for the digital human's first expression description information corresponding to the first call assistant service. Then, a digital human video is generated based on the prompt information replied by the call assistant, the emotion information matching the prompt information, and the first expression description information. Since the digital human's video is generated based on the first expression description information and the emotion information matching the prompt information, the expression of the digital human in the video is more natural. In addition, the generation of the digital human's expression does not introduce a model with complex processing, only requires a small amount of computing cost, and can run with high concurrency. Based on this, when using the call assistant service, the voice-driven digital human's expression is more natural, which can improve the user's service experience.

[0112] To better illustrate the solution of this application, the following will be described with specific examples. The following takes the first network element as MF, the second network element as SAAF, the third network element as the material warehouse, the data channel application server as DC AS, and the terminal devices as UE1 and UE2 as examples for illustration. The following Figures 6 to 9 In this case, MF and SAAF are separate network elements. Figure 10 In this case, MF and SAAF are integrated network elements. Optionally, in specific applications, there may also be interactions between other network elements, which are not limited here.

[0113] Refer to Figure 6 , assuming that the digital humans corresponding to the call assistant services subscribed by UE1 and UE2 are different and are digital humans specially customized for UE1 and UE2 respectively. The following takes UE1 as an example for illustration, and UE2 can refer to it for understanding. Figure 6 In this case, MF generates the fourth information by itself. The following operations are performed:

[0114] Step 601, UE1 responds to the user's call, accesses the IMS network, triggers the call assistant service 1 subscribed by the user, and sends a first notification to MF.

[0115] Step 602, MF sends a second request message to the material warehouse. The second request message includes: the digital human 1 identifier corresponding to the call assistant service 1.

[0116] Step 603, the material warehouse feeds back the expression description file corresponding to digital human 1 to MF.

[0117] Among them, the expression description file corresponding to digital human 1 is an index of a certain emotion file of digital human 1, which can be understood by referring to the description in step 302 above and will not be elaborated here.

[0118] Step 604, after UE1 responds to the user's multimedia information, it sends the multimedia information to SAAF.

[0119] For example, the multimedia information is the audio information "Completed an important task today!".

[0120] Step 605: SAAF generates second information based on the LLM service according to the multimedia information.

[0121] Among them, the prompt information in the second message is the text "Congratulations! I'm glad that you can make such progress", as well as the emotional information <happy>.

[0122] Step 606: SAAF sends second information to MF.

[0123] Step 607: MF generates fourth information according to the second information.

[0124] Step 608: MF determines the second expression description information according to the fourth information and the first expression description information.

[0125] This can be understood by referring to the description in the above step 304 and will not be described in detail here.

[0126] Step 609: The MF sends a first request message to the material warehouse, where the first request message includes the second expression description information.

[0127] Step 610: The material warehouse sends third information to the MF.

[0128] Step 609 and step 610 may be understood by referring to the description in step 304 above, and will not be described in detail here.

[0129] Step 611: MF generates first video information according to the third information and the second information.

[0130] Step 612: MF sends the first video information to UE1.

[0131] In this example, expression materials are produced offline, and a variety of expression materials are stored in an expression warehouse. During a call, MF arranges expressions in real time to achieve the goal of lightweight expression drive, so that the expressions are associated with the reply content of the LLM.

[0132] Reference Figure 7 , assuming that the call assistant services signed by UE1 and UE2 correspond to different digital people, which are digital people customized for UE1 and UE2 respectively. The following is explained by taking UE1 as an example, and UE2 can refer to it for understanding. Figure 7 In the process, SAAF generates the fourth information. The execution is as follows:

[0133] Step 701, UE1 responds to a user making a call, accesses the IMS network, triggers the call assistant service 1 that the user has subscribed to, and sends a first notification to the MF.

[0134] In addition, the DC AS determines that the start time of the call assistant service 1 scheduled by UE1 has arrived, and can also send a first notification to the MF.

[0135] Step 702, the MF sends a second request message to the material warehouse. The second request message includes: the digital human 1 identifier corresponding to the call assistant service 1.

[0136] Step 703, the material warehouse feeds back to the MF the expression description file corresponding to the digital human 1.

[0137] Among them, the expression description file corresponding to the digital human 1 is an index of a certain emotion file of the digital human 1, which can be understood by referring to the description in step 302 above and will not be elaborated here.

[0138] Step 704, after the UE1 responds to the user's multimedia information, it sends the multimedia information to the SAAF.

[0139] For example, the multimedia information is the audio information "Completed an important task today!".

[0140] Step 705, the SAAF generates a second piece of information based on the multimedia information using the LLM service.

[0141] Among them, the prompt information in the second piece of information is the text "Congratulations! I'm glad you've made such progress", and the emotion information is <happy>.

[0142] Step 706, the SAAF generates a fourth piece of information based on the second piece of information.

[0143] Step 707, the SAAF sends the second piece of information and the fourth piece of information to the MF.

[0144] Exemplarily, the SAAF sends the prompt information in the second piece of information and the fourth piece of information to the MF.

[0145] Step 708, the MF determines the second expression description information based on the fourth piece of information and the first expression description information.

[0146] It can be understood by referring to the description in step 304 above and will not be elaborated here.

[0147] Step 709, the MF sends a first request message to the material warehouse. The first request message includes the second expression description information.

[0148] Step 710, the material warehouse sends a third piece of information to the MF.

[0149] Among them, steps 709 and 710 can be understood by referring to the description in step 304 above and will not be elaborated here.

[0150] Step 711, the MF generates a first video information based on the third piece of information and the second piece of information.

[0151] Step 712, the MF sends the first video information to the UE1.

[0152] In this example, expression materials are produced offline, and various expression materials are stored in the expression repository. During a call, the MF performs real-time expression choreography based on the fourth piece of information fed back by the SAAF, achieving the goal of lightweight expression driving and associating the expressions with the reply content of the LLM.

[0153] Refer to Figure 8 , assuming that the digital humans corresponding to the call assistant services subscribed by UE1 and UE2 are the same, for example, a certain artist. The following takes UE1 as an example for illustration, and UE2 can refer to it for understanding. Figure 8 In, the MF generates the fourth piece of information by itself. The following operations are performed:

[0154] Step 801, UE1 responds to the user's call, accesses the IMS network, triggers the call assistant service 1 subscribed by the user, and sends a first notification to the MF.

[0155] Step 802, the MF sends a second request message to the material repository. The second request message includes: the identifier of digital human 1 corresponding to the call assistant service 1.

[0156] Step 803, the material repository feeds back the expression description file corresponding to digital human 1 to the MF.

[0157] Among them, since the artist image data is small, all emotion files can be downloaded at one time when starting the digital human service. Therefore, the expression description file corresponding to this digital human 1 is all the emotion files of digital human 1, and can be understood by referring to the description in step 302 above and will not be elaborated here.

[0158] Step 804, after UE1 responds to the user's multimedia information, it sends the multimedia information to the SAAF.

[0159] For example, the multimedia information is the audio information "Completed an important task today!".

[0160] Step 805, the SAAF generates the second piece of information based on the multimedia information using the LLM service.

[0161] Among them, the prompt information in the second piece of information is the text "Congratulations! I'm glad you've made such progress", and the emotion information is <happy>.

[0162] Step 808, the SAAF sends the second piece of information to the MF.

[0163] Step 807, the MF generates the fourth piece of information based on the second piece of information.

[0164] Step 808, the MF determines the second expression description information based on the fourth piece of information and the first expression description information.

[0165] It can be understood by referring to the description in step 304 above and will not be elaborated here.

[0166] Step 809, the MF generates the first video information according to the second expression description information and the second information.

[0167] Step 810, the MF sends the first video information to UE1.

[0168] In this example, expression materials are produced offline, and various expression materials are stored in the expression repository. During a call, the MF performs real-time expression arrangement to achieve the goal of lightweight expression driving, and associates the expressions with the reply content of the LLM.

[0169] Refer to Figure 9 , assuming that the digital humans corresponding to the call assistant services subscribed by UE1 and UE2 are the same, for example, a certain artist. The following takes UE1 as an example to illustrate, and UE2 can refer to it for understanding. Figure 9 In, the SAAF generates the fourth information. The following is executed:

[0170] Step 901, UE1 responds to the user's call, accesses the IMS network, triggers the call assistant service 1 subscribed by the user, and sends the first notification to the MF.

[0171] Step 902, the MF sends a second request message to the material repository. The second request message includes: the digital human 1 identifier corresponding to the call assistant service 1.

[0172] Step 903, the material repository feeds back the expression description file corresponding to the digital human 1 to the MF.

[0173] Among them, the artist image data is less. When starting the digital human service, all emotion files can be downloaded at one time. Therefore, the expression description file corresponding to the digital human 1 is all the emotion files of the digital human 1, and it can be understood by referring to the description in step 302 above and will not be elaborated here.

[0174] Step 904, after UE1 responds to the user's multimedia information, it sends the multimedia information to the SAAF.

[0175] For example, the multimedia information is the audio information "Completed an important task today!".

[0176] Step 905, the SAAF generates the second information based on the multimedia information using the LLM service.

[0177] Among them, the prompt information in the second information is the text "Congratulations! I'm glad you've made such progress", and the emotion information is <happy>.

[0178] Step 906, the SAAF generates the fourth information according to the second information.

[0179] Step 907, the SAAF sends the second information and the fourth information to the MF.

[0180] Step 908, the MF determines the second expression description information according to the fourth information and the first expression description information.

[0181] It can be understood with reference to the description in step 304 above, and will not be elaborated here.

[0182] Step 909, the MF generates the first video information according to the second expression description information and the second information.

[0183] Step 910, the MF sends the first video information to the UE1.

[0184] In this example, the expression materials are produced offline, and various expression materials are stored in the expression warehouse. During the call, the MF performs real-time expression arrangement based on the fourth information fed back by the SAAF, achieving the goal of lightweight expression drive and associating the expressions with the reply content of the LLM.

[0185] Refer to Figure 10 , assuming that the digital humans corresponding to the call assistant services subscribed by the UE1 and UE2 are different and are digital humans specially customized for the UE1 and UE2 respectively. The following takes the UE1 as an example to illustrate, and the UE2 can refer to it for understanding. The MF and the SAAF are co-located and perform as follows:

[0186] Step 1001, the UE1 responds to the user's call, accesses the IMS network, triggers the call assistant service 1 subscribed by the user, and sends the first notification to the MF.

[0187] Step 1002, the MF sends a second request message to the material warehouse. The second request message includes: the digital human 1 identifier corresponding to the call assistant service 1.

[0188] Step 1003, the material warehouse feeds back the expression description file corresponding to the digital human 1 to the MF.

[0189] Among them, the expression description file corresponding to the digital human 1 is an index of a certain emotion file of the digital human 1, and can be understood with reference to the description in step 302 above. It will not be elaborated here.

[0190] Step 1004, after the UE1 responds to the user's multimedia information, it sends the multimedia information to the MF.

[0191] For example, the multimedia information is the audio information "Completed an important task today!".

[0192] Step 1005, the MF generates the second information based on the multimedia information using the LLM service.

[0193] Among them, the text with the prompt message "Congratulations! I'm glad you've made such progress" in the second piece of information, and the emotion information <happy>.

[0194] Step 1006, the MF generates a fourth piece of information based on the second piece of information.

[0195] Step 1007, the MF determines the second expression description information based on the fourth piece of information and the first expression description information.

[0196] It can be understood by referring to the description in step 304 above, and will not be elaborated here.

[0197] Step 1008, the MF sends a first request message to the material warehouse, and the first request message includes the second expression description information.

[0198] Step 1009, the material warehouse sends the third piece of information to the MF.

[0199] Among them, steps 609 and 610 can be understood by referring to the description in step 304 above and will not be elaborated here.

[0200] Step 1010, the MF generates the first video information based on the third piece of information and the second piece of information.

[0201] Step 1011, the MF sends the first video information to UE1.

[0202] In this example, expression materials are produced offline, and various expression materials are stored in the expression warehouse. During a call, the MF performs real-time expression arrangement to achieve the goal of lightweight expression drive, and associates the expressions with the reply content of the LLM.

[0203] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of device interaction. It can be understood that, in order to implement the above functions, each device may include corresponding hardware structures and / or software modules for performing each function. Those skilled in the art should easily realize that, combined with the units and algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0204] The embodiments of the present application can divide the functions of the device according to the above method examples. For example, each function unit can be divided corresponding to each function, or two or more functions can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software function unit.

[0205] In the case of adopting an integrated unit, Figure 11 A possible exemplary block diagram of the communication device involved in the embodiments of the present application is shown. As Figure 11 shown, the communication device 1100 may include: a processing unit 1101 and a transceiver unit 1102. The processing unit 1101 is used to control and manage the operations of the communication device 1100. The transceiver unit 1102 is used to support the communication between the communication device 1100 and other devices. Optionally, the transceiver unit 1102 may include a receiving unit and / or a transmitting unit, which are respectively used to perform receiving and transmitting operations. Optionally, the communication device 1100 may further include a storage unit, which is used to store the program code and / or data of the communication device 1100. The transceiver unit may be referred to as an input / output unit, a communication unit, etc., and the transceiver unit may be a transceiver; the processing unit may be a processor. When the communication device is a module (such as a chip) in a communication device, the transceiver unit may be an input / output interface, an input / output circuit, or an input / output pin, etc., and may also be referred to as an interface, a communication interface, or an interface circuit, etc.; the processing unit may be a processor, a processing circuit, or a logic circuit, etc. Exemplarily, the device may be the above-mentioned first network element or second network element.

[0206] For a more detailed description of the above-mentioned processing unit 1101 and transceiver unit 1102, reference can be directly made to the relevant descriptions in the above-mentioned various method embodiments, and details are not repeated here.

[0207] As Figure 12 shown, a communication device 1200 provided by the present application is also shown. The communication device 1200 may be a chip or a chip system. The communication device may be located in the device involved in any of the above-mentioned method embodiments, such as a terminal or a network device, etc., to perform the actions corresponding to the device.

[0208] Optionally, the chip system may be composed of chips, or may include chips and other discrete devices.

[0209] The communication device 1200 includes a processor 1210.

[0210] The processor 1210 is used to execute the computer program stored in the memory 1220 to implement the actions of each device in any of the above-mentioned method embodiments.

[0211] The communication device 1200 may further include a memory 1220, which is used to store the computer program.

[0212] Optionally, there is a coupling between the memory 1220 and the processor 1210. The coupling is an indirect coupling or communication connection between devices, units or modules, which can be electrical, mechanical or other forms for information interaction between devices, units or modules. Optionally, the memory 1220 and the processor 1210 are integrated together.

[0213] Wherein, both the processor 1210 and the memory 1220 can be one or more, without limitation.

[0214] Optionally, in practical applications, the communication device 1200 may or may not include a transceiver 1230. In the figure, it is indicated by a dashed box. The communication device 1200 can interact with other devices through the transceiver 1230. The transceiver 1230 can be a circuit, a bus, a transceiver or any other device that can be used for information interaction.

[0215] In a possible implementation manner, the communication device 1200 can be a terminal and a network device, etc. in the above method implementations.

[0216] In the embodiments of the present application, the specific connection medium between the transceiver 1230, the processor 1210 and the memory 1220 is not limited. In the embodiments of the present application Figure 12 it is shown that the memory 1220, the processor 1210 and the transceiver 1230 are connected through a bus. The bus is Figure 12 shown by a thick line. The connection manners between other components are only for illustrative purposes and are not to be construed as limiting. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 12 only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. In the embodiments of the present application, the processor can be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field-programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly implemented by a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0217] In the embodiments of the present application, the memory can be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), etc., or can also be a volatile memory, such as a random-access memory (RAM). The memory can also be any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing computer programs, program instructions, and / or data.

[0218] Based on the above embodiments, refer to Figure 13 , the embodiments of the present application further provide another communication device 1300, including: an interface circuit 1310 and a logic circuit 1320; the interface circuit 1310, which can be understood as an input / output interface, can be used to execute the transceiver steps of each device in any of the above method embodiments, and the logic circuit 1320 can be used to run code or instructions to execute the methods executed by each device in any of the above embodiments, which will not be elaborated herein.

[0219] Based on the above embodiments, the embodiments of the present application further provide a computer-readable storage medium, which stores instructions that, when executed, cause the methods executed by each device in any of the above method embodiments to be implemented. The computer-readable storage medium can include: various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory, a random-access memory, a magnetic disk, or an optical disc.

[0220] Based on the above embodiments, the embodiments of the present application provide a communication system, which includes the terminals, network devices, etc. mentioned in any of the above method embodiments and can be used to execute the methods executed by each device in any of the above method embodiments.

[0221] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0222] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a device for implementing the functions specified in a flow or multiple flows in the flowchart and / or a block or multiple blocks in the block diagram.

[0223] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in a flow or multiple flows in the flowchart and / or a block or multiple blocks in the block diagram.

[0224] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in a flow or multiple flows in the flowchart and / or a block or multiple blocks in the block diagram.

Claims

1. A communication method, characterized in that: Applied to the first network element, including: receiving a first notification, the first notification being used to indicate that a first call assistant service has been started, the first notification including: an identifier of a digital person corresponding to the first call assistant service; Requesting first information according to the first notification, the first information being used to indicate first expression description information of the digital person corresponding to the first call assistant service; Acquire second information, where the second information includes: prompt information replied by the call assistant and emotion information matching the prompt information; Generate first video information according to the first information and the second information, where the first video information is used to indicate a video of a digital person corresponding to the first call assistant service; The first video information is sent.

2. The method according to claim 1, characterized in that The generating first video information according to the first information and the second information includes: Determine second expression description information according to the second information and the first information, the second expression description information is associated with the emotion information matched with the prompt information, and the second expression description information belongs to the first expression description information; Acquire third information according to the second expression description information, where the third information is used to indicate a video frame or a picture frame of a digital person corresponding to the first call assistant service that matches the second expression description information; The first video information is generated according to the second information and the third information.

3. The method according to claim 1, characterized in that The generating first video information according to the first information and the second information includes: Determine second expression description information according to the second information and the first information, the second expression description information is associated with the emotion information matched with the prompt information, and the second expression description information belongs to the first expression description information; The first video information is generated according to the second expression description information and the third information.

4. The method according to claim 2 or 3, characterized in that: The determining the second expression description information according to the second information and the first information comprises: generating fourth information according to the second information, wherein the fourth information is used to indicate a label of the emotion information matching the prompt information; The second expression description information is determined according to the fourth information and the first information.

5. The method according to claim 2 or 3, characterized in that: The method further comprises: receiving fourth information from the second network element, where the fourth information is used to indicate a label of the emotion information matching the prompt information; The determining the second expression description information according to the second information and the first information includes: The second expression description information is determined according to the fourth information and the first information.

6. The method according to any one of claims 2 to 5, characterized in that: The acquiring the third information according to the second expression description information comprises: Sending a first request message to a third network element according to the second expression description information, the first request message including: the second expression description information, the first request message being used to request a video frame or a picture frame of a digital person corresponding to the first call assistant service matching the second expression description information; The third information is received from the third network element.

7. The method according to any one of claims 4 to 6, characterized in that: The fourth information is associated with the following information: The duration of the prompt information is associated with the frame rate of the digital human video corresponding to the first call assistant service.

8. The method according to any one of claims 4 to 7, characterized in that: The fourth information is a sequence used to indicate the emotion information.

9. The method according to any one of claims 1 to 8, characterized in that: The requesting first information according to the first notification includes: Sending a second request message to a third network element according to the first notification, the second request message including: an identifier of a digital person corresponding to the first call assistant service, the second request message being used to request the first expression description information; The first information is received from the third network element.

10. The method according to any one of claims 1 to 9, characterized in that: The obtaining of the second information includes: Second information is received from a second network element.

11. A communication method, characterized in that: Applied to the second network element, including: receiving multimedia information; Acquire second information according to the multimedia information, the second information comprising: prompt information replied by the call assistant and emotion information matching the prompt information; Send the second information to the first network element.

12. The method according to claim 11, characterized in that: The method further comprises: generating fourth information according to the second information, wherein the fourth information is used to indicate a label of the emotion information matching the prompt information; Send the fourth information to the first network element.

13. The method according to claim 11 or 12, characterized in that: The fourth information is associated with the following information: The duration of the prompt information is associated with the frame rate of the digital human video corresponding to the first call assistant service.

14. The method according to any one of claims 11 to 13, characterized in that: The fourth information is a sequence used to indicate the emotion information.

15. A communication method, characterized in that: Applied to the third network element, including: receiving a second request message from the first network element, the second request message including: an identifier of a digital person corresponding to the first call assistant service, the second request message being used to request first expression description information; Sending first information to the first network element, where the first information is used to indicate the first expression description information of the digital person corresponding to the first call assistant service.

16. The method according to claim 15, characterized in that The method further comprises: Receiving a first request message from a first network element, the first request message including: second expression description information, the first request message being used to request a video frame or a picture frame of a digital person corresponding to the first call assistant service matching the second expression description information; Send third information to the first network element, where the third information is used to indicate a video frame or a picture frame of a digital person corresponding to the first call assistant service that matches the second expression description information.

17. A communication device, characterized in that: include: A first network element for executing the method as claimed in any one of claims 1 to 10, or a second network element for executing the method as claimed in any one of claims 11 to 14, or a third network element for executing the method as claimed in any one of claims 15 to 16.

18. A communication device, characterized in that: The method comprises at least one processor; and a communication interface communicatively connected to the at least one processor; the at least one processor executes instructions stored in a memory so that the method according to any one of claims 1 to 16 is executed.

19. A computer-readable storage medium, characterized in that: A computer program or instruction is stored, and when the computer program or instruction is executed on a computer, the computer is caused to implement the method according to any one of claims 1 to 16.

20. A computer program product, characterized in that When a computer reads and executes the computer program product, the method according to any one of claims 1 to 16 is executed.

Citation Information

Patent Citations

  • Digital human interactive expression generation method and device, equipment and storage medium

    CN117409115A

  • Video call control method, communication equipment and storage medium

    CN117412254A

  • Control method of digital human and training method and system of large text model

    CN118227082A

  • 3D virtual character interaction system accessed to large visual screen

    CN118312073A

  • Chat interaction with multiple virtual assistants at the same time

    US20230013828A1

Cited By

  • Communication method and apparatus

    WO2026153011A1