A digital human interaction method and system in a native call process

Through the collaborative work of the core network and edge rendering nodes, the delay problem of digital human interaction in traditional communication networks is solved, low-latency real-time interaction is achieved, and the user experience and immersive interaction effect are improved.

CN120567835BActive Publication Date: 2025-10-10CHINA UNICOM ONLINE INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511073354.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-10-10
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

Existing digital human interaction technologies have latency and compatibility issues due to the limitations of browser and terminal hardware performance, and cannot achieve low-latency real-time interaction in traditional communication networks.

Method used

A digital human service request is generated through the core network. The digital human engine loads the pre-customized model and sends it to the edge rendering node. The edge rendering node generates synchronized video and voice streams, which are forwarded to the user terminal by the core network, realizing low-latency interaction based on traditional communication networks.

Benefits of technology

It achieves low-latency, real-time digital human interaction, improves user experience and immersive interaction effects, and is suitable for traditional communication network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567835B_ABST
    Figure CN120567835B_ABST
Patent Text Reader

Abstract

The application provides a digital human interaction method and system in a native call process, and relates to the technical field of digital humans. When a user terminal initiates a call event, a core network generates and outputs a digital human service request based on the call content corresponding to the call event. A digital human engine loads a digital human model in response to the digital human service request. An edge rendering node generates and outputs a digital human video stream and a voice stream to the core network based on the digital human model. The core network establishes a call with the user terminal while forwarding the digital human video stream and the voice stream to the user terminal to display a virtual digital human image capable of real-time interaction in the user terminal. The application realizes a low-latency digital human interaction method that can be implemented based on a traditional communication network. The core network intelligently schedules audio and video streams to ensure that the user terminal presents a high-fidelity, interactive digital human image in real time, thereby improving the immersive experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital human technology, and in particular to a digital human interaction method and system during a native call process. Background Art

[0002] With the development of artificial intelligence and computer graphics, digital humans and intelligent agent interaction technologies have been widely used in various fields, such as online customer service, virtual assistants, and identity authentication. The current technical architecture mainly relies on client-side models, including web pages (H5) and mobile apps.

[0003] In web-based interactive solutions, digital humans typically implement real-time communication based on WebRTC or WebSocket protocols, and rely on HTML5 Canvas or WebGL for 3D model rendering. This solution requires cloud-based APIs for computing support. Typical applications include virtual shopping guides in e-commerce or customer service websites. However, due to browser performance limitations, rendering and interacting with complex digital humans can face latency and compatibility issues.

[0004] In mobile interaction solutions, digital humans are typically integrated through SDKs like Unity or Unreal Engine. Model resources must be pre-downloaded and cached locally to improve rendering efficiency and interactive fluidity. This type of solution is common in digital human identity verification scenarios within financial apps, but due to limitations in terminal hardware performance, it can lead to excessive resource usage and compatibility issues.

[0005] In addition, the communication links of existing digital human interaction solutions are mainly based on Internet application layer protocols (such as SIP and RTMP), which are only applicable to IP network environments and cannot be directly connected to the circuit domain (CS) or packet domain (PS) of traditional communication networks, limiting their expanded application in telecommunications-level services (such as voice calls and SMS interactions).

[0006] Therefore, there is an urgent need for a low-latency digital human interaction method that can be implemented based on traditional communication networks. Summary of the Invention

[0007] The problem solved by the present invention is how to provide a low-latency digital human interaction method that can be implemented based on a traditional communication network.

[0008] To solve the above problems, the present invention provides a method and system for digital human interaction during a native call.

[0009] In a first aspect, the present invention provides a method for digital human interaction during a native call, comprising:

[0010] When a user terminal initiates a call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event; the digital human service request includes at least digital human service identification information, and the digital human service identification information is used to identify the digital human service type corresponding to the call content;

[0011] The digital human engine responds to the digital human service request, loads a pre-customized digital human model according to the digital human service identification information, and sends the loaded digital human model to the edge rendering node;

[0012] The edge rendering node is used to generate and output a digital human video stream and a voice stream based on the digital human model to the core network, wherein the digital human video stream and the voice stream are synchronized in time;

[0013] The core network establishes a call with the user terminal in response to the received digital human video stream and voice stream, and simultaneously forwards the digital human video stream and voice stream to the user terminal, so as to display a virtual digital human image capable of real-time interaction in the user terminal.

[0014] Optionally, when the user terminal initiates a call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event, including:

[0015] When the user terminal initiates a call event, a SIP signaling channel is established between the user terminal and the core network, and the core network receives the call content corresponding to the call event initiated by the user terminal through the SIP signaling channel;

[0016] When the core network identifies the need for digital human service based on the call content corresponding to the call event, it determines the digital human engine corresponding to the call content, establishes a WebSocket signaling channel, and the digital human engine receives the digital human service request through the WebSocket signaling channel.

[0017] Optionally, the call content includes at least a called number; when the core network identifies that a digital human service is needed based on the call content corresponding to the call event, determining a digital human engine corresponding to the call content includes:

[0018] The core network determines the preset digital human service type corresponding to the called number based on the called number corresponding to the call event, and determines the corresponding digital human engine according to the preset digital human service type corresponding to the called number; wherein the called number and the preset digital human service type have a corresponding mapping relationship.

[0019] Optionally, the preset digital human service types include at least: bank customer service digital human, e-commerce shopping guide digital human and personal telephone assistant digital human.

[0020] Optionally, the edge rendering node is configured to generate and output a digital human video stream and a voice stream to a core network based on the digital human model, including:

[0021] The edge rendering node encapsulates the digital human video stream based on the RTP protocol, encapsulates the voice stream based on the RTCP protocol, and transmits the encapsulated digital human video stream and the voice stream to the core network through a media channel.

[0022] Optionally, the edge rendering node is deployed in a predetermined area of ​​the core network.

[0023] Optionally, after forwarding the digital human video stream and voice stream to the user terminal, the method further includes:

[0024] During a call, the core network receives voice commands sent by the user terminal, converts the voice commands into voice data, and outputs the data to the digital human engine;

[0025] The digital human engine converts the voice data into text data, performs intent analysis based on the text data, determines a response action instruction corresponding to the voice data, and outputs the response action instruction to the edge rendering node;

[0026] The edge rendering node generates and outputs the corresponding digital human video stream and voice stream to the core network based on the response action instruction;

[0027] The core network forwards the digital human video stream and voice stream corresponding to the response action instruction to the user terminal to update the virtual digital human image displayed in the user terminal.

[0028] In a second aspect, the present invention provides a digital human interaction system during a native call, comprising:

[0029] The core network is configured to generate and output a digital human service request based on the call content corresponding to the call event when the user terminal initiates a call event;

[0030] A digital human engine, in response to the digital human service request, loads and outputs a pre-customized digital human model;

[0031] An edge rendering node, configured to generate and output a digital human video stream and a voice stream based on the digital human model, wherein the digital human video stream and the voice stream are synchronized in time;

[0032] The core network, in response to the received digital human video stream and voice stream, establishes a call with the user terminal and simultaneously forwards the digital human video stream and voice stream to the user terminal so as to display a virtual digital human image capable of real-time interaction in the user terminal.

[0033] Optionally, when the user terminal initiates a call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event, including:

[0034] When the user terminal initiates a call event, a SIP signaling channel is established between the user terminal and the core network, and the core network receives the call content corresponding to the call event initiated by the user terminal through the SIP signaling channel;

[0035] When the core network identifies the need for digital human service based on the call content corresponding to the call event, it determines the digital human engine corresponding to the call content, establishes a WebSocket signaling channel, and the digital human engine receives the digital human service request through the WebSocket signaling channel.

[0036] Optionally, the call content includes at least a called number; and when the core network identifies that a digital human service is needed based on the call content corresponding to the call event, determining the digital human engine corresponding to the call content includes:

[0037] The core network determines the preset digital human service type corresponding to the called number based on the called number corresponding to the call event, and determines the corresponding digital human engine according to the preset digital human service type corresponding to the called number; wherein the called number and the preset digital human service type have a corresponding mapping relationship.

[0038] The digital human interaction method and system in the native call process of the application have the following beneficial effects: when a user terminal initiates a call event, a core network generates and outputs a digital human service request to a digital human engine based on the call content corresponding to the call event; the native call initiated by the user terminal can directly trigger the digital human service request to the digital human engine through the core network, without the aid of an Internet network, and the triggering of the digital human service request can be directly realized through a traditional communication network. The digital human engine loads a pre-customized digital human model in response to the digital human service request and sends the loaded digital human model to an edge rendering node. The digital human engine can load a corresponding digital human model according to the digital human service request to provide a model basis for subsequent digital human interaction. The edge rendering node is used to generate and output a digital human video stream and a voice stream to the core network based on the digital human model. The digital human video stream and the voice stream are synchronized in time, and the digital human video stream and the voice stream are dynamically loaded and optimized by the edge rendering node, realizing millisecond-level digital human starting and interaction and significantly improving user experience. The core network establishes a call with the user terminal in response to the received digital human video stream and voice stream, and simultaneously forwards the digital human video stream and voice stream to the user terminal to display a virtual digital human image capable of real-time interaction in the user terminal, realizing a low-delay digital human interaction method capable of being realized based on a traditional communication network. Through intelligent scheduling of the core network, the audio and video streams are ensured to be presented in real time on the user terminal in a high-fidelity and interactive manner, and immersive experience is improved. BRIEF DESCRIPTION OF DRAWINGS

[0039] Figure 1 A digital human interaction method flowchart in a native call process of an embodiment of the application;

[0040] Figure 2 A core network generates a digital human service request flowchart of an embodiment;

[0041] Figure 3 A digital human interaction method flowchart in a native call process of another embodiment of the application;

[0042] Figure 4 A digital human interaction flowchart among a user terminal, a core network, a digital human engine and an edge rendering node;

[0043] Figure 5 A structure schematic diagram of a digital human interaction system in a native call process of an embodiment of the application;

[0044] Figure 6 A structure schematic diagram of an electronic device of an embodiment of the application. DETAILED DESCRIPTION

[0045] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.

[0046] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.

[0047] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to"; the term "based on" means "based at least in part on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments"; the term "optionally" means "optional embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0048] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0049] The names of the messages or information exchanged between multiple devices in the embodiments of the present invention are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0050] like Figure 1 As shown, an embodiment of the present invention provides a digital human interaction method during a native call, including the following steps:

[0051] S100: When a user terminal initiates a call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event.

[0052] Specifically, the user terminal can be a terminal device such as a mobile phone, a computer, or an AR device. The user terminal initiates a call event, which can be a user dialing a corporate hotline number (such as a bank hotline) through a mobile phone, or a user dialing another person's mobile phone number through a mobile phone.

[0053] In some embodiments, the call content corresponding to the call event may be the called number, that is, the number dialed by the user through the user terminal.

[0054] Specifically, the core network is the hub connecting user terminals and the digital human engine, and is responsible for call routing, signaling processing, and resource scheduling.

[0055] Specifically, when a user terminal initiates a call, if the called number is a registered Digital Human service number, the core network determines the corresponding Digital Human service identification information based on the called number and generates a corresponding Digital Human service request. The Digital Human service request includes the Digital Human service identification information and the user terminal's capability parameters (such as supported video resolution and encoding format). The Digital Human service request is then sent to the Digital Human engine. In one embodiment, the Digital Human service identification information can be a Digital Human service ID, which is used to identify the Digital Human service type corresponding to the call content.

[0056] S200: The digital human engine responds to the digital human service request, loads a pre-customized digital human model according to the digital human service identification information, and sends the loaded digital human model to the edge rendering node.

[0057] Specifically, the digital human engine is the "brain" of the digital human, responsible for parsing the digital human service request, determining the corresponding digital human service type based on the digital human service identification information in the digital human service request, and loading the pre-customized digital human model. For example, when the digital human service type is the bank customer service type, the corresponding digital human model is the bank customer service image.

[0058] S300: The edge rendering node is used to generate and output a digital human video stream and a voice stream based on the digital human model to the core network; wherein the digital human video stream and the voice stream are synchronized in time.

[0059] Specifically, the edge rendering node is deployed in a predetermined area of ​​the core network, and is close to the core network.

[0060] Specifically, the edge rendering node is the real-time audio and video generation hub of the digital human, based on the digital human model, using a graphics processing unit (GPU) to accelerate the rendering of high-fidelity digital human video streams, synchronously synthesizing spatial voice streams, and transmitting them to the core network through low-latency encryption protocols. Among them, the intelligent resource scheduling and quality of service guarantee mechanism of the edge rendering node ensures that 60FPS video and 48kHz audio can still be stably output in complex network environments, realizing an immersive interactive experience with an end-to-end delay of <100ms.

[0061] Specifically, the digital human video stream refers to the facial animation, body movements, etc. of the digital human, and the voice stream refers to the synthesized voice or real voice of the digital human. Both need to be synchronized in time, and any lack of synchronization will cause problems such as the digital human's appearance not matching, delayed movements, etc. For the synchronization of digital human video streams and voice streams in time, the digital human video stream and voice stream can be aligned in time stamp, i.e., each video frame in the video stream and each language frame in the language stream are strictly matched on the time axis.

[0062] S400: The core network responds to the received digital human video stream and voice stream, establishes a call with the user terminal, and simultaneously forwards the digital human video stream and voice stream to the user terminal to display a virtual digital human image capable of real-time interaction in the user terminal.

[0063] Specifically, after receiving the video stream (containing real-time rendering pictures of facial expressions, body movements, etc. of the virtual image) and voice stream (synthesized or real voice data) sent by the edge rendering node, the core network first establishes a two-way communication channel with the user terminal through a signaling protocol, and simultaneously starts a quality of service guarantee mechanism to mark the priority of the audio and video stream and shape the traffic. Subsequently, the core network synchronously encapsulates the video stream and voice stream based on timestamp alignment technology (such as the RTP / RTCP protocol), and forwards the encapsulated video stream and voice stream to the user terminal through a low-latency transmission path. The user terminal analyzes the video stream and voice stream in real time through a decoder, and finally presents a virtual digital human image with precise mouth shape matching (based on phoneme-visual position mapping) and natural and coherent movements on the screen of the user terminal, and ensures that the interaction delay between the user's voice input and the digital human's feedback is controlled within 200ms, realizing a natural and smooth real-time conversation experience.

[0064] In this embodiment, when a user terminal initiates a call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event. The digital human service request can be directly triggered to the digital human engine through the core network through the native call initiated by the user terminal. Without the help of the Internet network, the triggering of the digital human service request can be achieved directly through the traditional communication network. In response to the digital human service request, the digital human engine loads the pre-customized digital human model and sends the loaded digital human model to the edge rendering node. The digital human engine can load the corresponding digital human model according to the digital human service request to provide a model basis for subsequent digital human interaction. The edge rendering node is used to generate and output the digital human video stream and voice stream to the core network based on the digital human model. The digital human video stream and voice stream are synchronized in time. Through dynamic loading and rendering optimization of the edge rendering node, millisecond-level digital human startup and interaction are achieved, significantly improving the user experience. In response to the received digital human video stream and voice stream, the core network establishes a call with the user terminal and forwards the digital human video stream and voice stream to the user terminal to display a virtual digital human image capable of real-time interaction in the user terminal, thereby realizing a low-latency digital human interaction method that can be implemented based on traditional communication networks. Through the core network's intelligent scheduling of audio and video streams, it ensures that the user terminal presents a high-fidelity, interactive digital human image in real time, enhancing the immersive experience.

[0065] Alternatively, as Figure 2 As shown, when a user terminal initiates a call event, the core network generates and outputs a Digital Human service request to the Digital Human engine based on the call content corresponding to the call event, including the following steps:

[0066] S210: When the user terminal initiates a call event, a SIP signaling channel is established between the user terminal and the core network, and the core network receives the call content corresponding to the call event initiated by the user terminal through the SIP signaling channel.

[0067] Specifically, the SIP (Session Initiation Protocol) signaling channel is a communication path used to establish, modify, and terminate multimedia sessions (such as voice, video, instant messaging, etc.).

[0068] S220: When the core network recognizes that a digital human service is needed based on the call content corresponding to the call event, it determines the corresponding digital human engine and establishes a WebSocket signaling channel. The digital human engine receives the digital human service request through the WebSocket signaling channel.

[0069] Among them, the call content corresponding to the call event is transmitted in the form of ISUP signaling or BICC signaling, and the digital human service request is transmitted in the form of JSON-RPC instructions.

[0070] Specifically, the WebSocket signaling channel is a bidirectional communication channel based on the WebSocket protocol, which is used to efficiently transmit signaling messages in the Web and real-time communications. It makes up for the delay problem of traditional HTTP polling or long polling and provides low-latency, full-duplex communication capabilities.

[0071] Specifically, the core network can call the digital human engine corresponding to the call content through the API (application programming interface).

[0072] In this optional embodiment, a SIP signaling channel is used between the user terminal and the core network to achieve seamless connection with the traditional communication network, and key signaling such as the call content corresponding to the call event is reliably transmitted, so that digital human interaction can be realized in the native call; in addition, a WebSocket signaling channel is used between the core network and the digital human engine to provide the digital human engine with low-latency, two-way real-time interaction capabilities; and the SIP signaling channel processes standardized signaling, and the WebSocke signaling channel efficiently transmits high-frequency service requests, which can improve the overall performance.

[0073] Optionally, the call content corresponding to the call event initiated by the user terminal includes at least the called number. When the core network identifies that a digital human service is needed based on the call content corresponding to the call event, determining the digital human engine corresponding to the call content includes:

[0074] The core network determines the preset digital human service type corresponding to the called number based on the called number corresponding to the call event, and determines the corresponding digital human engine according to the preset digital human service type corresponding to the called number; wherein the called number and the preset digital human service type have a corresponding mapping relationship.

[0075] In some embodiments, the preset digital human service types include at least: bank customer service digital humans, e-commerce shopping guide digital humans, and personal telephone assistant digital humans.

[0076] In this optional embodiment, the preset digital human service type is directly associated with the called number to achieve millisecond-level service identification and digital human engine matching, thereby improving service response efficiency.

[0077] Optionally, the edge rendering node is used to generate and output a digital human video stream and voice stream based on the digital human model to the core network, including:

[0078] The edge rendering node encapsulates the digital human video stream based on the RTP protocol and the voice stream based on the RTCP protocol, and transmits the encapsulated digital human video stream and voice stream to the core network through the media channel.

[0079] Specifically, the encapsulated digital human video stream and voice stream are transmitted to the core network through independent media channels, separated from the above-mentioned signaling channels (SIP signaling channel and WebSocket signaling channel). The signaling channel is used to transmit control instructions (such as call establishment, service type determination, etc.), and the media channel is used to transmit video stream and voice stream.

[0080] In this optional embodiment, the encapsulated digital human video stream and voice stream are transmitted through the media channel and physically or logically isolated from the signaling channel to ensure the real-time interaction of the digital human.

[0081] Optionally, after forwarding the digital human video stream and voice stream to the user terminal, Figure 3 As shown, the following steps are also included:

[0082] S310: During a call, the core network receives voice commands sent by the user terminal, converts the voice commands into voice data, and outputs the data to the digital human engine.

[0083] Specifically, after receiving the voice command sent by the user terminal (usually a voice stream transmitted through the RTP media channel), the core network will first pre-process the voice data (such as noise reduction and format conversion), and then forward the voice data to the digital human engine through an internal interface (such as an API based on HTTP / REST or WebSocket).

[0084] S320: The digital human engine converts the voice data into text data, performs intent analysis based on the text data, determines the response action instructions corresponding to the voice data, and outputs the response action instructions to the edge rendering node; the text data is a structured string that can be recognized by a computer.

[0085] Specifically, the digital human engine converts voice data into text data through ASR (automatic speech recognition) technology, and then analyzes user intentions through natural language processing to generate corresponding response action instructions.

[0086] S330: The edge rendering node generates and outputs the corresponding digital human video stream and voice stream to the core network based on the response action instruction.

[0087] Specifically, after receiving the response action instructions (such as expressions, lip shapes, and body movements) issued by the digital human engine, the edge rendering node generates synchronized digital human video streams (based on RTP encapsulation) and voice streams (based on RTCP encapsulation) in real time, and transmits them to the core network through the media channel.

[0088] S340: The core network forwards the digital human video stream and voice stream corresponding to the response action instruction to the user terminal to update the virtual digital human image displayed in the user terminal.

[0089] Specifically, the core network forwards the digital human video and voice streams corresponding to the action instructions to the user terminal, achieving low-latency, high-fidelity interaction of the digital human. At the same time, it maintains real-time synchronization of control instructions through the signaling channel to ensure an end-to-end coherent experience.

[0090] In this optional embodiment, low-latency, high-fidelity interaction of digital humans is achieved during calls, while real-time synchronization of control instructions is maintained through signaling channels to ensure an end-to-end consistent experience.

[0091] Based on the relevant description of the above embodiment, Figure 4 As shown, the digital human interaction between the user terminal, core network, digital human engine and edge rendering node includes the following steps:

[0092] S401: The user terminal initiates a call event (eg, dials an enterprise's hotline number), and the user terminal sends the call content corresponding to the call event to the core network.

[0093] S402: The core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event.

[0094] S403: The digital human engine responds to the digital human service request, loads a pre-customized digital human model according to the digital human service identification information, and sends the loaded digital human model to the edge rendering node.

[0095] S404: The edge rendering node is used to generate and output the digital human video stream and voice stream based on the digital human model to the core network.

[0096] S405: The core network establishes a call with the user terminal in response to the received digital human video stream and voice stream, and forwards the digital human video stream and voice stream to the user terminal to display a virtual digital human image capable of real-time interaction in the user terminal.

[0097] S406: During the call, the user terminal receives a voice command input by the user.

[0098] S407: The user terminal sends the voice command to the core network, and the core network converts the voice command into voice data and outputs it to the digital human engine.

[0099] S408: The digital human engine converts the voice data into text data, performs intent analysis based on the text data, determines the response action instructions corresponding to the voice data, and outputs the response action instructions to the edge rendering node.

[0100] S409: The edge rendering node generates and outputs the corresponding digital human video stream and voice stream to the core network based on the response action instruction.

[0101] S410: The core network forwards the digital human video stream and the voice stream corresponding to the response action instruction to the user terminal to update the virtual digital human image displayed in the user terminal.

[0102] As shown in the figure, the embodiment of the application provides a digital human interaction system 500 in a native call process, which comprises: Figure 5

[0103] The core network 510 is configured to generate and output a digital human service request based on the call content corresponding to the call event when the user terminal initiates the call event.

[0104] The digital human engine 520 is configured to load and output a pre-customized digital human model in response to the digital human service request.

[0105] The edge rendering node 530 is configured to generate and output a digital human video stream and a voice stream based on the digital human model, wherein the digital human video stream and the voice stream are synchronized in time.

[0106] The core network 510 establishes a call with the user terminal in response to the received digital human video stream and voice stream, and forwards the digital human video stream and voice stream to the user terminal to display a virtual digital human image capable of real-time interaction in the user terminal.

[0107] Optionally, when the user terminal initiates the call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event, which comprises:

[0108] When the user terminal initiates the call event, a SIP signaling channel is established between the user terminal and the core network, and the core network receives the call content corresponding to the call event initiated by the user terminal through the SIP signaling channel.

[0109] When the core network identifies that the digital human service is needed based on the call content corresponding to the call event, the digital human engine corresponding to the call content is determined, a WebSocket signaling channel is established, and the digital human engine receives the digital human service request through the WebSocket signaling channel.

[0110] Optionally, the call content at least comprises a called number; and when the core network identifies that the digital human service is needed based on the call content corresponding to the call event, the digital human engine corresponding to the call content is determined, which comprises:

[0111] ​The core network determines the preset digital human service type corresponding to the called number based on the called number corresponding to the call event, and determines the corresponding digital human engine according to the preset digital human service type corresponding to the called number; wherein the called number and the preset digital human service type have a corresponding mapping relationship.

[0112] Optionally, the preset digital human service types include at least: bank customer service digital human, e-commerce shopping guide digital human and personal telephone assistant digital human.

[0113] Optionally, the edge rendering node is configured to generate and output a digital human video stream and a voice stream to a core network based on the digital human model, including:

[0114] The edge rendering node encapsulates the digital human video stream based on the RTP protocol, encapsulates the voice stream based on the RTCP protocol, and transmits the encapsulated digital human video stream and the voice stream to the core network through a media channel.

[0115] Optionally, the edge rendering node is deployed in a predetermined area of ​​the core network.

[0116] Optionally, after forwarding the digital human video stream and voice stream to the user terminal, the method further includes:

[0117] During a call, the core network receives voice commands sent by the user terminal, converts the voice commands into voice data, and outputs the data to the digital human engine;

[0118] The digital human engine converts the voice data into text data, performs intent analysis based on the text data, determines a response action instruction corresponding to the voice data, and outputs the response action instruction to the edge rendering node;

[0119] The edge rendering node generates and outputs the corresponding digital human video stream and voice stream to the core network based on the response action instruction;

[0120] The core network forwards the digital human video stream and voice stream corresponding to the response action instruction to the user terminal to update the virtual digital human image displayed in the user terminal.

[0121] like Figure 6 As shown, an electronic device 600 provided by an embodiment of the present invention includes a memory 610 and a processor 620; the memory 610 is used to store a computer program; the processor 620 is used to implement the digital human interaction method in the native call process as described above when executing the computer program.

[0122] In other words, an electronic device 600 includes a memory 610 and a processor 620 coupled to the memory 610; the memory 610 is configured to store a computer program; and the processor 620 is configured to perform the following operations when executing the computer program:

[0123] When a user terminal initiates a call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event; the digital human service request includes at least digital human service identification information, and the digital human service identification information is used to identify the digital human service type corresponding to the call content;

[0124] The digital human engine responds to the digital human service request, loads a pre-customized digital human model according to the digital human service identification information, and sends the loaded digital human model to the edge rendering node;

[0125] The edge rendering node is used to generate and output a digital human video stream and a voice stream based on the digital human model to the core network, wherein the digital human video stream and the voice stream are synchronized in time;

[0126] The core network establishes a call with the user terminal in response to the received digital human video stream and voice stream, and forwards the digital human video stream and voice stream to the user terminal to display a virtual digital human image capable of real-time interaction in the user terminal.

[0127] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the digital human interaction method in the native call process as described above is implemented.

[0128] In other words, a non-volatile computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the following operations:

[0129] When a user terminal initiates a call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event; the digital human service request includes at least digital human service identification information, and the digital human service identification information is used to identify the digital human service type corresponding to the call content;

[0130] The digital human engine responds to the digital human service request, loads a pre-customized digital human model according to the digital human service identification information, and sends the loaded digital human model to the edge rendering node;

[0131] The edge rendering node is used to generate and output a digital human video stream and a voice stream based on the digital human model to the core network, wherein the digital human video stream and the voice stream are synchronized in time;

[0132] The core network establishes a call with the user terminal in response to the received digital human video stream and voice stream, and forwards the digital human video stream and voice stream to the user terminal to display a virtual digital human image capable of real-time interaction in the user terminal.

[0133] An electronic device 600 that can serve as a server or client of the present invention will now be described, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device 600 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device 600 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0134] Electronic device 600 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). The RAM can also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. An input / output (I / O) interface is also connected to the bus.

[0135] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing relevant hardware, and the program can be stored in a computer readable storage medium. When the program is executed, the program can include the processes of the above-mentioned embodiment methods. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like. In this application, the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment of the present application. In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically independently, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0136] Although the present application is disclosed as above, the protection scope of the present application is not limited to this. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present application, and these changes and modifications will fall within the protection scope of the present application.

Claims

1. A digital human interaction method during a native call, characterized in that: include: When a user terminal initiates a call event, the core network generates and outputs a Digital Human service request to the Digital Human engine based on the call content corresponding to the call event; The digital human service request at least includes digital human service identification information, and the digital human service identification information is used to identify the digital human service type corresponding to the call content; The digital human engine responds to the digital human service request, loads a pre-customized digital human model according to the digital human service identification information, and sends the loaded digital human model to the edge rendering node; The edge rendering node is used to generate and output a digital human video stream and a voice stream based on the digital human model to the core network, wherein the digital human video stream and the voice stream are synchronized in time; The core network establishes a call with the user terminal in response to the received digital human video stream and voice stream, and simultaneously forwards the digital human video stream and voice stream to the user terminal, so as to display a virtual digital human image capable of real-time interaction in the user terminal.

2. The digital human interaction method during a native call according to claim 1, characterized in that: When a user terminal initiates a call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event, including: When the user terminal initiates a call event, a SIP signaling channel is established between the user terminal and the core network, and the core network receives the call content corresponding to the call event initiated by the user terminal through the SIP signaling channel; When the core network identifies the need for digital human service based on the call content corresponding to the call event, it determines the digital human engine corresponding to the call content, establishes a WebSocket signaling channel, and the digital human engine receives the digital human service request through the WebSocket signaling channel.

3. The digital human interaction method during a native call according to claim 2, characterized in that: The call content includes at least a called number; and when the core network identifies that a digital human service is needed based on the call content corresponding to the call event, determining a digital human engine corresponding to the call content includes: The core network determines the preset digital human service type corresponding to the called number based on the called number corresponding to the call event, and determines the corresponding digital human engine according to the preset digital human service type corresponding to the called number; wherein the called number and the preset digital human service type have a corresponding mapping relationship.

4. The digital human interaction method during a native call according to claim 3, characterized in that: The preset digital human service types include at least: bank customer service digital humans, e-commerce shopping guide digital humans, and personal telephone assistant digital humans.

5. The digital human interaction method during a native call according to claim 1, characterized in that: The edge rendering node is used to generate and output a digital human video stream and a voice stream to the core network based on the digital human model, including: The edge rendering node encapsulates the digital human video stream based on the RTP protocol, encapsulates the voice stream based on the RTCP protocol, and transmits the encapsulated digital human video stream and the voice stream to the core network through a media channel.

6. The digital human interaction method during a native call according to any one of claims 1 to 5, characterized in that: The edge rendering node is deployed in a predetermined area of ​​the core network.

7. The digital human interaction method during a native call according to any one of claims 1 to 5, characterized in that: After forwarding the digital human video stream and voice stream to the user terminal, the method further includes: During a call, the core network receives voice commands sent by the user terminal, converts the voice commands into voice data, and outputs the data to the digital human engine; The digital human engine converts the voice data into text data, performs intent analysis based on the text data, determines a response action instruction corresponding to the voice data, and outputs the response action instruction to the edge rendering node; The edge rendering node generates and outputs the corresponding digital human video stream and voice stream to the core network based on the response action instruction; The core network forwards the digital human video stream and voice stream corresponding to the response action instruction to the user terminal to update the virtual digital human image displayed in the user terminal.

8. A digital human interaction system during a native call, characterized in that: include: The core network is configured to generate and output a digital human service request based on the call content corresponding to the call event when the user terminal initiates a call event; A digital human engine, in response to the digital human service request, loads and outputs a pre-customized digital human model; An edge rendering node, configured to generate and output a digital human video stream and a voice stream based on the digital human model, wherein the digital human video stream and the voice stream are synchronized in time; The core network, in response to the received digital human video stream and voice stream, establishes a call with the user terminal and simultaneously forwards the digital human video stream and voice stream to the user terminal so as to display a virtual digital human image capable of real-time interaction in the user terminal.

9. The digital human interaction system during a native call according to claim 8, characterized in that: When a user terminal initiates a call event, the core network generates and outputs a digital human service request to the digital human engine based on the call content corresponding to the call event, including: When the user terminal initiates a call event, a SIP signaling channel is established between the user terminal and the core network, and the core network receives the call content corresponding to the call event initiated by the user terminal through the SIP signaling channel; When the core network identifies the need for digital human service based on the call content corresponding to the call event, it determines the digital human engine corresponding to the call content, establishes a WebSocket signaling channel, and the digital human engine receives the digital human service request through the WebSocket signaling channel.

10. The digital human interaction system during a native call according to claim 9, characterized in that: The call content includes at least a called number; and when the core network identifies that a digital human service is needed based on the call content corresponding to the call event, determining a digital human engine corresponding to the call content includes: The core network determines the preset digital human service type corresponding to the called number based on the called number corresponding to the call event, and determines the corresponding digital human engine according to the preset digital human service type corresponding to the called number; wherein the called number and the preset digital human service type have a corresponding mapping relationship.

Citation Information

Patent Citations

  • Digital human driving method, device, system, equipment, program product and medium

    CN119131206A

  • Virtual human interaction method and device, related equipment and computer program product

    CN119883006A

Cited By

  • Personalized native call digital human interaction method based on AI dynamic modulation

    CN121585763A