Digital human communication method, communication system, and communication device
By negotiating and exchanging information between terminal devices and network devices, the digital human driving mode is determined and corresponding resources are configured, which solves the problem of digital human communication failure caused by terminal device hardware limitations and realizes the successful execution of digital human services and resource matching.
Patent Information
- Application Number
- PCT/CN2025/075869
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2025-02-06
- Publication Date
- 2025-12-04
AI Technical Summary
In existing technologies, due to limitations in the hardware resources of terminal devices, the failure rate of digital human-driven services is high, making it impossible to effectively realize digital human communication on terminal devices.
Through information exchange between the first device and the second device, the driving mode of the digital human is determined through negotiation, enabling the second device to select the driving mode desired or supported by the first device, configure the corresponding driving resources, and achieve the successful execution of the digital human service.
It improved the success rate of digital human services, ensured the matching of digital human driving methods with terminal devices, reduced signaling overhead, and improved the success rate of digital human communication.
Smart Images

Figure CN2025075869_04122025_PF_FP_ABST
Abstract
Description
A method, communication system, and communication device for digital human communication
[0001] This application claims priority to Chinese Patent Application No. 202410685187.7, filed on May 29, 2024, entitled “A method, communication system and communication device for digital human communication”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of communication technology, and more specifically, to a method, communication system, and communication device for digital human communication. Background Technology
[0003] A digital human is a virtual character with a digital appearance, presented through display devices such as mobile phones, televisions, and virtual reality (VR) glasses. Driven by digital technology, digital humans can possess human-like behaviors, including language, facial expressions, and gestures, and can interact with people.
[0004] In real-time communication of digital humans, the computational resources required to drive the digital human are substantial. Some terminals, limited by factors such as hardware resources, size, weight, power consumption, and heat dissipation, are unable to drive the digital human. Currently, this is primarily achieved through network-side servers. For example, these terminals collect driving data (such as audio and video) and send it to the network-side server. The network-side server then uses this data to drive the digital human image and finally sends the generated, driven digital human to the other end of the terminal in video format.
[0005] However, digital human businesses based on existing technologies may fail. Summary of the Invention
[0006] This application provides a method, system, and device for digital human communication, which can help improve the success rate of executing digital human services.
[0007] In a first aspect, a method for digital human communication is provided, comprising: receiving first information and second information from a first device, the first information being used to identify a first digital human and the second information being used to indicate capabilities related to driving data, the driving data being used to drive the first digital human; and sending third information to the first device according to the second information, the third information being used to indicate a first driving mode, the first driving mode being a mode for driving the first digital human.
[0008] The solution described in the first aspect can be executed by a second device, which can be a terminal device or a network device, or a module (such as a chip system) within the terminal device or network device, or a logical node, logical module, or software capable of implementing all or part of the functions of the terminal device or network device. For ease of description, the second device will be used as an example below.
[0009] The statement that the third information is sent according to the second information can be understood as: the first driving mode is determined by the second device according to the second information, or the second device determines the first driving mode according to the capability related to the driving data indicated by the second information.
[0010] In the above scheme, the first device and the second device can exchange first information and second information. The second device selects the driving method desired or supported by the first device based on the above information, thereby driving the first digital human. This allows the first and second devices to negotiate the driving method of the first digital human, ensuring that the driving method adopted by the second device is the one desired or supported by the first device. This, in turn, enables the successful execution of digital human services initiated by the first device, thus improving the success rate of digital human service execution.
[0011] Simultaneously, this also supports the coordination or matching between the digital human business model selected by the first device and the digital human driving method adopted by the second device. For example, the processing requirements of 2D digital humans and 3D digital humans are different. For instance, 3D digital humans require 3D rendering, while 2D digital humans do not. When the first device selects a 3D digital human (which can be indicated through the digital human's identification information), the device used to drive the first digital human can use 3D rendering technology. When the first device selects a 2D digital human (which can be indicated through the digital human's identification information), the device used to drive the first digital human does not need to use 3D rendering technology. For another example, when the first device wants to drive a digital human based on facial features (which can be indicated through the second information that the first device supports generating driving data corresponding to the facial feature driving method), the device used to drive the first digital human can select the facial feature driving method.
[0012] In the first aspect, sending third information to the first device based on second information includes: sending third information to the first device based on the second information and the driving mode of the first digital human supported by the fourth device, wherein the fourth device is a device for driving the first digital human.
[0013] The fourth device can be either the second device or the third device.
[0014] Thus, the second device can determine the first driving mode based on the second information and the driving mode of the first digital human supported by the fourth device, which can support the successful execution of digital human services.
[0015] In the first aspect, the fourth device is the second device, and the method further includes: receiving first driving data from the first device, the first driving data being determined according to a first driving mode; and driving the first digital human according to the first driving data.
[0016] In this way, it can support the completion of the first digital human, and thus enable the completion of digital human business.
[0017] In the first aspect, sending third information to the first device based on second information includes: sending third information to the first device based on the second information and the driving mode of the first digital human supported by the second device.
[0018] In this way, the second device can combine the driving mode of the first digital human it supports and the second information to determine the driving mode corresponding to the digital human service between the first device and the peer device of the first device, which can support the successful execution of the digital human service.
[0019] In the first aspect, the method further includes: receiving fourth information, the fourth information being used to indicate the type of the first digital human. Specifically, sending information indicating a first driving mode to the first device based on second information includes: sending the information indicating the first driving mode to the first device based on the second information and the fourth information.
[0020] Thus, the second device can determine the first driving mode based on the second and fourth information, which can support the matching of the first driving mode with the type of the first digital human.
[0021] In the first aspect, the fourth device is the third device, and the method further includes: sending a request message to the third device, the request message including first information and third information, the request message being used to request driving resources for the first digital human; and receiving a response from the third device to the request message.
[0022] Based on the above process, the embodiments of this application can support a third device to configure corresponding driving resources for a first digital human.
[0023] In the first aspect, the request information also includes fourth information, which indicates the type of the first digital human.
[0024] Thus, embodiments of this application can support the third device in configuring corresponding driving resources according to the type of the first digital human. For example, if the fourth information indicates that the type of the first digital human is a 3D digital human, the third device configures 3D rendering resources for the first digital human; if the fourth information indicates that the type of the first digital human is a 2D digital human, the third device configures artificial intelligence (AI) inference resources for the first digital human, etc.
[0025] In a second aspect, a method for digital human communication is provided, comprising: sending first information and second information to a second device, the first information being used to identify a first digital human and the second information being used to indicate capabilities related to driving data (so that the second device can determine a driving mode for the first digital human), the driving data being used to drive the first digital human; and receiving third information from the second device, the third information being used to indicate a first driving mode, the first driving mode being a mode for driving the first digital human.
[0026] The solution described in the second aspect can be executed by a first device, which can be a terminal device, a module within the terminal device (such as a chip system), or a logical node, logical module, or software capable of implementing all or part of the functions of the terminal device. For ease of description, the first device will be used as an example in the following description.
[0027] The capabilities related to the drive data indicated by the second information described above can be used by the second device to determine the first drive mode.
[0028] In the above scheme, the first device and the second device can exchange first information and second information. The second device selects the driving method desired or supported by the first device based on the above information, thereby driving the first digital human. In this way, it can support the negotiation between the first device and the second device on the driving method of the first digital human, so that the driving method of the first digital human adopted by the second device is the driving method desired or supported by the first device, thereby enabling the successful execution of the digital human service initiated by the first device.
[0029] In the second aspect, the method further includes: sending first driving data, the first driving data being determined according to a first driving mode.
[0030] Optionally, the first device sends first driving data to a device for driving the first digital human. For example, if the second device is a device for driving the first digital human, the first device sends the first driving data to the second device; or if the third device is a device for driving the first digital human, the first device sends the first driving data to the third device.
[0031] Thus, the embodiments of this application can support the driving of the first digital human, thereby enabling the completion of digital human services.
[0032] In a second aspect, the method further includes: (to the second device) sending fourth information, the fourth information being used to indicate the type of the first digital human.
[0033] Thus, embodiments of this application can support the second device in determining the first driving mode based on the second information and the fourth information, which can support the matching of the first driving mode with the type of the first digital human.
[0034] In the second aspect, the determination of the first driving mode is related to the second information and the driving mode of the first digital human supported by the fourth device, which is a device for driving the first digital human. The fourth device can be either the second device or the third device.
[0035] Thus, embodiments of this application can support the second device in determining the driving mode corresponding to the digital human service between the first device and the peer device of the first device by combining the driving mode of the first digital human supported by the fourth device and the second information.
[0036] Thirdly, a method for digital human communication is provided, comprising: receiving a request message from a second device, the request message including first information and third information, the first information being used to identify a first digital human, the third information being used to indicate a first driving mode, the first driving mode being a mode for driving the first digital human, the determination of the first driving mode being related to the second information, the second information being used to indicate capabilities related to driving data, the driving data being used to drive the first digital human, the request message being used to request driving resources for the first digital human; and sending a response to the request message to the second device.
[0037] The solution described in the third aspect can be executed by a third device, which can be a network device, a module within a network device (such as a chip system), or a logical node, logical module, or software capable of implementing all or part of the functions of the network device. For ease of description, the following description uses a third device as an example.
[0038] The capabilities related to the drive data indicated by the second information described above can be used by the second device to determine the first drive mode.
[0039] In the above scheme, the third device can configure corresponding driving resources for the first digital human based on the information exchange with the second device regarding the driving mode of the first digital human, and can drive the first digital human based on these driving resources. Since the third device configures the corresponding driving resources according to the driving mode selected by the second device, these corresponding driving resources can be used to drive the first digital human. This enables the successful execution of digital human services between the first device and its peer device. For example, if the third information indicates that the driving mode of the first digital human is voice-driven, the third device can configure driving resources for voice-driven operation accordingly, but will not configure driving resources for text-driven operation.
[0040] In the third aspect, the determination of the first driving mode is related to the second information and the driving mode of the first digital human supported by the third device.
[0041] Thus, the embodiments of this application can support the second device in determining the driving mode corresponding to the digital human service between the first device and the peer device of the first device by combining the driving mode of the first digital human supported by the third device and the second information, which can support the successful execution of the digital human service.
[0042] In the third aspect, the request information also includes fourth information, which is used to indicate the type of the first digital human.
[0043] Thus, the third device configures the corresponding driving resources based on the fourth information. For example, when the third device determines that the type of the first digital human is a 3D digital human, the third device configures 3D rendering resources for the driving of the first digital human; or, when the third device determines that the type of the first digital human is a 2D digital human, the third device configures AI inference resources for the driving of the first digital human, etc.
[0044] In a third aspect, the method further includes: receiving first driving data from a first device, the first driving data being determined according to a first driving mode; and driving a first digital human according to the first driving data.
[0045] In this way, it is possible to support the completion of the first digital human, and thus support the completion of digital human business.
[0046] In conjunction with any one of the first to third aspects, the capability related to the driving data includes: the data type of the driving data, the driving method associated with the driving data, or the driving technology associated with the driving data.
[0047] Thus, embodiments of this application can support determining the first driving method based on one or more of the above.
[0048] Combining any one of the first to third aspects, the first driving method includes any one of the following: text-driven, voice-driven, video-driven, or facial feature-driven.
[0049] In combination with any one of the first to third aspects, the identifier of the first digital human can also be used to indicate the type of the first digital human.
[0050] Thus, embodiments of this application support determining the type of a first digital person based on its identifier, which can reduce the signaling overhead used to indicate the type of the first digital person.
[0051] In conjunction with any one of the first to third aspects, the second device includes at least one of the following: an over-the-top server, a multimedia telephony application server, a data channel application server, or a terminal device. The first device is a terminal device.
[0052] Fourthly, a method for digital human communication is provided, comprising: a first device sending first information and second information to a second device, the first information being used to identify a first digital human, and the second information being used to indicate capabilities related to driving data (so that the second device can determine the driving mode of the digital human), the driving data being used to drive the first digital human; the second device sending third information to the first device according to the second information, the third information being used to indicate a first driving mode, the first driving mode being a mode for driving the first digital human.
[0053] For a description of the fourth aspect, please refer to the foregoing descriptions of the first to third aspects, which will not be repeated here.
[0054] Fifthly, a communication system is provided, comprising: a first device and a second device. The first device is configured to send first information and second information to the second device, the first information being used to identify a first digital human, and the second information being used to indicate capabilities related to driving data (for the second device to determine a driving mode for the first digital human), the driving data being used to drive the first digital human; the second device is configured to receive the first information and the second information, and is further configured to send third information to the first device based on the second information, the third information being used to indicate a first driving mode, the first driving mode being a mode for driving the first digital human; the first device is configured to receive the third information.
[0055] For a description of the fifth aspect, please refer to the foregoing descriptions of the first to third aspects, which will not be repeated here.
[0056] In a sixth aspect, a communication device is provided, which may be a second device, or a device or module for performing the functions of the second device.
[0057] One possible implementation is that the communication device may include modules or units corresponding to the methods / operations / steps / actions described in the first aspect, which may be hardware circuits, software, or a combination of hardware circuits and software.
[0058] In a seventh aspect, a communication device is provided, which may be a first device, or a device or module for performing the functions of the first device.
[0059] One possible implementation is that the communication device may include modules or units corresponding to the methods / operations / steps / actions described in the second aspect, which may be hardware circuits, software, or a combination of hardware circuits and software.
[0060] Eighthly, a communication device is provided, which may be a third device, or a device or module for performing the functions of a third device.
[0061] One possible implementation is that the communication device may include modules or units corresponding to the methods / operations / steps / actions described in the third aspect, which may be hardware circuits, software, or a combination of hardware circuits and software.
[0062] A ninth aspect provides a communication device including a processor configured to, by executing a computer program or instructions, or by logic circuitry, cause the communication device to perform the method described in the first aspect and any possible mode of the first aspect; or cause the communication device to perform the method described in the second aspect and any possible mode of the second aspect; or cause the communication device to perform the method described in the third aspect and any possible mode of the third aspect; or cause the communication device to perform the method described in the fourth aspect and any possible mode of the fourth aspect.
[0063] In one possible implementation, the communication device also includes a memory for storing the computer program or instructions.
[0064] In one possible implementation, the communication device also includes a communication interface for inputting and / or outputting signals.
[0065] A tenth aspect provides a communication device including logic circuitry and an input / output interface for inputting and / or outputting signals. The input / output interface is configured to perform the method described in the first aspect and any possible mode of the first aspect; or, the logic circuitry is configured to perform the method described in the second aspect and any possible mode of the second aspect; or, the logic circuitry is configured to perform the method described in the third aspect and any possible mode of the third aspect; or, the logic circuitry is configured to perform the method described in the fourth aspect and any possible mode of the fourth aspect.
[0066] Eleventhly, a computer-readable storage medium is provided, on which a computer program or instructions are stored, which, when executed on a computer, cause the method described in the first aspect and any possible manner of the first aspect to be executed; or cause the method described in the second aspect and any possible manner of the second aspect to be executed; or cause the method described in the third aspect and any possible manner of the third aspect to be executed; or cause the method described in the fourth aspect and any possible manner of the fourth aspect to be executed.
[0067] In a twelfth aspect, a computer program product is provided, comprising instructions that, when executed on a computer, cause the method described in the first aspect and any possible mode of the first aspect to be executed; or cause the method described in the second aspect and any possible mode of the second aspect to be executed; or cause the method described in the third aspect and any possible mode of the third aspect to be executed; or cause the method described in the fourth aspect and any possible mode of the fourth aspect to be executed.
[0068] In a thirteenth aspect, a chip or chip system is provided, comprising: one or more processors configured to execute computer programs or instructions in the memory, such that the chip or chip system implements the methods of the first aspect and any possible implementation thereof; or, such that the chip or chip system implements the methods of the second aspect and any possible implementation thereof; or, such that the chip or chip system implements the methods of the third aspect and any possible implementation thereof; or, such that the chip or chip system implements the methods of the fourth aspect and any possible implementation thereof.
[0069] For a description of the beneficial effects of any of the fifth to thirteenth aspects, please refer to the description of the beneficial effects of the first to fourth aspects, which will not be repeated here. Attached Figure Description
[0070] Figure 1 is a schematic diagram of the architecture of the communication system 100 applicable to the embodiments of this application.
[0071] Figure 2 is a schematic diagram of the architecture of the communication system 200 applicable to the embodiments of this application.
[0072] Figure 3 is a schematic diagram of the architecture of the communication system 300 applicable to the embodiments of this application.
[0073] Figure 4 is a schematic diagram of the architecture of the communication system 400 applicable to the embodiments of this application.
[0074] Figure 5 is a schematic diagram of the interaction flow of the digital human communication method 500 according to an embodiment of this application.
[0075] Figure 6 is a schematic diagram of the interaction flow of the digital human communication method 600 according to an embodiment of this application.
[0076] Figure 7 is a schematic diagram of the interaction flow of the digital human communication method 700 according to an embodiment of this application.
[0077] Figure 8 is a schematic diagram of the interaction flow of the digital human communication method 800 according to an embodiment of this application.
[0078] Figure 9 is a schematic diagram of the interaction flow of a digital human communication method 900 according to an embodiment of this application.
[0079] Figure 10 is a schematic diagram of the interaction flow of the digital human communication method 1000 according to an embodiment of this application.
[0080] Figure 11 is a schematic block diagram of a communication device 1100 according to an embodiment of this application.
[0081] Figure 12 is a schematic block diagram of a communication device 1200 according to an embodiment of this application. Detailed Implementation
[0082] To facilitate understanding of the embodiments of this application, the following points will be explained first.
[0083] 1. Unless otherwise stated, “multiple” means two or more.
[0084] 2. Unless otherwise specified or in case of logical conflict, the terms and / or descriptions in different embodiments of this application are consistent and can be referenced in each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0085] III. The various numerical designations used in this application are merely for descriptive convenience and are not intended to limit the scope of protection of this application. The magnitude of the serial numbers used in this application does not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic. For example, the terms "first," "second," "third," "fourth," and other various terminology (if present) in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein.
[0086] Furthermore, any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner for ease of understanding.
[0087] IV. The terms “comprising” and “having” and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are expressly listed, but may include other steps or units that are not expressly listed or that are inherent to such process, method, product or device.
[0088] V. In this application, "for indicating" can be understood as "enabling", and "enabling" includes direct enabling and indirect enabling. When describing information for enabling A, it may include whether the information directly enables A or indirectly enables A, but it does not mean that the information necessarily carries A.
[0089] The information that enables the information is called the information to be enabled. In the specific implementation process, there are many ways to enable the information to be enabled, such as, but not limited to, directly enabling the information to be enabled, such as the information to be enabled itself or its index. It can also be indirectly enabled by enabling other information, where there is a relationship between the other information and the information to be enabled. It can also enable only a part of the information to be enabled, while the other parts are known or pre-agreed upon. For example, enabling specific information can be achieved by using a pre-agreed (e.g., protocol-defined) arrangement of various pieces of information, thereby reducing enabling overhead to some extent. Simultaneously, common parts of various pieces of information can be identified and enabled uniformly to reduce the enabling overhead caused by individually enabling the same information.
[0090] In addition, "instruction" can include direct instruction, indirect instruction, explicit instruction, and implicit instruction. When describing a certain instruction information to indicate A, it can be understood that the instruction information carries A, directly indicates A, or indirectly indicates A.
[0091] In this application, the information indicated by the instruction information is called the information to be instructed. In specific implementations, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is a relationship between the other information and the information to be instructed. It can also indicate only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction of specific information can be achieved by using a pre-agreed (e.g., protocol-defined) arrangement of various pieces of information, thereby reducing instruction overhead to some extent. Furthermore, the information to be instructed can be sent as a whole or divided into multiple sub-information pieces, and the sending period and / or timing of these sub-information pieces can be the same or different.
[0092] VI. In this application, "pre-configuration" may include pre-defined terms, such as protocol definitions. These "pre-defined terms" can be implemented by pre-storing corresponding codes, tables, or other means of indicating relevant information in the device (e.g., including various network elements). This application does not limit the specific implementation method.
[0093] VII. The term "storage" or "preservation" in this application can refer to storage in one or more memory devices. These memory devices can be separately configured or integrated into an encoder, decoder, processor, or communication device. Alternatively, some memory devices can be separately configured, while others can be integrated into a decoder, processor, or communication device. The type of memory can be any form of storage medium, and this is not limited.
[0094] 8. The “protocol” involved in this application may refer to standard protocols in the field of communications, such as fourth-generation (4G) network protocols, fifth-generation (5G) network protocols, 5.5G network protocols, and related protocols applied in future communication networks. This application does not limit the scope of the term.
[0095] 9. The arrows or boxes indicated by dashed lines in the schematic diagrams in the accompanying drawings of this application represent optional steps or optional modules.
[0096] 10. Unless otherwise stated, " / " indicates that the objects before and after are in an "or" relationship. For example, A / B can mean A or B. In this application, "and / or" is merely a description of the relationship between the related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. A and B can be singular or plural.
[0097] XI. In this application, "send" and "receive" indicate the direction of signal transmission. For example, "send information to XX" can be understood as the destination of the information being XX, which may include direct transmission via the air interface or indirect transmission by other units or modules via the air interface. "Receive information from YY" can be understood as the source of the information being YY, which may include direct reception from YY via the air interface or indirect reception from YY by other units or modules via the air interface. "Send" can also be understood as the "output" of a chip interface, and "receive" can also be understood as the "input" of a chip interface. In other words, sending and receiving can occur between devices, such as between network devices and terminal devices, or within a device, such as between components, modules, chips, software modules, or hardware modules within the device via a bus, wiring, or interface.
[0098] First, the communication system to which the embodiments of this application are applicable will be described.
[0099] Figure 1 is a schematic diagram of the architecture of a communication system 100 applicable to an embodiment of this application. As shown in Figure 1, the communication system 100 includes a first device and a second device. Optionally, the communication system 100 may further include a third device.
[0100] The first device is a device for determining or generating driving data for driving the digital human. The second device is a device with an information negotiation function, based on which the second device can determine the driving mode of the digital human through information interaction with the first device. The third device is a device for implementing the driving of the digital human; for example, the third device has a digital human driving function. The second device can also interact with the third device, and this interaction facilitates the second device in determining the driving mode of the digital human.
[0101] When the second device has the function of driving a digital human, the communication system 100 may not include the third device. When the second device does not have the function of driving a digital human, the communication system 100 may include the third device.
[0102] When the second device has the function of driving a digital human, it can send the video media stream corresponding to the driven digital human to the peer device of the first device. When the second device does not have the function of driving, the third device can send the video media stream corresponding to the driven digital human to the peer device of the first device. This can support the completion of digital human services between the first device and its peer device.
[0103] In other words, the second device can possess both negotiation and digital human driving functions simultaneously, or it can possess only negotiation functions. When the second device does not possess digital human driving functions, the second device and the third device work together to complete the digital human business between the first device and its peer device. When the second device possesses both digital human driving and negotiation functions, the second device independently completes the digital human business between the first device and its peer device.
[0104] One possible implementation is that the second device may include any of the following: an over-the-top (OTT) server, a multimedia telephony application server (MMTEL AS), a data channel application server (DC AS) (the aforementioned devices may be collectively referred to as network devices), or a terminal device.
[0105] It should be noted that when the second device is a terminal device, the terminal device may have a driving function. When the terminal device does not have a driving function, the terminal device may drive the digital human through other devices (such as an OTT server), and there are no restrictions on this.
[0106] In this embodiment, the communication system 100 can be applied to multiple different communication network architectures, such as IMS network or OTT network, as shown in Figures 2 to 4.
[0107] Figure 2 is a schematic diagram of the architecture of the communication system 200 applicable to the embodiments of this application. As shown in Figure 2, the communication system 200 includes, but is not limited to:
[0108] οDC AS can provide the function of digital human negotiation between the terminal and the IMS network during digital human calls;
[0109] The data channel signaling function (DCSF) provides DC channel management functionality.
[0110] οMMTEL AS can provide multimedia services;
[0111] The proxy-call session control function (P-CSCF), located in the visited network, is the entry node for SIP users to access the IMS network. It is mainly responsible for forwarding SIP signaling between SIP users and the home network.
[0112] The interrogating-call session control function (I-CSCF) is located in the home network and serves as the unified entry point for the home network. It is responsible for allocating or querying the S-CSCF that serves the user.
[0113] The serving-call session control function (S-CSCF) is located in the home network and is the central node of the IMS network. It is responsible for user registration, authentication, session management, routing, and service triggering.
[0114] The Internet Protocol Multimedia Subsystem Access Gateway (IMS-AGW) provides an anchor point for accessing the IMS network.
[0115] Optionally, the communication system 200 may also include a media function (MF) (by way of example only, and not limited to other terms), which can provide DC media processing capabilities and digital human driving capabilities.
[0116] Optionally, the communication system 200 may also include terminal equipment.
[0117] In the communication system 200, DC AS can be a second device, MF can be a third device, and terminal equipment can be a first device.
[0118] Figure 3 is a schematic diagram of the architecture of the communication system 300 applicable to the embodiments of this application. As shown in Figure 3, the communication system 300 includes, but is not limited to:
[0119] οMMTEL AS;
[0120] οP-CSCF;
[0121] οCSCF;
[0122] οS-CSCF;
[0123] οIMS-AGW.
[0124] Optionally, the communication system 300 may also include a media resource function (MRF) (by way of example only, and not limited to other terms), which can provide digital human-driven functionality.
[0125] Optionally, the communication system 300 may also include terminal equipment.
[0126] In the communication system 300, MMTEL AS can be a second device, MRF can be a third device, and terminal equipment can be a first device.
[0127] For information on the interfaces between network elements in Figures 2 and 3, please refer to the existing standards for the description of IMS networks; further details will not be provided here.
[0128] Figure 4 is a schematic diagram of the architecture of the communication system 400 applicable to the embodiments of this application. As shown in Figure 4, the communication system 400 includes: a terminal device and an OTT server. The OTT server can provide OTT services, including: OTT signaling processing, routing and media processing functions, providing digital human negotiation functions between the terminal device and IMS based on digital human calls, and digital human driving functions, etc.
[0129] In the communication system 400, the OTT server can be a second device (with information negotiation function and digital human driving function), and the terminal device can be a first device.
[0130] The network elements mentioned in the embodiments of this application, such as P-CSCF, I-CSCF, and S-CSCF, can be understood as network elements used to implement different functions, which can be combined into network slices as needed. These network elements can be independent hardware devices, integrated into the same hardware device, or network components within a hardware device. They can also be software functions running on dedicated hardware, or virtualized functional modules instantiated on a platform (e.g., a cloud platform). This application does not limit the specific form of the aforementioned network elements. It should also be understood that the above naming is only defined to facilitate the differentiation of different functions and should not constitute any limitation on this application.
[0131] The network elements or devices listed in the above architecture are merely illustrative examples. The communication architecture applicable to this application may also include other network elements or devices, and this application does not limit them.
[0132] Figures 2 and 3 illustrate the network architecture when the communication network is an IMS network. When the communication network is a circuit-switched (CS) network or other communication networks, some network elements in the above architecture can be replaced with other network elements, or the above architecture may also include other network elements. For example, in a CS network, communication system 200 and communication system 300 may include a mobile service switching center (MSC), and the IMS-AGW can be replaced with a media gateway (MGW), etc.
[0133] In this application embodiment, the terminal device is a device with wireless transceiver function, which may refer to user equipment (UE), access terminal, subscriber unit, user station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication device, user agent, or user device.
[0134] In this application embodiment, the terminal device can also be a satellite phone, cellular phone, smartphone, wireless data card, wireless modem, machine-type communication device, cordless phone, session initiation protocol (SIP) phone, wireless local loop (WLL) station, personal digital assistant (PDA), customer-premises equipment (CPE), point of sale (POS) machine, handheld device with wireless communication function, computing device or other processing device connected to a wireless modem, vehicle-mounted device, communication device mounted on a high-altitude aircraft, wearable device, drone, robot, terminal in device-to-device (D2D) communication, terminal in vehicle-to-everything (V2X) communication, virtual reality (VR) terminal device, augmented reality (AR) terminal device, wireless terminal in industrial control, wireless terminal in self-driving, telemedicine (or telehealth). Wireless terminals in services, smart grids, transportation safety, smart cities, smart homes, or terminal devices in communication networks that evolve after 5G are not subject to any restrictions.
[0135] In this embodiment of the application, the terminal device may also be a device with communication function in a future communication network, and the form or type of the terminal device in the future communication network is not limited.
[0136] In this application embodiment, the communication device used to implement the functions of the terminal device can be the terminal device itself, or it can be a device capable of supporting the terminal device in implementing the functions, such as a chip system. This device can be installed in the terminal device or used in conjunction with the terminal device. In this application, the chip system can be composed of chips, or it can include chips and other discrete components.
[0137] In this embodiment, the network device (which may include an OTT server, MMTEL AS, or DC AS, etc.) is a device with wireless transceiver capabilities used to communicate with terminal devices. The network device can be a node in the RAN, also known as a base station or RAN node. It can be a long-term evolution (LTE) eNB; or a base station in a 5G network such as a gNB; or a base station in a public land mobile network (PLMN) evolved after 5G; a broadband network gateway (BNG); an aggregation switch; or a 3rd Generation Partnership Project (3GPP) network. rd Network equipment, etc. in the Generation Partnership Project (3GPP).
[0138] Network equipment can also include various forms of base stations, such as: macro base stations, micro base stations (also known as small stations), relay stations, TRPs, transmission points (TPs), mobile switching centers, and equipment that performs base station functions in device-to-device (D2D), vehicle-to-everything (V2X), and machine-to-machine (M2M) communications, as well as network equipment in non-terrestrial networks (NTNs), etc., without specific limitations.
[0139] In this embodiment, the communication device used to implement the functions of the network device can be the network device itself, or it can be a device that supports the network device in implementing those functions, such as a chip system. This device can be installed in the network device or used in conjunction with the network device. The chip system in this embodiment can be composed of chips, or it can include chips and other discrete components.
[0140] In the communication system 100, the first device and the second device can exchange information regarding the driving mode of the first digital human. For example, the first device sends information to the second device indicating the driving mode of the first digital human that the first device expects or supports, and the second device determines the corresponding driving mode of the first digital human based on this information. This allows the first and second devices to negotiate the driving mode of the first digital human, ensuring that the driving mode adopted by the second device is the driving mode expected or supported by the first device, thereby enabling the successful execution of digital human services initiated by the first device.
[0141] For information exchange between the first device and the second device, please refer to Figure 5.
[0142] For ease of understanding and explanation, the following description of the digital human communication method of this application uses the interaction between the first device and the second device as an example, but this should not constitute any limitation on the subject executing the digital human communication method. For example, the method executed by the device (such as the first device and / or the second device) can also be executed by a module in the device (such as a circuit, chip, or chip system, etc.), or by a logic node, logic module, or software that can implement all or part of the functions of the device, without limitation.
[0143] Figure 5 is a schematic diagram of the interaction flow of a digital human communication method 500 according to an embodiment of this application. As shown in Figure 5, method 500 includes:
[0144] S501, the first device sends first information and second information to the second device. Correspondingly, the second device receives the first information and second information.
[0145] The first information is used to identify the first digital person, or the first information can be the identification information of the first digital person, or the first information can be the first identification information, which is used to identify the first digital person. The first digital person is the digital person selected by the first device. The first information can be used by the second device to determine that the digital person selected by the first device is the first digital person. That is, the first digital person is used for the execution of digital person services between the first device and its peer device. In other words, the peer device of the first device can display the driving video media stream corresponding to the first digital person, thereby enhancing the fun of communication between the first device and its peer device.
[0146] The second information is used to indicate capabilities related to the animation data (hereinafter referred to as animation data S, which includes one or more animation data sets, without limitation), or, the second information is used to indicate capabilities related to the animation data S for the second device to determine the driving mode of the first digital human. That is, the second device determines the first driving mode based on the second information or based on the capabilities related to the animation data S indicated by the second information. The first driving mode is the mode used to drive the first digital human. The animation data S is the data used to drive the first digital human. The animation data S is generated or determined by the first device. Different animation data sets are associated with different driving modes. For example, animation data 1 is associated with driving mode 1, animation data 2 is associated with driving mode 2, etc. The second device can determine the first driving mode based on the association between the animation data and the driving mode.
[0147] In one possible implementation, the second information may also indicate relevant information used by the second device to determine the driving mode (such as the first driving mode) of the first digital human. For example, the relevant information may include, but is not limited to: capabilities related to the driving data S, the driving mode of the first digital human supported by the first device, or the driving technology of the first digital human supported by the first device, or the type of driving data of the first digital human supported by the first device, etc.
[0148] In summary, the second device can select the driving mode of the first digital human based on the second information, and the driving mode of the first digital human selected by the second device can meet the requirements or expectations of the first device.
[0149] A description of how the second information is used to determine the driving mode of the first digital human can be found in Table 1. The content of Table 1 is for illustrative purposes only and is not intended as a final limitation. As an example, the correspondence between the capabilities associated with the driving data S and the driving mode of the first digital human can exist in the form of tables, functions, text, or strings, such as storage or transmission.
[0150] Table 1
[0151] As shown in Table 1:
[0152] The second information indicates that the capability related to the drive data S is capability 1, and the second device determines the drive mode a1 accordingly.
[0153] The second information indicates that the capability related to the drive data S is capability 2, and the second device determines the drive mode a2 accordingly.
[0154] The second information indicates that the capability related to the drive data S is capability 3, and the second device determines the drive mode a3 accordingly;
[0155] The second information indicates the capability related to the drive data S as capability 4, and the second device determines the drive mode a4 accordingly.
[0156] Thus, the second device can determine the corresponding driving mode based on the capability related to the driving data S indicated by the second information.
[0157] One possible implementation, the capabilities related to the driving data S, include, but are not limited to:
[0158] The data type that drives the data S;
[0159] The driving method associated with the driving data S; or,
[0160] The driving technology that associates data with the driver.
[0161] When the second information indicates the type of drive data S, the second device determines the drive mode of the first digital human supported or desired by the first device based on the type of drive data S, and can determine the first drive mode accordingly.
[0162] When the second information indicates the driving mode associated with the driving data S, the second device determines the driving mode of the first digital human supported or desired by the first device based on the driving mode associated with the driving data S, and can determine the first driving mode accordingly.
[0163] When the second information indicates the driving technology associated with the driving data S, the second device determines the driving technology corresponding to the driving data S based on the driving technology associated with the driving data S, and determines the corresponding driving method based on the driving technology. For example, the first driving technology needs to use the first driving data, and the type of the first driving data is related to the first driving method. For example, text data corresponds to text driving, and voice driving corresponds to voice driving, etc.
[0164] Taking the data type S driven by capability as an example:
[0165] For example, capability 1 indicates that the data type of the driving data S is text data, and the second device determines that the first driving mode is text driving;
[0166] For example, capability 2 indicates that the data type of the driving data S is voice data, and the second device determines that the first driving mode is voice driving;
[0167] For example, capability 3 indicates that the data type of the driving data S is video data, and the second device determines that the first driving mode is video driving;
[0168] For example, capability 4 indicates that the data type of the driving data S is facial feature data, and the second device determines that the first driving mode is facial feature driving.
[0169] Taking the capability-driven data S association method as an example:
[0170] For example, capability 1 indicates that the drive type associated with drive data S is text drive, and the second device determines that the first drive mode is text drive;
[0171] For example, if capability 2 indicates that the drive type associated with drive data S is voice drive, the second device determines that the first drive mode is voice drive.
[0172] For example, capability 3 indicates that the drive type associated with the drive data S is video drive, and the second device determines that the first drive mode is video drive;
[0173] For example, capability 4 indicates that the driving type associated with the driving data S is facial feature driving, and the second device determines that the first driving mode is facial feature driving.
[0174] Taking the technology that uses capability as the driving force to associate data S as an example:
[0175] For example, capability 1 indicates that the driving technology associated with the driving data S is 3D rendering, and the second device determines that the first driving method is voice driving.
[0176] For example, if capability 2 indicates that the driving technology associated with the driving data S is speech recognition, the second device determines that the first driving method is speech driving.
[0177] For example, capability 3 indicates that the driving technology associated with the driving data S is deep neural network (DNN) inference, and the second device determines that the first driving method is one of video driving, voice driving, and text driving.
[0178] For example, capability 4 indicates that the driving technology associated with the driving data S is DDN inference + speech recognition, and the second device determines that the first driving mode is either speech driving or text driving.
[0179] In this embodiment, the driving data (which may be driving data S or the first driving data hereinafter, without limitation) may also include source data. That is, the device for driving the first digital human can drive the first digital human based on the driving data or the source data. Source data refers to the originally acquired data, such as audio data acquired through a microphone or video data acquired through a camera, which does not require additional conversion processing. Driving data is generally data obtained after converting the source data, such as facial feature point data converted from video data.
[0180] One possible implementation, method 500, may also include:
[0181] S501a, the first device sends the fourth information to the second device. Correspondingly, the second device receives the fourth information.
[0182] The fourth piece of information is used to indicate the type of the first digital human, for example, whether the first digital human is a 2D digital human or a 3D digital human. Different types of digital humans can be associated with different driving methods.
[0183] Accordingly, the second device can determine the first driving mode based on the second and fourth information. This allows for matching the first driving mode with the type of the first digital human.
[0184] For example, the second device determines driving mode 1 and driving mode 2 based on the second information. Driving mode 1 is applicable to type 1, such as a 2D digital human, and driving mode 2 is applicable to type 2, such as a 3D digital human. The fourth information indicates that the type of the first digital human is type 1, and the second device determines the first driving mode as driving mode 1.
[0185] In this embodiment of the application, the main difference between 2D digital humans and 3D digital humans is:
[0186] Technical implementation: 2D digital humans are mainly "paper-thin" figures generated from images and video materials, and the production process is relatively simple; 3D digital humans, on the other hand, require complex modeling and 3D rendering technologies to achieve higher realism and three-dimensionality.
[0187] Visual effects and realism: 2D digital humans usually have simple flat shapes and styles, and can present unique artistic styles and cartoon characteristics; 3D digital humans can move and be observed freely in three-dimensional space, with higher realism and three-dimensionality, and can present more details, lighting effects and realism.
[0188] Production complexity and technical requirements: The production of 2D digital humans usually has lower technical requirements and can be created using simple drawing tools and software; the production of 3D digital humans has higher technical requirements and requires the use of professional 3D modeling, animation and rendering techniques, as well as more powerful computing resources to create and present them.
[0189] ο Dynamism and interactivity: 2D digital humans may be limited in terms of dynamism and interactivity, and are usually presented as static images or simple animations; 3D digital humans have greater dynamism and interactivity, can perform complex movements and expressions, interact with the audience in real time, and provide a richer interactive experience in the virtual environment.
[0190] Application scenarios: 2D digital humans are more widely used in short videos, live streaming, real-time communication and other scenarios, which can significantly reduce the human and time costs of video creation; 3D digital humans are more used in occasions that require high realism and interactivity, such as virtual reality applications, games, movies and so on.
[0191] One possible implementation is that the identifier of the first digital human can also be used to indicate the type of the first digital human.
[0192] For example, the format of the identifier of a digital human corresponding to a 2D digital human is different from the format of the identifier of a digital human corresponding to a 3D digital human. The second device determines the type of the first digital human based on the format of the identifier of the first digital human. The identifier of the first digital human is identifier 1, and the format of identifier 1 corresponds to type 1; the identifier of the digital human is identifier 2, and the format of identifier 2 corresponds to type 2.
[0193] For example, different digital human identifiers are associated with different types. For instance, the digital human identifiers corresponding to 2D digital humans include identifier 1 and identifier 2, and the digital human identifiers corresponding to 3D digital humans include identifier 3 and identifier 4. The second device can determine the type of the corresponding digital human based on the identifier of the first digital human.
[0194] Thus, embodiments of this application support determining the type of a first digital person based on its identifier, which can reduce the signaling overhead used to indicate the type of the first digital person.
[0195] In summary, the second device determines the type of the first digital human based on its identification information, and determines the corresponding driving method based on the type of the first digital human.
[0196] S502, the second device sends third information to the first device based on the second information. Correspondingly, the first device receives the third information.
[0197] The third information is used to indicate the first driving mode. The first driving mode is determined by the second device based on the second information, or the first driving mode is determined by the second device based on the capabilities related to the driving data S indicated by the second information. For details, please refer to the above description of the capabilities related to the driving data S, which will not be repeated here.
[0198] Through S501 and S502, the first device and the second device can interact with the first information and the second information. The second device can select the driving mode desired or supported by the first device according to the above information, thereby realizing the driving of the first digital human. This can make the driving mode of the first digital human desired or supported by the first device supported or achievable by the second device.
[0199] Simultaneously, this also supports the coordination or matching between the digital human business model selected by the first device and the digital human driving method adopted by the second device. For example, the processing requirements of 2D digital humans and 3D digital humans are different. For instance, 3D digital humans require 3D rendering, while 2D digital humans do not. When the first device selects a 3D digital human (which can be indicated through the digital human's identification information), the device used to drive the first digital human can use 3D rendering technology. When the first device selects a 2D digital human (which can be indicated through the digital human's identification information), the device used to drive the first digital human does not need to use 3D rendering technology. For another example, when the first device wants to drive a digital human based on facial features (which can be indicated through the second information that the first device supports generating driving data corresponding to the facial feature driving method), the device used to drive the first digital human can select the facial feature driving method.
[0200] One possible implementation is that the second device sends third information to the first device based on the second information, which may include:
[0201] The second device can send third information to the first device based on the second information and the driving mode of the first digital human supported by the fourth device. Alternatively, the first driving mode is determined by the second device based on the second information and the driving mode of the first digital human supported by the fourth device. The fourth device can be either the second device or the third device.
[0202] For example, taking the fourth device as the second device:
[0203] After the second device determines at least one driving mode of the first digital human based on the second information, the second device determines the first driving mode from the above at least one driving mode based on the driving mode of the first digital human supported by the second device.
[0204] For example, the second device determines driving mode 1 and driving mode 2 based on the second information, and the driving mode of the first digital human supported by the second device is driving mode 1. The second device determines the first driving mode as driving mode 1.
[0205] For example, taking the fourth device as the second device:
[0206] The second device supports the first digital human driving mode 1 and driving mode 2. The second device determines driving mode 1 and driving mode 3 according to the second information. The second device determines the first driving mode as driving mode 1.
[0207] In summary, the embodiments of this application can support the second device in determining the driving mode corresponding to the digital human service between the first device and the peer device of the first device by combining the driving mode of the first digital human it supports and the second information, which can support the successful execution of the digital human service.
[0208] The above description uses the second device performing the digital human driving function as an example. When the second device does not perform the digital human driving function, the second device can determine the first driving mode based on the driving mode of the first digital human supported by the device performing the digital human driving function.
[0209] For example, taking the fourth device as the third device:
[0210] After the second device determines at least one driving mode of the first digital human based on the second information, the second device determines the first driving mode based on the driving mode of the first digital human supported by the third device.
[0211] For example, the second device determines driving mode 1 and driving mode 2 based on the second information, and the third device supports driving mode 1 for the first digital human. The second device then determines the first driving mode as driving mode 1. The second and third devices can exchange information regarding the driving mode of the first digital human supported by the third device, thus allowing the second device to determine the driving mode of the first digital human supported by the third device.
[0212] For example, let's take the fourth device as the third device:
[0213] The third device supports two driving modes for the first digital human: driving mode 1 and driving mode 2. The second device determines driving mode 1 based on the second information, and the second device determines the first driving mode as driving mode 1. The second device and the third device can exchange information regarding the driving modes of the first digital human supported by the third device, thus allowing the second device to determine the driving modes of the first digital human supported by the third device.
[0214] In this embodiment, the method and execution order of information exchange between the second device and the third device regarding the driving mode of the first digital human supported by the third device are not limited.
[0215] Thus, embodiments of this application can support the second device in determining the driving mode corresponding to the digital human service between the first device and the peer device of the first device by combining the driving mode of the first digital human supported by the fourth device and the second information, which can support the successful execution of the digital human service.
[0216] Optionally, method 500 may also include:
[0217] S503a, the first device sends first drive data to the second device. Correspondingly, the second device receives the first drive data.
[0218] After receiving the third information, the first device determines the first driving mode based on the third information, and determines or generates the first driving data corresponding to the first driving mode accordingly.
[0219] For example, the first driving method is text-driven, and the first driving data is text data; the first driving method is voice-driven, and the first driving data is voice data; the first driving method is video-driven, and the first driving data is video data; the first driving method is facial feature-driven, and the first driving data is facial feature data.
[0220] S503b, the second device drives the first digital human according to the first driving data.
[0221] For example, the second device can drive the digital human based on the user's voice, text, or video to generate the digital human's movements. The movements can be partial, such as facial movements, or full-body, such as limb movements.
[0222] For a description of how the second device drives the first digital human based on the first driving data, please refer to existing standards, which will not be repeated here.
[0223] Through S503a and S503b, the embodiments of this application support the completion of driving the first digital human, thereby enabling the completion of digital human services.
[0224] When the second device is the peer device of the first device, the second device displays the driving video media stream corresponding to the first digital human. When the second device is not the peer device of the first device, the second device sends the driving video media stream corresponding to the first digital human to the peer device of the first device.
[0225] The above S503a and S503b are examples of the second device driving the first digital human, but the first digital human can also be driven by a third device, as can be seen in the following content.
[0226] Optionally, method 500 may also include:
[0227] S504a, the second device sends a request message to the third device. Correspondingly, the third device receives the request message.
[0228] The aforementioned request message may include first information and third information. The request message can be used to request a third device to reserve or pre-configure driving resources for the first digital human. The driving resources reserved or pre-configured by the third device for the first digital human can be used by the third device to drive the first digital human.
[0229] In one possible implementation, the aforementioned request message may also include a fourth piece of information. In this way, the third device configures the corresponding driving resources based on the fourth piece of information. For example, when the third device determines that the type of the first digital human is a 3D digital human, the third device configures 3D rendering resources for the first digital human's driving; or, when the third device determines that the type of the first digital human is a 2D digital human, the third device configures AI inference resources for the first digital human's driving, etc.
[0230] In one possible implementation, the request message may further include the address information of the first device. Correspondingly, the response message to the request message includes the address information of the third device.
[0231] Thus, the third device establishes a connection with the first device based on the address information of the first device, and the first device establishes a connection with the third device based on the address information of the third device. Once a connection is established between the first device and the third device, the first device can send driving data to the third device to drive the first digital human.
[0232] S504b, the third device sends a response (indicated by a response message) to the second device in response to the request message. Accordingly, the second device receives the response message.
[0233] For example, the response message may indicate that the third device agrees to reserve driving resources for the first digital human, such as by including an acknowledgment (ACK); or, the response message may indicate that the third device does not agree to reserve driving resources for the first digital human, such as by including a non-acknowledgment (NACK).
[0234] Through S504a and S504b, the third device configures corresponding driving resources for the first digital human based on information exchange with the second device regarding the driving mode of the first digital human, and drives the first digital human based on these driving resources. Since the third device configures the corresponding driving resources according to the driving mode selected by the second device, these driving resources can be used to drive the first digital human. This enables the successful execution of digital human services between the first device and its peer device. For example, if the third information indicates that the driving mode of the first digital human is voice-driven, the third device configures driving resources for voice-driven operation accordingly; the third device will not configure driving resources for text-driven operation, etc.
[0235] The solutions described in S503a and S503b above are two different implementations of the solutions described in S504a and S504b above. In specific implementation, one of them can be selected for execution. Therefore, the above numbering is only for example and not for limitation. That is, the solution described in S504 (including S504a and S504b) is not a further solution of the solution described in S503 (including S503a and S503b).
[0236] When the third device drives the first digital human, S504a can be performed before or after S502, without limitation.
[0237] Optionally, method 500 may also include:
[0238] S504c, the first device sends first drive data to the third device. Correspondingly, the third device receives the first drive data.
[0239] S504d, the third device drives the first digital human according to the first driving data.
[0240] Through the above-described S504a and S504b, the embodiments of this application can support the completion of driving the first digital human, thereby enabling the completion of digital human services between the first device and its peer device.
[0241] The following description, in conjunction with other accompanying drawings, further describes the information interaction between the first and second devices in Figure 5.
[0242] Figure 6 is a schematic diagram of the interaction flow of a digital human communication method 600 according to an embodiment of this application. As shown in Figure 6, the first device is a first terminal device, and the second device is a network device, such as an IMS network, an OTT server, etc. Method 600 includes:
[0243] S601, the first terminal device sends a digital human call request message to the network device. Correspondingly, the network device receives the digital human call request message.
[0244] The digital human call request message is used to request a digital human call between the first terminal device and the second terminal device. The digital human call request message may include the aforementioned first information and second information.
[0245] Optionally, the digital human call request message may also include a fourth piece of information.
[0246] S602, The network device determines the first driving mode.
[0247] For example, the network device determines the first driving mode based on the first information, the second information, and the driving mode of the first digital human supported by the network device. See Figure 5 for details, which will not be elaborated further.
[0248] S603, the network device sends a digital human call response message to the first terminal device. Correspondingly, the first terminal device receives the digital human call response message.
[0249] The digital human call response message may include third information. This digital human call response information is used to indicate that the network device can support digital human services between the first terminal device and the second terminal device.
[0250] S604, The first terminal device sends the first driving data to the network device.
[0251] For example, the first terminal device generates corresponding first driving data according to the first driving method and sends the first driving data to the network device.
[0252] S605, The network device drives the first digital human according to the first driving data.
[0253] For details, please refer to the descriptions in S503b or S504d, which will not be elaborated upon here.
[0254] S606, the network device sends the driving video media stream corresponding to the first digital human to the second terminal device. Correspondingly, the second terminal device receives the driving video media stream corresponding to the first digital human.
[0255] Using the above method, the digital human service between the first terminal device and the second terminal device can be successfully executed.
[0256] Figure 6 illustrates an example of a first terminal device initiating a digital human call request, but it is not limited to the scenario where the network device initiates the digital human call request. For instance, the network device sends a digital human call request message to the first terminal device, and the first terminal device sends a digital human call response message to the network device. This digital human call response message includes first information and second information. In this way, the network device and the first terminal device can complete information exchange regarding the driving method of the first digital human.
[0257] Figure 7 is a schematic diagram of the interaction flow of a digital human communication method 700 according to an embodiment of this application. As shown in Figure 7, the first device is a first terminal device, the second device is a DC AS, and the first terminal device and the DC AS need to communicate through P / S-CSCF, MMTEL AS, and DCSF. The third device is MF. Figure 7 takes the first terminal device interacting with information via SIP messages as an example. Method 700 includes:
[0258] Audio and video media channels and DC channels can be established between S701, the first terminal device, P / S-CSCF, MMTEL AS, DSCF, MF, and the second terminal device.
[0259] For details, please refer to the description in the existing standards, which will not be repeated here.
[0260] S702, the first terminal device sends an Update / re-INVITE message to the MMTE AS. Correspondingly, the MMTE AS receives the Update / re-INVITE message.
[0261] The Update / re-INVITE message is used to request a digital human call, or to request digital human services between a first terminal device and a second terminal device.
[0262] The process of the first terminal device sending an Update / re-INVITE message to the MMTEL AS may include: the first terminal device sending an Update / re-INVITE message to the P / S-CSCF, and the P / S-CSCF sending an Update / re-INVITE message to the MMTEL AS.
[0263] The Update / re-INVITE message includes an Avatar_Communication field. The header of the Avatar_Communication field indicates a digital human call request. The Avatar_Communication field can have the following parameters:
[0264] οAvatar_ID (first information), indicating the digital human identifier used in this call;
[0265] `οAvatar_Type` (fourth information) (optional parameter) indicates the type of digital human used in this call. Examples of values are: 2D / 3D. The type of digital human can also be determined based on its identifier.
[0266] οSupported_Animation_Mode (secondary information) indicates the driver mode supported by the terminal, i.e. the driver data types that can be sent. This parameter can have multiple values, which are arranged in order of priority. Examples of values: text, audio, facial_expression.
[0267] S703 and MMTEL AS / DCSF send Hypertext Transfer Protocol (HTTP) POST messages to DC AS. Correspondingly, DC AS receives the HTTP POST messages.
[0268] In this embodiment of the application, the MMTEL AS / DCSF sends an HTTP POST message to the DC AS, including: the MMTEL AS converts the Update / re-INVITE request message into an HTTP POST message, and forwards it to the DC AS via the DCSF. The HTTP POST message includes first information and second information.
[0269] For a description of HTTP POST messages, please refer to Tables 2 and 3.
[0270] Table 2
[0271] Table 2 describes the process using an HTTP POST message in JSON format as an example. Thus, the DC AS can select the driving method for the first digital human based on the information shown in Table 2.
[0272] Table 3
[0273] Table 3 describes the process using an HTTP POST message in XML format as an example. Thus, the DC AS can select the driving method for the first digital human based on the information shown in Table 2.
[0274] S704 and DC AS determine the first driving mode.
[0275] For example, DC AS determines the first driving mode based on the first and second information in the HTTP POST message, as described above.
[0276] Optionally, DC AS can also determine the first driving mode based on the driving mode of the first digital human supported by MF, the first information, and the second information, as described above.
[0277] S705, DC AS and MF exchange request and response messages.
[0278] For example, DC AS sends a request message to MF through DCSF / MMTEL AS, and MF sends a response message to DC AS through MMTEL AS / DCSF.
[0279] For a description of the request and response messages, please refer to the previous text; it will not be repeated here.
[0280] S706, DC AS sends a 200 for POST message to DCSF / MMTEL AS. Correspondingly, DCSF / MMTEL AS receives the 200 for POST message.
[0281] A 200 for POST message includes third-party information.
[0282] The 200 for POST message includes an Avatar_Communication field, which contains the following parameters:
[0283] Required_Animation_Mode = audio, indicating that the first driving mode is voice driving.
[0284] When DC AS indicates the first driving mode, the first terminal device can determine the first driving data according to the first driving mode.
[0285] S707 and MMTEL AS send a 200 for Update / re-INVITE message to the first terminal device. Correspondingly, the first terminal device receives the 200 for Update / re-INVITE message.
[0286] The 200 for Update / re-INVITE message includes third-party information. The 200 for Update / re-INVITE message is determined based on the 200 for POST message. This application does not limit the format of the 200 for Update / re-INVITE message.
[0287] S708, the first terminal device sends first drive data to the MF. Correspondingly, the MF receives the first drive data.
[0288] S709 and MF drive the first digital human based on the first drive data.
[0289] S710 and MF send the driving video media stream corresponding to the first digital human to the second terminal device. Correspondingly, the second terminal device receives the driving video media stream corresponding to the first digital human.
[0290] Using the above method, the digital human service between the first terminal device and the second terminal device can be successfully executed.
[0291] Figure 7 illustrates the scenario where the first terminal device sends an Update / re-INVITE message to the DC AS, but it is not limited to the scenario where the first terminal device passively sends an Update / re-INVITE message to the DC AS. For example, the DC AS sends a re-INVITE message to the first terminal device, and the first terminal device sends a response message to the DC AS for the re-INVITE message. The response message for the re-INVITE message includes first information and second information, etc.
[0292] Figure 8 is a schematic diagram of the interaction flow of a digital human communication method 800 according to an embodiment of this application. As shown in Figure 8, the first device is a first terminal device, the second device is a DC AS, and the first terminal device and the DC AS need to communicate through MF. The third device is the MF. Figure 8 takes the first terminal device interacting with information through a DC channel as an example. Method 800 includes:
[0293] Audio and video media channels and DC channels can be established between S801, the first terminal device, P / S-CSCF, MMTEL AS, DSCF, MF, and the second terminal device.
[0294] For details, please refer to the description in the existing standards, which will not be repeated here.
[0295] S802, the first terminal device sends a digital human call request message to the MF. Correspondingly, the MF receives the digital human call request message.
[0296] The Digital Human Call Request message is used to request a digital human call. The message includes first information and second information.
[0297] The digital human call request message includes an Avatar_Communication field, the header of which indicates the purpose of the digital human call request. See Figure 7 for a description of the Avatar_Communication field.
[0298] S803 and MF send a digital human call request message to DC AS. Correspondingly, DC AS receives the digital human call request message.
[0299] S804 and DC AS determine the first driving mode.
[0300] For example, DC AS determines the first driving mode based on the first and second information in the digital human request message, as can be seen above.
[0301] Optionally, DC AS can also determine the first driving mode based on the driving mode of the first digital human supported by MF, the first information, and the second information, as described above.
[0302] S805, DC AS and MF exchange request and response messages.
[0303] For example, DC AS sends a request message to MF through DCSF / MMTEL AS, and MF sends a response message to DC AS through MMTEL AS / DCSF.
[0304] For a description of the request and response messages, please refer to the previous text; it will not be repeated here.
[0305] S806 and DC AS send a digital human call response message to MF. Correspondingly, MF receives the digital human call response message.
[0306] The digital human call response message includes third-party information.
[0307] S807 and MF send a digital human call response message to the first terminal device. Correspondingly, the first terminal device receives the digital human call response message.
[0308] S808: The first terminal device sends the first driving data to the MF. Correspondingly, the MF receives the first driving data.
[0309] S809 and MF drive the first digital human based on the first driving data.
[0310] S810 and MF send the driving video media stream corresponding to the first digital human to the second terminal device. Correspondingly, the second terminal device receives the driving video media stream corresponding to the first digital human.
[0311] Through the above, the digital human service between the first terminal device and the second terminal device can be successfully executed.
[0312] Figure 9 is a schematic diagram of the interaction flow of a digital human communication method 900 according to an embodiment of this application. As shown in Figure 9, the first device is a first terminal device, the second device is an MMTEL AS, the first terminal device and the MMTEL AS communicate via P / S-CSCF, and the third device is an MRF. Figure 9 takes the first terminal device interacting with information via SIP messages as an example. Method 900 includes:
[0313] S901, the first terminal device sends an INVITE message (or an Update / re-INVITE message, etc.) to the MMTEL AS. Correspondingly, the MMTEL AS receives the INVITE message.
[0314] The INVITE message is used to request a digital human conversation. The INVITE message includes the aforementioned first and second information.
[0315] The INVITE message includes an Avatar_Communication field, whose header indicates a request for a digital human call. See Figure 7 for a description of the Avatar_Communication field.
[0316] S902 and MMTEL AS determine the first drive mode.
[0317] For a detailed description, please refer to the previous text, which will not be repeated here.
[0318] S903, MMTEL AS, and MRF exchange request and response messages.
[0319] For a description of the request and response messages, please refer to the previous text; it will not be repeated here.
[0320] S904, MMTEL AS sends an 18X / 200 for INVITE message to the first terminal device. Correspondingly, the first terminal device receives the 18X / 200 for INVITE message.
[0321] The 18X / 200 for INVITE message includes third-party information.
[0322] S905, the first terminal device sends the first driving data to the MRF. Correspondingly, the MRF receives the first driving data.
[0323] S906 and MRF drive the first digital human based on the first driving data.
[0324] S907 and MRF send the driving video media stream corresponding to the first digital human to the second terminal device. Correspondingly, the second terminal device receives the driving video media stream corresponding to the first digital human.
[0325] Through the above, the digital human service between the first terminal device and the second terminal device can be successfully executed.
[0326] Figure 9 illustrates an example of a first terminal device actively sending an INVITE message to the MMTEL AS, but the scenario is not limited to the MMTEL AS sending an INVITE message to the first terminal device. For instance, when the MMTEL AS sends an INVITE message to the first terminal device, the first terminal device sends a response message to the MMTEL AS for that INVITE message, such as the 18X / 200 for INVITE message, which includes first and second information. Correspondingly, the MMTEL AS can send a provisional response acknowledgement (PRACK) / acknowledgement (ACK) message to the first terminal device, which includes third information.
[0327] Figure 10 is a schematic diagram of the interaction flow of a digital human communication method 1000 according to an embodiment of this application. As shown in Figure 10, the first device is a first terminal device, the second device is a second terminal device, and the network device is an OTT server or an IMS network. Method 1000 includes:
[0328] S1001, the first terminal device sends an INVITE message to the second terminal device. Correspondingly, the second terminal device receives the INVITE message.
[0329] The INVITE message is used to request a digital human conversation. The INVITE message includes first information and second information.
[0330] The first terminal device can exchange information with the second terminal device through network equipment.
[0331] The INVITE message includes an Avatar_Communication field, whose header indicates a request for a digital human call. See Figure 7 for a description of the Avatar_Communication field.
[0332] S1002, The second terminal device determines the first driving mode.
[0333] For a detailed description, please refer to the previous text, which will not be repeated here.
[0334] S1003, the second terminal device sends an 18X / 200 for INVITE message to the first terminal device. Correspondingly, the first terminal device receives the 18X / 200 for INVITE message.
[0335] The 18X / 200 message includes third-party information.
[0336] S1004, the first terminal device sends first driving data to the second terminal device. Correspondingly, the second terminal device receives the first driving data.
[0337] S1005. The second terminal device drives the first digital human according to the first driving data.
[0338] S1006, The second terminal device displays the driving video media stream corresponding to the first digital human.
[0339] Through the above, the digital human service between the first terminal device and the second terminal device can be successfully executed.
[0340] Figure 10 illustrates information interaction between a first terminal device and a second terminal device via SIP messages. However, it is not limited to the first terminal device interacting with the second terminal device via a DC (Distributed Data Center). For example, the first terminal device can send a digital human call request message to the second terminal device via the DC, which includes the aforementioned first and second information. The second terminal device can then send a digital human call response message to the first terminal device via the DC, which includes the aforementioned third information.
[0341] As shown in Figures 5 to 10, embodiments of this application can support the interaction between the first device and the second device regarding the driving mode of the first digital human. The second device can determine the first driving mode based on this information. This can support the negotiation between the first device and the second device regarding the driving mode of the first digital human, ensuring that the driving mode of the first digital human adopted by the second device is the driving mode expected or supported by the first device, thereby enabling the successful execution of the digital human service initiated by the first device.
[0342] The content shown in Figures 6 to 10 is an example of how the first device achieves information interaction with the second device. However, the content shown in Figures 6 to 10 is only an example and is not a final limitation.
[0343] Finally, the device embodiments of this application will be described.
[0344] To achieve the functions of the methods provided in this application, the aforementioned devices (such as the first device, the second device, or the third device) may include hardware structures and / or software modules, implementing the aforementioned functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular function is executed in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.
[0345] Figure 11 is a schematic block diagram of a communication device 1100 according to an embodiment of this application. The communication device 1100 includes a processing circuit 1110 and a transceiver circuit 1120, which can be interconnected or coupled to each other, for example, through a bus 1130. This communication device can be a first device, a second device, or a third device.
[0346] Optionally, the communication device 1100 may also include a memory 1140. The memory 1140 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), or compact disc read-only memory (CD-ROM), which is used for related instructions and data.
[0347] The processing circuit 1110 can be all or part of the processing circuitry in one or more processors, or it can be one or more processors. The processor can be a central processing unit (CPU). If the processing circuit 1110 is a CPU, the CPU can be a single-core CPU or a multi-core CPU.
[0348] The processing circuit 1110 may be a signal processor, a chip, or other integrated circuit that can implement the method of this application, or a portion of the aforementioned processor, chip, or integrated circuit used for processing functions.
[0349] The transceiver circuit 1120 can also be a transceiver, or an input / output interface. The input / output interface is used for the input or output of signals or data, and can also be called an input / output circuit.
[0350] When the communication device 1100 is the first device, the processing circuit 1110 is used to perform the following operations: send first information and second information to the second device; receive third information from the second device.
[0351] When the communication device 1100 is the second device, the processing circuit 1110 is used to perform the following operations: receive first information and second information from the first device; send third information to the first device, etc.
[0352] When the communication device 1100 is a third device, the processing circuit 1110 is used to perform the following operations: receive a request message from the second device; send a response message to the second device, etc.
[0353] The above description is for illustrative purposes only.
[0354] When the communication device 1100 is a first device, a second device, or a third device, it will be responsible for executing the methods or steps related to the first device, the second device, or the third device in the foregoing method embodiments.
[0355] When the communication device 1100 is a first device, a second device, or a third device, the transceiver circuit 1120 can be a transceiver.
[0356] When the communication device 1100 is a chip used in the first device, the second device, or the third device, the transceiver circuit 1120 can be an input / output circuit.
[0357] For details, please refer to the content shown in the above method embodiments.
[0358] The implementation of each operation in Figure 11 can also be described in the corresponding description of the method embodiments shown in Figures 5 to 10.
[0359] Figure 12 is a schematic block diagram of a communication device 1200 according to an embodiment of this application. The communication device 1200 can be a first device, a second device, or a third device, used to implement the methods involved in the above embodiments.
[0360] The communication device 1200 includes a transceiver unit 1210 and a processing unit 1220. The transceiver unit 1210 may include a sending unit and a receiving unit. The sending unit is used to perform the sending action of the communication device, and the receiving unit is used to perform the receiving action of the communication device. For ease of description, the sending unit and the receiving unit are combined into one transceiver unit in this embodiment. This will be explained uniformly here and will not be repeated later.
[0361] When the communication device 1200 is the first device, exemplarily, the transceiver unit 1210 is used to: send first information and second information to the second device, and receive third information from the second device; the processing unit 1220 is used to determine the first driving data or the first information, etc.
[0362] When the communication device 1200 is a second device, exemplarily, the transceiver unit 1210 is used to receive first information and second information from the first device, and to send third information to the first device; the processing unit 1220 is used to determine the third information, etc.
[0363] When the communication device 1200 is a third device, for example, the transceiver unit 1210 is used to receive request messages from the second device and to send response messages to the second device; the processing unit 1220 is used to drive the first digital human, etc.
[0364] When the communication device 1200 is a first device, a second device, or a third device, it will be responsible for executing one or more of the methods or steps related to the first device, the second device, or the third device in the aforementioned method embodiments.
[0365] Optionally, the communication device 1200 further includes a storage unit 1230 for storing programs or code for executing the aforementioned methods.
[0366] The transceiver unit in Figure 12 can correspond to the transceiver circuit in Figure 11, and the processing unit in Figure 12 can correspond to the processing circuit in Figure 11.
[0367] The apparatus embodiments shown in Figures 11 and 12 are used to implement the contents described in Figures 5 to 10. The specific execution steps and methods of the apparatus shown in Figures 11 and 12 can be found in the foregoing method embodiments.
[0368] This application also provides a chip, including a processor, for calling and executing instructions stored in a memory, causing a communication device on which the chip is installed to perform the methods described in the examples above. The memory may be integrated within the chip or located externally.
[0369] This application also provides another chip, including: an input interface, an output interface, and a processing circuit. The input interface, the output interface, and the processor are connected via an internal connection path. The processing circuit is used to execute code in a memory. When the code is executed, the processing circuit is used to execute the methods in the examples described above. Optionally, the chip also includes a memory for storing computer programs or code. The input interface and the output interface can be independent of each other, or they can be integrated into a single input / output interface.
[0370] The processing circuitry can be all or part of the processing circuitry in one or more processors, or one or more processors.
[0371] This application also provides a processor for coupling with a memory for performing the methods and functions of a network device or terminal device involved in any of the above embodiments.
[0372] In another embodiment of this application, a computer program product containing instructions is provided, which, when run on a computer, enables the implementation of the methods described in the foregoing embodiments.
[0373] This application also provides a computer program that, when run on a computer, enables the implementation of the methods described in the foregoing embodiments.
[0374] In another embodiment of this application, a computer-readable storage medium is provided, which stores a computer program that, when executed by a computer, implements the methods described in the foregoing embodiments.
[0375] It should be understood that in the embodiments of this application, the processor can be a central processing unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0376] It should also be understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced synchronous SDRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0377] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0378] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0379] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0380] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the above functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0381] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
Claims
1. A method of digital human communication, characterized in that, Applied to a second device, comprising: Receive first information and second information from a first device, wherein the first information is used to identify a first digital human and the second information is used to indicate capabilities related to driving data, wherein the driving data is used to drive the first digital human; Based on the second information, a third information is sent to the first device, the third information being used to indicate a first driving mode, the first driving mode being a mode for driving the first digital human.
2. The method of claim 1, wherein, Sending third information to the first device based on the second information includes: Based on the second information and the driving mode of the first digital human supported by the second device, the third information is sent to the first device.
3. The method of claim 1, wherein, Sending third information to the first device based on the second information includes: Based on the second information and the driving mode of the first digital human supported by the fourth device, the third information is sent to the first device, wherein the fourth device is a device for driving the first digital human.
4. The method of claim 3, wherein, The fourth device is the second device, and the method further includes: Receive first drive data from the first device, wherein the first drive data is determined according to the first drive mode; The first digital human is driven according to the first driving data.
5. The method of claim 3, wherein, The fourth device is the third device, and the method further includes: Send a request message to the third device, the request message including the first information and the third information, the request message being used to request the driving resources of the first digital human; Receive a response from the third device to the request message.
6. The method of claim 5, wherein, The request message also includes fourth information, which indicates the type of the first digital human.
7. The method according to any one of claims 1 to 6, characterized in that, The capabilities related to driving data include at least one of the following: The data type of the driver data, the driving method associated with the driver data, or the driving technology associated with the driver data.
8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Receive fourth information, the fourth information being used to indicate the type of the first digital human; The step of sending third information to the first device based on the second information includes: Based on the second information and the fourth information, the third information is sent to the first device.
9. The method according to any one of claims 1 to 8, characterized in that, The first driving method includes any one of the following: Text-driven, voice-driven, video-driven, or facial feature-driven.
10. The method according to any one of claims 1 to 9, characterized in that, The second device includes at least one of the following: Overhead server, multimedia telephony application server, data channel application server, or terminal device; The first device is a terminal device.
11. The method according to any one of claims 1 to 10, characterized in that, The identifier of the first digital human is also used to indicate the type of the first digital human.
12. A method for digital human communication, characterized in that, Applied to the first device, comprising: Send first information and second information to the second device, wherein the first information is used to identify the first digital human and the second information is used to indicate capabilities related to driving data for the second device to determine the driving mode of the first digital human, and the driving data is used to drive the first digital human. Receive third information from the second device, the third information being used to indicate a first driving mode, the first driving mode being a mode for driving the first digital human.
13. The method according to claim 12, characterized in that, The method further includes: Send first driving data, which is determined according to the first driving method.
14. The method according to claim 12 or 13, characterized in that, The capabilities related to driving data include at least one of the following: The data type of the driver data, the driving method associated with the driver data, or the driving technology associated with the driver data.
15. The method according to any one of claims 12 to 14, characterized in that, The first driving method includes any one of the following: Text-driven, voice-driven, video-driven, or facial feature-driven.
16. The method according to any one of claims 13 to 15, characterized in that, The method further includes: Send a fourth message, which indicates the type of the first digital human.
17. The method according to any one of claims 13 to 16, characterized in that, The determination of the first driving mode is related to the second information and the driving mode of the first digital human supported by the fourth device, wherein the fourth device is a device for driving the first digital human.
18. The method according to any one of claims 13 to 17, characterized in that, The second device includes at least one of the following: Overhead server, multimedia telephony application server, data channel application server, or terminal device; The first device is a terminal device.
19. The method according to any one of claims 12 to 18, characterized in that, The identifier of the first digital human is also used to indicate the type of the first digital human.
20. A method for digital human communication, characterized in that, Applied to a third device, including: A request message is received from a second device. The request message includes first information and third information. The first information is used to identify a first digital human. The third information is used to indicate a first driving mode. The first driving mode is a mode used to drive the first digital human. The determination of the first driving mode is related to the second information. The second information is used to indicate capabilities related to driving data. The driving data is used to drive the first digital human. The request message is used to request driving resources for the first digital human. Send a response to the request message to the second device.
21. The method according to claim 20, characterized in that, The method further includes: Receive first drive data from the first device, wherein the first drive data is determined according to the first drive mode; The first digital human is driven according to the first driving data.
22. The method according to claim 20 or 21, characterized in that, The capabilities related to driving data include at least one of the following: The data type of the driver data, the driving method associated with the driver data, or the driving technology associated with the driver data.
23. The method according to any one of claims 20 to 22, characterized in that, The first driving method includes any one of the following: Text-driven, voice-driven, video-driven, or facial feature-driven.
24. The method according to any one of claims 20 to 23, characterized in that, The determination of the first driving mode is related to the second information and the driving mode of the first digital human supported by the third device.
25. The method according to any one of claims 20 to 24, characterized in that, The request message also includes fourth information, which indicates the type of the first digital human.
26. The method according to any one of claims 20 to 25, characterized in that, The second device includes at least one of the following: Overhead server, multimedia telephony application server, data channel application server, or terminal device; The first device is a terminal device.
27. The method according to any one of claims 20 to 26, characterized in that, The identifier of the first digital human is also used to indicate the type of the first digital human.
28. A method for digital human communication, characterized in that, include: The first device sends first information and second information to the second device. The first information is used to identify the first digital human, and the second information is used to indicate the capabilities related to the driving data so that the second device can determine the driving mode of the first digital human. The driving data is used to drive the first digital human. The second device sends third information to the first device based on the second information. The third information is used to indicate a first driving mode, which is a mode for driving the first digital human.
29. A communication system, characterized in that, include: A first device is configured to send first information and second information to a second device, wherein the first information is used to identify a first digital human, and the second information is used to indicate capabilities related to driving data for the second device to determine the driving mode of the first digital human, and the driving data is used to drive the first digital human. The second device is configured to receive the first information and the second information, and, based on the second information, send third information to the first device, the third information being used to indicate a first driving mode, the first driving mode being a mode for driving the first digital human; The first device is used to receive the third information.
30. A communication device, characterized in that, Includes a processor, the processor being configured to cause the communication device to perform the method of any one of claims 1 to 28 by executing a computer program or instructions, or by using logic circuitry.
31. The communication device according to claim 30, characterized in that, The communication device further includes a memory for storing the computer program or instructions.
32. The communication device according to claim 30 or 31, characterized in that, The communication device further includes a communication interface for inputting and / or outputting signals.
33. A communication device, characterized in that, It includes logic circuitry and input / output interfaces, the input / output interfaces being used to input and / or output signals, and the logic circuitry being used to perform the method of any one of claims 1 to 28.
34. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program or instructions that, when executed on a computer, cause the method of any one of claims 1 to 28 to be performed.
35. A computer program product, characterized in that, It includes instructions that, when run on a computer, cause the method of any one of claims 1 to 28 to be performed.
Citation Information
Patent Citations
Interaction method and device, electronic equipment and storage medium
CN113050791A
Live broadcast method and apparatus, and electronic device
CN114245155A
Multi-modal interaction method and device
CN114995636A
Remote digital human driving method, device and system
CN116301382A
Mobile user borne brain activity data and surrounding environment data correlation system
US20170061034A1