Digital object presentation method, apparatus, device, medium, and program product

CN122597597APending Publication Date: 2026-08-18JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510152144.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]各个模态下的对象内容不能有效地进行融合,数字对象在表现形式上常出现口齿不清、动作僵硬、情感表现欠佳等问题,使得数字对象展示效果不协调

Benefits of technology

[0025]The above embodiments of this disclosure have the following beneficial effects: Through the digital object display method of some embodiments of this disclosure, digital object display information can be accurately generated through feature semantic alignment processing, so as to display it accurately and vividly in the form of digital objects. Specifically, the reason why the related digital object display effect is not accurate and vivid is that the object content under different modalities cannot be effectively integrated. Digital objects often have problems such as unclear speech, stiff movements, and poor emotional expression in their expression, resulting in an uncoordinated digital object display effect. Based on this, the digital object display method of some embodiments of this disclosure first obtains multimodal digital object driving information for the target object to obtain the object content of the target object in multiple modalities, so that the subsequent digital object display method is more in line with the behavior of the target object and the display effect is better. Then, a multimodal feature information set corresponding to the above multimodal digital object driving information is generated. Among them, the multimodal feature information in the above multimodal feature information set corresponds to the modality in the multimodality. Here, by generating a multimodal feature information set, it is convenient to extract the semantic content of the target object in different modalities, thereby facilitating the alignment of the semantic content and improving the consistency of content across modalities. Next, feature semantic alignment processing is performed on the aforementioned multimodal feature information set to generate an aligned multimodal feature information set. This semantic alignment process ensures content alignment across modalities, preventing misalignment and mismatch, and further guaranteeing the accurate display of subsequent digital objects, avoiding various obvious problems in presentation. Furthermore, based on the aligned multimodal feature information set, accurate digital object display information can be generated, enabling the corresponding digital object display of the target object. Finally, based on the digital object display information, digital object display is performed on the target object to generate and accurately display it in the form of a digital object. In summary, by performing feature semantic alignment processing on the feature information of each modality, we can avoid content misalignment and mismatch in different modalities, further ensuring the accurate display effect of subsequent digital objects, avoiding various obvious problems in the presentation, and making the subsequent display effect of digital objects better.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597597A_ABST
    Figure CN122597597A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a digital object display method, device, equipment, medium and program product. A specific embodiment of the method comprises: obtaining multi-modal digital object driving information for a target object; generating a multi-modal feature information set corresponding to the multi-modal digital object driving information, wherein the multi-modal feature information in the multi-modal feature information set corresponds to a modality in the multi-modal; performing feature semantic alignment processing on the multi-modal feature information set to generate an aligned multi-modal feature information set; generating digital object display information according to the aligned multi-modal feature information set; and performing digital object display for the target object according to the digital object display information. The embodiment is related to a digital object, and through feature semantic alignment processing, the digital object display information can be accurately generated to accurately and vividly display in the form of a digital object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of this disclosure relate to the field of computer technology, and more specifically to digital object display methods, apparatuses, devices, media, and program products. Background Technology

[0002] Currently, with the development of technologies such as artificial intelligence and virtual reality, digital humans are frequently appearing in the public eye, entering people's lives in different identities, marking the presentation of digital humans' functions in social practice. The typical approach to displaying digital objects involves stages such as fixing the object's IP address, driver generation, and post-processing to generate and display the digital object.

[0003] However, the inventors discovered that when using the above method to generate and display digital objects, the following technical problems often arise:

[0004] The content of objects in different modalities cannot be effectively integrated. Digital objects often exhibit problems such as unclear speech, stiff movements, and poor emotional expression, resulting in inconsistent display effects.

[0005] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0007] Some embodiments of this disclosure provide digital object display methods, apparatuses, devices, media, and program products to address the technical problems mentioned in the background section above.

[0008] In a first aspect, some embodiments of this disclosure provide a digital object display method, including: acquiring multimodal digital object driving information for a target object; generating a multimodal feature information set corresponding to the multimodal digital object driving information, wherein the multimodal feature information in the multimodal feature information set corresponds to the modality in the multimodality; performing feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set; generating digital object display information based on the aligned multimodal feature information set; and performing digital object display for the target object based on the digital object display information.

[0009] Optionally, the aforementioned multimodal digital object driving information includes: first digital object driving information in image mode, second digital object driving information in audio mode, third digital object driving information in text mode, and object optical flow information corresponding to the target object; and the generation of the multimodal feature information set corresponding to the aforementioned multimodal digital object driving information includes: generating image mode feature information corresponding to the aforementioned first digital object driving information; generating audio mode feature information corresponding to the aforementioned second digital object driving information; generating text mode feature information corresponding to the aforementioned third digital object driving information; generating optical flow feature information corresponding to the aforementioned object optical flow information; and determining the aforementioned image mode feature information, the aforementioned audio mode feature information, the aforementioned text mode feature information, and the aforementioned optical flow feature information as the aforementioned multimodal feature information set.

[0010] Optionally, the above-mentioned feature semantic alignment processing for the multimodal feature information set to generate an aligned multimodal feature information set includes: performing feature semantic alignment processing for the audio modal feature information, the text modal feature information, and the optical flow feature information to generate an aligned feature information set; and determining the aligned feature information set and the image modal feature information as the aligned multimodal feature information set.

[0011] Optionally, the above-mentioned feature semantic alignment processing for the audio modal feature information, the text modal feature information, and the optical flow feature information to generate an aligned feature information set includes: adding noise to the optical flow feature information to obtain noisy feature information; and using a pre-trained feature semantic alignment model to perform feature semantic alignment processing on the audio modal feature information, the text modal feature information, and the noisy feature information to generate the aligned feature information set.

[0012] Optionally, the above-mentioned use of a pre-trained feature semantic alignment model to perform feature semantic alignment processing on the audio modal feature information, the text modal feature information, and the noisy feature information to generate the alignment feature information set includes: obtaining display effect control information for the target object; and inputting the display effect control information, the audio modal feature information, the text modal feature information, and the noisy feature information into the feature semantic alignment model to generate the alignment feature information set.

[0013] Optionally, generating digital object display information based on the aligned multimodal feature information set includes: fusing the feature information of each multimodal feature information in the aligned multimodal feature information set to generate fused feature information; and inputting the fused feature information into a pre-trained feature decoding model to generate a sequence of digital object display image frames as the digital object display information.

[0014] Optionally, the above method further includes: generating digital object audio information for the target object based on the multimodal feature information set; and performing digital object display for the target object based on the digital object display information and the digital object audio information.

[0015] Secondly, some embodiments of this disclosure provide a digital object display device, including: an acquisition unit configured to acquire multimodal digital object driving information for a target object; a first generation unit configured to generate a multimodal feature information set corresponding to the multimodal digital object driving information, wherein the multimodal feature information in the multimodal feature information set corresponds to the modality in the multimodality; a first execution unit configured to perform feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set; a second generation unit configured to generate digital object display information based on the aligned multimodal feature information set; and a second execution unit configured to execute digital object display for the target object based on the digital object display information.

[0016] Optionally, the aforementioned multimodal digital object driving information includes: first digital object driving information in image mode, second digital object driving information in audio mode, third digital object driving information in text mode, and object optical flow information corresponding to the target object; and the first generation unit can be configured to: generate image modal feature information corresponding to the aforementioned first digital object driving information; generate audio modal feature information corresponding to the aforementioned second digital object driving information; generate text modal feature information corresponding to the aforementioned third digital object driving information; generate optical flow feature information corresponding to the aforementioned object optical flow information; and determine the aforementioned image modal feature information, the aforementioned audio modal feature information, the aforementioned text modal feature information, and the aforementioned optical flow feature information as the aforementioned multimodal feature information set.

[0017] Optionally, the first execution unit may be configured to: perform feature semantic alignment processing on the audio modal feature information, the text modal feature information and the optical flow feature information to generate an aligned feature information set; and determine the aligned feature information set and the image modal feature information as an aligned multimodal feature information set.

[0018] Optionally, the first execution unit can be configured to: add noise to the optical flow feature information to obtain noisy feature information; and use a pre-trained feature semantic alignment model to perform feature semantic alignment on the audio modal feature information, the text modal feature information, and the noisy feature information to generate the alignment feature information set.

[0019] Optionally, the first execution unit can be configured to: acquire display effect control information for the target object; input the display effect control information, the audio modal feature information, the text modal feature information, and the noise-added feature information into the feature semantic alignment model to generate the alignment feature information set.

[0020] Optionally, the second generation unit can be configured to: fuse the multimodal feature information in the above-aligned multimodal feature information set to generate fused feature information; and input the fused feature information into a pre-trained feature decoding model to generate a sequence of digital object display image frames as the digital object display information.

[0021] Optionally, the apparatus further includes: generating digital object audio information for the target object based on the multimodal feature information set; and performing digital object display for the target object based on the digital object display information and the digital object audio information.

[0022] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, such that when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any implementation of the first aspect.

[0023] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0024] Fifthly, some embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0025] The above embodiments of this disclosure have the following beneficial effects: Through the digital object display method of some embodiments of this disclosure, digital object display information can be accurately generated through feature semantic alignment processing, so as to display it accurately and vividly in the form of digital objects. Specifically, the reason why the related digital object display effect is not accurate and vivid is that the object content under different modalities cannot be effectively integrated. Digital objects often have problems such as unclear speech, stiff movements, and poor emotional expression in their expression, resulting in an uncoordinated digital object display effect. Based on this, the digital object display method of some embodiments of this disclosure first obtains multimodal digital object driving information for the target object to obtain the object content of the target object in multiple modalities, so that the subsequent digital object display method is more in line with the behavior of the target object and the display effect is better. Then, a multimodal feature information set corresponding to the above multimodal digital object driving information is generated. Among them, the multimodal feature information in the above multimodal feature information set corresponds to the modality in the multimodality. Here, by generating a multimodal feature information set, it is convenient to extract the semantic content of the target object in different modalities, thereby facilitating the alignment of the semantic content and improving the consistency of content across modalities. Next, feature semantic alignment processing is performed on the aforementioned multimodal feature information set to generate an aligned multimodal feature information set. This semantic alignment process ensures content alignment across modalities, preventing misalignment and mismatch, and further guaranteeing the accurate display of subsequent digital objects, avoiding various obvious problems in presentation. Furthermore, based on the aligned multimodal feature information set, accurate digital object display information can be generated, enabling the corresponding digital object display of the target object. Finally, based on the digital object display information, digital object display is performed on the target object to generate and accurately display it in the form of a digital object. In summary, by performing feature semantic alignment processing on the feature information of each modality, we can avoid content misalignment and mismatch in different modalities, further ensuring the accurate display effect of subsequent digital objects, avoiding various obvious problems in the presentation, and making the subsequent display effect of digital objects better. Attached Figure Description

[0026] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0027] Figure 1 This is a schematic diagram of an application scenario of a digital object display method according to some embodiments of the present disclosure;

[0028] Figure 2 This is a flowchart of some embodiments of the digital object display method according to the present disclosure;

[0029] Figure 3 This is a schematic diagram illustrating the process of generating digital object display information according to some embodiments of the digital object display method of this disclosure;

[0030] Figure 4 These are flowcharts of other embodiments of the digital object display method according to this disclosure;

[0031] Figure 5 These are schematic diagrams illustrating the structure of some embodiments of the digital object display device according to this disclosure;

[0032] Figure 6 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0033] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0034] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0035] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0036] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0037] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0038] Before performing any of the operations involving the collection, storage, and use of user personal information (such as multimodal digital object-driven information) disclosed in this disclosure, the relevant organizations or individuals shall fulfill their obligations, including conducting personal information security impact assessments, informing personal information subjects, and obtaining prior authorization and consent from personal information subjects.

[0039] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0040] Figure 1 This is a schematic diagram of an application scenario of a digital object display method according to some embodiments of the present disclosure.

[0041] exist Figure 1 In this application scenario, firstly, the electronic device 101 can acquire multimodal digital object driving information for the target object. In this application scenario, the multimodal digital object driving information may include: object image 102 and audio 103. Then, the electronic device 101 can generate a multimodal feature information set 104 corresponding to the aforementioned multimodal digital object driving information. The multimodal feature information in the multimodal feature information set 104 corresponds to the modalities in the multimodal model. In this application scenario, the multimodal feature information set 104 includes: multimodal feature information 1041 corresponding to the image modality and multimodal feature information 1042 corresponding to the audio modality. Next, the electronic device 101 can perform feature semantic alignment processing on the aforementioned multimodal feature information set 104 to generate an aligned multimodal feature information set 105. In this application scenario, the aligned multimodal feature information set 105 may include: aligned multimodal feature information 1051 corresponding to multimodal feature information 1041, and aligned multimodal feature information 1052 corresponding to multimodal feature information 1042. Furthermore, the electronic device 101 can generate digital object display information 106 based on the aligned multimodal feature information set 105. Finally, the electronic device 101 can perform digital object display for the target object based on the digital object display information 106.

[0042] It should be noted that the aforementioned electronic device 101 can be either hardware or software. When the electronic device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the electronic device is software, it can be installed in the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here.

[0043] It should be understood that Figure 1 The number of electronic devices shown is merely illustrative. Any number of electronic devices can be used depending on the implementation requirements.

[0044] Continue to refer to Figure 2 The flowchart 200 illustrates some embodiments of a digital object display method according to the present disclosure. The digital object display method includes the following steps:

[0045] Step 201: Obtain multimodal digital object driving information for the target object.

[0046] In some embodiments, the execution entity of the above-described digital object display method (e.g.) Figure 1 The electronic device 101 shown can acquire multimodal digital object driving information for a target object via a wired or wireless connection. The target object can be an object to be displayed as a digital identity. In practice, the target object can be a user to be displayed as a digital human. The multimodal digital object driving information can be object information used to drive the display of digital objects in multiple modalities. That is, the multimodal digital object driving information can be the data foundation for the display of digital objects. For example, the multimodal digital object driving information can include: object information corresponding to the image modality, object information corresponding to the video modality, and biometric information corresponding to the target object in the biometric modality.

[0047] Step 202: Generate the multimodal feature information set corresponding to the multimodal digital object driving information mentioned above.

[0048] In some embodiments, the executing entity can generate a multimodal feature information set corresponding to the multimodal digital object driving information. The multimodal feature information in the multimodal feature information set corresponds to a modality within the multimodal context. For example, for multimodal contexts including image modality, video modality, and biometric modality, the corresponding multimodal feature information set may include: multimodal feature information corresponding to the image modality, multimodal feature information corresponding to the video modality, and multimodal feature information corresponding to the biometric modality. The multimodal feature information can characterize the semantic content of the object information within the corresponding modality. The multimodal feature information can be in vector form.

[0049] As an example, the aforementioned execution entity can utilize the corresponding multimodal feature extraction model to extract the semantic content of object information in each modality within the multimodal digital object-driven information, thereby generating multimodal feature information and obtaining a multimodal feature information set. There is a one-to-one correspondence between the modality feature extraction model and the modality in the multimodal model. The modality feature extraction model can be a neural network model that extracts the semantic content of features in the corresponding modality. The network structure of the modality feature extraction model can vary depending on the corresponding modality.

[0050] In some optional implementations of certain embodiments, the aforementioned multimodal digital object driving information includes: first digital object driving information in image mode, second digital object driving information in audio mode, third digital object driving information in text mode, and object optical flow information corresponding to the target object. The first digital object driving information can be object information in image form for digital object display. In practice, the first digital object driving information can be a photograph of the object corresponding to the digital object. For example, the first digital object driving information can be a photograph of the digital person corresponding to the digital person. The second digital object driving information can be object information in audio form for digital object display. In practice, the second digital object driving information can be audio-related information that the digital object is expected to display. For example, audio-related information can include: audio timbre, audio frequency domain information, and audio volume. The third digital object driving information can be target-oriented information in text form for digital object display. In practice, the third digital object driving information can characterize the speech output by the target object during digital object display. For example, the third digital object driving information can be story text. The object optical flow information can be motion information representing the object corresponding to the target object during digital object display. Optical flow information can be in matrix form, representing the motion of the target object. It can include a sequence of image frames representing the motion of the target object.

[0051] Optionally, the aforementioned executing entity may generate a multimodal feature information set corresponding to the aforementioned multimodal digital object driving information, including the following steps:

[0052] The first step is to generate image modal feature information corresponding to the aforementioned first digital object driving information. This image modal feature information can be semantic content representing the image modality within the first digital object driving information. The image modal feature information can be in vector form.

[0053] As an example, the aforementioned execution entity can input the first digital object-driven information into a pre-trained image feature extraction model to generate image modal feature information. In practice, the image feature extraction model can be a multi-layered convolutional layer.

[0054] The second step is to generate the audio modal feature information corresponding to the aforementioned second digital object driving information. The audio modal feature information can be semantic content representing the audio modality within the second digital object driving information. The audio modal feature information can be in vector form.

[0055] As an example, the aforementioned execution entity can input the second digital object-driven information into a pre-trained audio feature extraction model to generate audio modal feature information. In practice, the audio feature extraction model can be a Mel Frequency Cepstrum Coefficient (MFCC) feature extraction model.

[0056] The third step is to generate the text modality feature information corresponding to the aforementioned third digital object-driven information. This text modality feature information can be the semantic content representing the text modality within the third digital object-driven information. The text modality feature information can be in vector form.

[0057] As an example, the aforementioned agent can input third-party digital object-driven information into a pre-trained text feature extraction model to generate text modal feature information. In practice, the text feature extraction model can be a recurrent neural network model.

[0058] The fourth step is to generate optical flow feature information corresponding to the aforementioned object optical flow information. This optical flow feature information can be semantic content representing the optical flow characteristics corresponding to the optical flow within the object's optical flow information. The optical flow feature information can be in vector form or in matrix sequence form.

[0059] As an example, the aforementioned execution entity can input object optical flow information into a pre-trained optical flow feature extraction model to generate optical flow feature information. In practice, the optical flow feature extraction model can be a Siamese network.

[0060] The fifth step is to determine the above-mentioned image modal feature information, audio modal feature information, text modal feature information and optical flow feature information as the above-mentioned multimodal feature information set.

[0061] Step 203: Perform feature semantic alignment processing on the above multimodal feature information set to generate an aligned multimodal feature information set.

[0062] In some embodiments, the execution entity may perform feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set.

[0063] As an example, the aforementioned execution entity can utilize the LLaVA (Large Language and Vision Assistant) model to perform feature semantic alignment processing on the aforementioned multimodal feature information set, in order to generate an aligned multimodal feature information set.

[0064] In some optional implementations of certain embodiments, the execution entity may perform feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set, including the following steps:

[0065] The first step is to perform feature semantic alignment processing on the aforementioned audio modal feature information, text modal feature information, and optical flow feature information to generate an alignment feature information set.

[0066] As an example, the aforementioned execution entity can utilize the LLaVA model to perform feature semantic alignment processing on the aforementioned audio modal feature information, the aforementioned text modal feature information, and the aforementioned optical flow feature information to generate an aligned feature information set.

[0067] The second step is to determine the above-mentioned alignment feature information set and the above-mentioned image modal feature information as the aligned multimodal feature information set.

[0068] Optionally, the aforementioned execution entity may perform feature semantic alignment processing on the aforementioned audio modal feature information, the aforementioned text modal feature information, and the aforementioned optical flow feature information to generate an aligned feature information set, including the following steps:

[0069] The first step is to add noise to the above optical flow feature information to obtain the noise-added feature information.

[0070] As an example, the aforementioned execution entity can fuse optical flow feature information with noise-corresponding feature information (e.g., splice them together) to generate noisy feature information.

[0071] As another example, the aforementioned execution entity can add the noise-corresponding feature information to the matrix subsequence following the second frame matrix corresponding to the optical flow feature information to generate the noise-added feature information.

[0072] The second step involves using a pre-trained feature semantic alignment model to perform feature semantic alignment processing on the aforementioned audio modal feature information, text modal feature information, and noisy feature information to generate the aligned feature information set. In practice, the aligned feature information set includes: aligned feature information corresponding to the audio modal feature information, aligned feature information corresponding to the text modal feature information, and aligned feature information corresponding to the noisy feature information. The feature semantic alignment model can be a neural network model for aligning feature semantic content. The feature semantic alignment model can be a model including a multi-head attention mechanism. For example, the feature semantic alignment model can be a Transformer model. Here, the feature semantic alignment model can implement a neural network model for motion estimation based on the feature information of each modality corresponding to the target object.

[0073] As an example, the aforementioned execution entity can input the aforementioned audio modal feature information, the aforementioned text modal feature information, and the aforementioned noise-added feature information into the feature semantic alignment model to generate alignment feature information corresponding to the audio modal feature information, alignment feature information corresponding to the text modal feature information, and alignment feature information corresponding to the noise-added feature information, thereby obtaining an alignment feature information set.

[0074] Optionally, the aforementioned execution entity may utilize a pre-trained feature semantic alignment model to perform feature semantic alignment processing on the aforementioned audio modal feature information, the aforementioned text modal feature information, and the aforementioned noisy feature information to generate the aforementioned aligned feature information set, including the following steps:

[0075] The first step is to obtain the display effect control information for the aforementioned target audience. This display effect control information can characterize the video style to be presented during the digital object display process. Specifically, for a user as the target audience, the display effect control information can characterize the video style requested by the user during the digital human display. For example, the video style can be, but is not limited to, at least one of the following: a happy digital human style, a sad digital human style, a serious digital human style, or a hip-hop digital human style. The display effect control information can be pre-set.

[0076] Here, the display effect control information can achieve the decoupling of the digital object visually from the corresponding display screen of the target object during the display process, so as to adjust the display effect of the feature information under various modalities.

[0077] The second step involves inputting the aforementioned display effect control information, audio modal feature information, text modal feature information, and noise-added feature information into the aforementioned feature semantic alignment model to generate the aforementioned alignment feature information set.

[0078] Step 204: Generate digital object display information based on the above-aligned multimodal feature information set.

[0079] In some embodiments, the aforementioned executing entity can generate digital object display information based on the aligned multimodal feature information set. This digital object display information can be the content displayed when a digital object is displayed. In practice, the digital object display information can be the content displayed by a digital human. For example, the digital object display information can be multiple consecutive image frames displayed by a digital human.

[0080] As an example, the aforementioned execution entity can input the aligned multimodal feature information set into the image generation layer to generate digital object display information. For instance, the image generation layer could be a series of cascaded residual layers.

[0081] In some optional implementations of certain embodiments, the aforementioned execution entity can generate digital object display information based on the aligned multimodal feature information set, including the following steps:

[0082] The first step is to fuse the multimodal feature information from the aligned multimodal feature information set to generate fused feature information.

[0083] As an example, the aforementioned execution entity can concatenate the feature information of each multimodal feature information in the aligned multimodal feature information set to generate concatenated feature information, which can then be used as fused feature information.

[0084] The second step involves inputting the fused feature information into a pre-trained feature decoding model to generate a sequence of digital object display image frames, which serve as the digital object display information. The feature decoding model can be a decoding model corresponding to the feature encoding model for each modality. The model structure corresponding to the feature decoding model can be pre-set based on practical experience. The decoding model can be an upsampling model. For example, the feature decoding model can be a multi-layered convolutional layer.

[0085] Step 205: Based on the above digital object display information, perform digital object display for the above target object.

[0086] In some embodiments, the aforementioned execution entity may perform digital object display for the aforementioned target object based on the aforementioned digital object display information.

[0087] As an example, the aforementioned executing entity can perform digital object-oriented display of information content corresponding to digital object display information.

[0088] As an example, for a digital object that is a digital human, the information displayed by the digital object is a sequence of digital human image frames corresponding to the digital human, and the executing entity can display the digital human.

[0089] like Figure 3 As shown, this illustrates the process of generating information for digital objects.

[0090] First, the system acquires the first digital object driving information in the image modality, the second digital object driving information in the audio modality, the third digital object driving information in the text modality, and the object optical flow information corresponding to the target object. Then, it generates image modality feature information corresponding to the first digital object driving information, audio modality feature information corresponding to the second digital object driving information, text modality feature information corresponding to the third digital object driving information, and optical flow feature information corresponding to the object optical flow information. Next, it adds noise to the optical flow feature information to generate noisy feature information. Then, it inputs the audio modality feature information, text modality feature information, noisy feature information, and style information (i.e., display effect control information) into a Transformer model (i.e., a feature semantic alignment model) to generate an alignment feature information set. Finally, using a feature decoding model, it performs decoding processing on the alignment feature information set and the image modality feature information to generate digital object display information.

[0091] The above embodiments of this disclosure have the following beneficial effects: Through the digital object display method of some embodiments of this disclosure, digital object display information can be accurately generated through feature semantic alignment processing, so as to display it accurately and vividly in the form of digital objects. Specifically, the reason why the related digital object display effect is not accurate and vivid is that the object content under different modalities cannot be effectively integrated. Digital objects often have problems such as unclear speech, stiff movements, and poor emotional expression in their expression, resulting in an uncoordinated digital object display effect. Based on this, the digital object display method of some embodiments of this disclosure first obtains multimodal digital object driving information for the target object to obtain the object content of the target object in multiple modalities, so that the subsequent digital object display method is more in line with the behavior of the target object and the display effect is better. Then, a multimodal feature information set corresponding to the above multimodal digital object driving information is generated. Among them, the multimodal feature information in the above multimodal feature information set corresponds to the modality in the multimodality. Here, by generating a multimodal feature information set, it is convenient to extract the semantic content of the target object in different modalities, thereby facilitating the alignment of the semantic content and improving the consistency of content across modalities. Next, feature semantic alignment processing is performed on the aforementioned multimodal feature information set to generate an aligned multimodal feature information set. This semantic alignment process ensures content alignment across modalities, preventing misalignment and mismatch, and further guaranteeing the accurate display of subsequent digital objects, avoiding various obvious problems in presentation. Furthermore, based on the aligned multimodal feature information set, accurate digital object display information can be generated, enabling the corresponding digital object display of the target object. Finally, based on the digital object display information, digital object display is performed on the target object to generate and accurately display it in the form of a digital object. In summary, by performing feature semantic alignment processing on the feature information of each modality, we can avoid content misalignment and mismatch in different modalities, further ensuring the accurate display effect of subsequent digital objects, avoiding various obvious problems in the presentation, and making the subsequent display effect of digital objects better.

[0092] Further reference Figure 4 The diagram illustrates a flow 400 of another embodiment of a digital object display method according to the present disclosure. This digital object display method includes the following steps:

[0093] Step 401: Obtain multimodal digital object driving information for the target object.

[0094] Step 402: Generate the multimodal feature information set corresponding to the multimodal digital object driving information mentioned above.

[0095] Step 403: Perform feature semantic alignment processing on the above multimodal feature information set to generate an aligned multimodal feature information set.

[0096] Step 404: Generate digital object display information based on the above-aligned multimodal feature information set.

[0097] Step 405: Based on the above digital object display information, perform digital object display for the above target object.

[0098] In some embodiments, the specific implementation of steps 401-405 and the resulting technical effects can be found in [reference needed]. Figure 2 Steps 201-205 in the corresponding embodiments will not be repeated here.

[0099] Step 406: Based on the above multimodal feature information set, generate digital object audio information for the above target object.

[0100] In some embodiments, the executing entity (e.g. Figure 1 The electronic device 101 shown can generate digital object audio information for the target object based on the aforementioned multimodal feature information set. The digital object audio information can be the audio content displayed when the digital object is shown. In practice, the digital object audio information can be the audio content displayed by a digital human. For example, the digital object audio information can be the audio displayed by a digital human.

[0101] As an example, firstly, the text audio of the target object is acquired based on the third digital object driving information. Then, the audio feature information corresponding to the text audio is extracted. Next, the image modality feature information is removed from the multimodal feature information set to obtain a set of modal feature information after removal. Then, the modal feature information after removal, the audio feature information, and the noisy feature information are input into a feature semantic alignment model to generate a first audio alignment feature information set. Furthermore, the various first audio alignment feature information sets in the first audio alignment feature information set are fused to generate a first audio fused feature information. Finally, the first audio fused feature information is input into an audio decoding model to generate digital object audio information. The audio decoding model can be a multi-layer cascaded LSTM model.

[0102] As another example, firstly, the aforementioned execution entity can acquire the text audio of the target object's third digital object driving information. Then, it extracts the audio feature information corresponding to the text audio. Next, the audio feature information, the aforementioned multimodal feature information set, and the forged feature information are input into a feature semantic alignment model to generate a second audio alignment feature information set. Furthermore, the various second audio alignment feature information sets in the second audio alignment feature information set are fused to generate a second audio fusion feature information. Finally, the second audio fusion feature information is input into an audio decoding model to generate digital object audio information. The audio decoding model can be a multi-layer cascaded LSTM model.

[0103] Step 407: Based on the above digital object display information and the above digital object audio information, perform digital object display for the above target object.

[0104] In some embodiments, the execution entity may perform digital object display for the target object based on the digital object display information and the digital object audio information.

[0105] As an example, the aforementioned execution entity can use the image frame sequence corresponding to the digital object display information as the display screen of the digital object, and the audio corresponding to the digital object audio information as the output audio, to execute the digital object display for the aforementioned target object.

[0106] from Figure 4 It can be seen from this that, with Figure 2 Compared to the description of some corresponding embodiments, Figure 4 In some corresponding embodiments, the process 400 of the digital object display method can accurately generate digital object audio information by using a multimodal feature information set and feature alignment processing, so as to combine digital object display information to display digital objects in a multimodal manner.

[0107] Further reference Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a digital object display device, which are similar to... Figure 2 Corresponding to the method embodiments shown, this digital object display device can be specifically applied to various electronic devices.

[0108] like Figure 5As shown, a digital object display device 500 includes: an acquisition unit 501, a first generation unit 502, a first execution unit 503, a second generation unit 504, and a second execution unit 505. The acquisition unit 501 is configured to acquire multimodal digital object driving information for a target object; the first generation unit 502 is configured to generate a multimodal feature information set corresponding to the multimodal digital object driving information, wherein the multimodal feature information in the multimodal feature information set corresponds to the modalities in the multimodality; the first execution unit 503 is configured to perform feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set; the second generation unit 504 is configured to generate digital object display information based on the aligned multimodal feature information set; and the second execution unit 505 is configured to execute digital object display for the target object based on the digital object display information.

[0109] In some optional implementations of certain embodiments, the aforementioned multimodal digital object driving information includes: first digital object driving information in image mode, second digital object driving information in audio mode, third digital object driving information in text mode, and object optical flow information corresponding to the target object; and the first generation unit 502 may be further configured to: generate image modal feature information corresponding to the aforementioned first digital object driving information; generate audio modal feature information corresponding to the aforementioned second digital object driving information; generate text modal feature information corresponding to the aforementioned third digital object driving information; generate optical flow feature information corresponding to the aforementioned object optical flow information; and determine the aforementioned image modal feature information, the aforementioned audio modal feature information, the aforementioned text modal feature information, and the aforementioned optical flow feature information as the aforementioned multimodal feature information set.

[0110] In some optional implementations of some embodiments, the first execution unit 503 may be further configured to: perform feature semantic alignment processing on the audio modal feature information, the text modal feature information and the optical flow feature information to generate an aligned feature information set; and determine the aligned feature information set and the image modal feature information as an aligned multimodal feature information set.

[0111] In some optional implementations of some embodiments, the first execution unit 503 may be further configured to: add noise to the optical flow feature information to obtain noisy feature information; and use a pre-trained feature semantic alignment model to perform feature semantic alignment processing on the audio modal feature information, the text modal feature information, and the noisy feature information to generate the alignment feature information set.

[0112] In some optional implementations of some embodiments, the first execution unit 503 may be further configured to: acquire display effect control information for the target object; input the display effect control information, the audio modal feature information, the text modal feature information and the noise-added feature information to the feature semantic alignment model to generate the alignment feature information set.

[0113] In some optional implementations of some embodiments, the second generation unit 504 may be further configured to: fuse the feature information of each multimodal feature information in the above-aligned multimodal feature information set to generate fused feature information; and input the above-mentioned fused feature information into a pre-trained feature decoding model to generate a digital object display image frame sequence as the above-mentioned digital object display information.

[0114] In some optional implementations of certain embodiments, the digital object display device 500 further includes a third generation unit and a third execution unit (not shown in the figure). The third generation unit can be configured to generate digital object audio information for the target object based on the aforementioned multimodal feature information set. The third execution unit can be configured to perform digital object display for the target object based on the aforementioned digital object display information and the aforementioned digital object audio information.

[0115] It is understandable that the units described in the digital object display device 500 are related to the reference. Figure 2 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the digital object display device 500 and the units contained therein, and will not be repeated here.

[0116] The following is for reference. Figure 6 It illustrates electronic devices suitable for implementing some embodiments of this disclosure (e.g., Figure 1 A schematic diagram of the structure of electronic device 101)600 in the middle. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0117] like Figure 6As shown, the electronic device 600 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory 602 or a program loaded from a storage device 608 into a random access memory 603. The random access memory 603 also stores various programs and data required for the operation of the electronic device 600. The processing unit 601, the read-only memory 602, and the random access memory 603 are interconnected via a bus 604. An input / output interface 605 is also connected to the bus 604.

[0118] Typically, the following devices can be connected to the input / output interface 605: input devices 606 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 608 including, for example, magnetic tape, hard disk, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 6 Each box shown can represent a device or multiple devices as needed.

[0119] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a read-only memory 602. When the computer program is executed by the processing device 601, it performs the functions defined above in the methods of some embodiments of this disclosure.

[0120] It should be noted that, in some embodiments of this disclosure, the computer-readable medium described above may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0121] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0122] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire multimodal digital object driving information for a target object; generate a multimodal feature information set corresponding to the multimodal digital object driving information, wherein the multimodal feature information in the multimodal feature information set corresponds to a modality in the multimodality; perform feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set; generate digital object display information based on the aligned multimodal feature information set; and perform digital object display for the target object based on the digital object display information.

[0123] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0125] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a first generation unit, a first execution unit, a second production unit, and a second execution unit. The names of these units do not necessarily limit the specific unit; for example, an acquisition unit may also be described as "a unit for acquiring multimodal digital object driving information for a target object."

[0126] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0127] Some embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the digital object display methods described above.

[0128] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A method for displaying digital objects, comprising: Obtain multimodal digital object driving information for the target object; Generate a multimodal feature information set corresponding to the multimodal digital object driving information, wherein the multimodal feature information in the multimodal feature information set corresponds to the mode in the multimodality; Perform feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set; Based on the aligned multimodal feature information set, digital object display information is generated; Based on the digital object display information, perform a digital object display for the target object.

2. The method according to claim 1, wherein, The multimodal digital object driving information includes: first digital object driving information in image mode, second digital object driving information in audio mode, third digital object driving information in text mode, and object optical flow information corresponding to the target object; and The generation of the multimodal feature information set corresponding to the multimodal digital object driving information includes: Generate image modal feature information corresponding to the first digital object driving information; Generate audio modal feature information corresponding to the second digital object driving information; Generate text modal feature information corresponding to the third digital object driving information; Generate optical flow feature information corresponding to the optical flow information of the object; The image modal feature information, the audio modal feature information, the text modal feature information, and the optical flow feature information are determined as the multimodal feature information set.

3. The method according to claim 2, wherein, The step of performing feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set includes: Perform feature semantic alignment processing on the audio modal feature information, the text modal feature information, and the optical flow feature information to generate an alignment feature information set; The alignment feature information set and the image modal feature information are determined as the aligned multimodal feature information set.

4. The method according to claim 3, wherein, The step of performing feature semantic alignment processing on the audio modal feature information, the text modal feature information, and the optical flow feature information to generate an alignment feature information set includes: The optical flow feature information is subjected to noise processing to obtain the noise-added feature information; Using a pre-trained feature semantic alignment model, feature semantic alignment processing is performed on the audio modal feature information, the text modal feature information, and the noisy feature information to generate the alignment feature information set.

5. The method according to claim 4, wherein, The step of using a pre-trained feature semantic alignment model to perform feature semantic alignment processing on the audio modal feature information, the text modal feature information, and the noisy feature information to generate the aligned feature information set includes: Obtain display effect control information for the target object; The display effect control information, the audio modal feature information, the text modal feature information, and the noise-added feature information are input into the feature semantic alignment model to generate the alignment feature information set.

6. The method according to claim 1, wherein, The step of generating digital object display information based on the aligned multimodal feature information set includes: The multimodal feature information in the aligned multimodal feature information set is fused to generate fused feature information; The fused feature information is input into a pre-trained feature decoding model to generate a sequence of digital object display image frames, which serve as the digital object display information.

7. The method according to claim 1, wherein, The method further includes: Based on the multimodal feature information set, digital object audio information for the target object is generated; Based on the digital object display information and the digital object audio information, perform digital object display for the target object.

8. A digital object display device, comprising: The acquisition unit is configured to acquire multimodal digital object driving information for a target object; The first generation unit is configured to generate a multimodal feature information set corresponding to the multimodal digital object driving information, wherein the multimodal feature information in the multimodal feature information set corresponds to the mode in the multimodality; The first execution unit is configured to perform feature semantic alignment processing on the multimodal feature information set to generate an aligned multimodal feature information set. The second generation unit is configured to generate digital object display information based on the aligned multimodal feature information set; The second execution unit is configured to perform a digital object display for the target object based on the digital object display information.

9. An electronic device, comprising: One or more processors; Storage device, on which one or more programs are stored, When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-7.

10. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-7.

11. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-7.