Digital human media stream arrangement method and device, equipment, storage medium and product
The method enhances digital human live streaming by precisely arranging multi-modal data through action-dragging and time-coordinate recording, improving efficiency and real-time performance.
Patent Information
- Application Number
- CN202510417371.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-15
AI Technical Summary
The existing digital human media stream orchestration method has inaccurate multimodal data orchestration and low data processing efficiency, which affects the real-time nature of digital human live broadcasts.
By introducing action drag and time coordinate recording functions in the multimodal editor, drag preset actions in text information according to user operations and record time coordinates and modal data type coordinates to generate digital human media streams, using asynchronous processing speech orchestration mechanism and efficient caching mechanism, and using Protobuf, WebSocket or HTTP/2 communication protocols to transmit data.
It realizes the precise orchestration of multimodal data, improves data processing efficiency and overall processing speed, and ensures the real-time and naturalness of digital live broadcasts.
Smart Images

Figure CN120321471A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a digital human media stream choreography method, device, equipment, storage medium and product. Background Art
[0002] The core idea of the multi-modal choreography mechanism for generating digital human media streams is that video materials are stored through an Object Storage Service (OSS) platform, and key data is persisted in a database.
[0003] In the existing technical solutions, although the multi-modal editor supports basic modal data association, it lacks detailed action dragging and time coordinate recording functions, resulting in inaccurate multi-modal data choreography; in addition, due to the lack of logic for matching actions and texts in the speech choreography system, the data processing efficiency is low, the overall processing speed is slow, and it is easy to affect the real-time performance of digital human live broadcasts. Summary of the Invention
[0004] Embodiments of this application provide a digital human media stream choreography method, device, equipment, storage medium and product, which are used to solve the technical problems of inaccurate multi-modal data choreography, low data processing efficiency, slow overall processing speed, and easy impact on the real-time performance of digital human live broadcasts in the existing digital human media stream choreography method.
[0005] In a first aspect, an embodiment of this application provides a digital human media stream choreography method, including: creating a digital human and determining the text information of the digital human; displaying the text information in a multi-modal editor according to the speaking speed of the digital human; the abscissa of the multi-modal editor represents the time axis, and the ordinate of the multi-modal editor represents the modal data type; in response to the user's dragging operation, dragging multiple preset actions to the lower part of the target text corresponding to each preset action respectively, and recording the time coordinate of each preset action and the modal data type coordinate of each preset action, so as to obtain the speech information and action information of the digital human; one preset action corresponds to one target text in the text information; generating a digital human media stream based on the speech information and action information.
[0006] In an embodiment, generating a digital human media stream based on the speech information and action information includes: obtaining metadata from an object storage service platform based on an asynchronous processing speech choreography mechanism; the metadata is the image video material metadata corresponding to the speech information and action information; converting the metadata into multiple target data with a preset format based on the speech information and action information; generating a digital human media stream based on the multiple target data.
[0007] In one embodiment, the metadata includes the action type of each preset action and the action ID of each preset action, the speech information includes the chronological order of each target word in the text information, and the action information includes the time coordinates of each preset action and the modal data type coordinates of each preset action; based on the speech information and the action information, the metadata is converted into multiple target data with a preset format, including: performing format conversion on the action type of each preset action and the action ID of each preset action to obtain multiple action data with a preset format; one action data is determined by the action type of one preset action and the action ID of one preset action; based on the time coordinates of each preset action, determining the text insertion position corresponding to each preset action; performing copywriting cutting processing on the text information based on the time coordinates of each preset action to obtain the target word corresponding to each preset action; based on the text insertion position corresponding to each preset action, respectively splicing the corresponding action data for each target word to obtain multiple target data with a preset format; one action data corresponds to one target word, and one target data is composed of one action data and one target word; storing each target data into the ZSet cache of the Redis database in sequence according to the chronological order of each target word; wherein, the preset format is any one of the VSML format, the JSON format, and the XML format.
[0008] In one embodiment, based on multiple target data, a digital human media stream is generated, including: sequentially obtaining each target data from the ZSet cache according to the chronological order of each target word; converting all target data into the form of a data stream based on the Protobuf protocol to generate a digital human media stream; transmitting the digital human media stream based on the target communication protocol; wherein, the target communication protocol is any one of the gRPC communication protocol, the WebSocket communication protocol, and the HTTP / 2 communication protocol.
[0009] In one embodiment, creating a digital human and determining the text information of the digital human includes: in response to a user's selection instruction, determining the preset image of the digital human and creating the digital human based on the preset image; adjusting the background image of the digital human and the position of the digital human in the background image based on the user interface; in response to the user's text input instruction, determining the text information of the digital human; wherein, the text information includes the image description copywriting and the product description copywriting of the digital human; the user interface is any one of a graphical interface, a command line tool, and an API interface.
[0010] Before creating a digital human and determining the text information of the digital human, it further includes: obtaining the metadata of the image video material; uploading the metadata of the image video material to the object storage service platform; storing the metadata of the image video material in the database.
[0011] Second aspect, an embodiment of the present application provides a digital human media stream choreographing device, including: a creation module, configured to create a digital human and determine the text information of the digital human; a display module, configured to display the text information in a multimodal editor according to the speech rate of the digital human; the abscissa of the multimodal editor represents the time axis, and the ordinate of the multimodal editor represents the modal data type; an editing module, configured to, in response to a user's dragging operation, drag multiple preset actions to the lower side of the target text corresponding to each preset action respectively, and record the time coordinate of each preset action and the modal data type coordinate of each preset action, to obtain the speech information and action information of the digital human; one preset action corresponds to one target text in the text information; a media stream generation module, configured to generate a digital human media stream based on the speech information and the action information.
[0012] Third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements any one of the above digital human media stream choreographing methods.
[0013] Fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the above digital human media stream choreographing methods.
[0014] Fifth aspect, an embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements any one of the above digital human media stream choreographing methods.
[0015] For the digital human media stream choreographing method, device, equipment, storage medium and product provided by the embodiments of the present application, after creating a digital human, the text information is displayed in a multimodal editor according to the speech rate of the digital human, and according to the user's dragging operation, multiple preset actions are respectively dragged to the lower side of the target text corresponding to each preset action, the time coordinate of each preset action and the modal data type coordinate of each preset action are recorded, to obtain the speech information and action information of the digital human, and then a digital human media stream is generated according to the speech information and the action information. Through the above method, the function of action dragging and time coordinate recording is introduced in the process of multimodal data choreographing. Through the dragging method, the precise pairing of each action and the corresponding text in the time coordinate is realized, providing the logic of action-text matching for the choreographing of multimodal data, which is beneficial to realizing the precise choreographing of multimodal data, improving the data processing efficiency and the overall processing speed, and further beneficial to ensuring the real-time performance of digital human live broadcast. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the present application or the prior art, the following will briefly introduce the drawings required for the embodiments or the prior art description. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is one of the schematic flowcharts of the digital human media stream choreography method provided by the embodiments of the present application.
[0018] Figure 2 It is the second of the schematic flowcharts of the digital human media stream choreography method provided by the embodiments of the present application.
[0019] Figure 3 It is the schematic structural diagram of the digital human media stream choreography device provided by the embodiments of the present application.
[0020] Figure 4 It is the schematic structural diagram of the electronic device provided by the embodiments of the present application. Specific Embodiments
[0021] To make the objectives, technical solutions, and advantages of the present application clearer, the following will clearly and completely describe the technical solutions in the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0022] Please refer to Figure 1 and Figure 2 , Figure 1 is one of the schematic flowcharts of the digital human media stream choreography method provided by the embodiments of the present application, Figure 2 is the second of the schematic flowcharts of the digital human media stream choreography method provided by the embodiments of the present application. As Figure 1 shown, in the embodiments of the present application, the digital human media stream choreography method includes steps S110 to S140, and the specific steps are as follows: S110: Create a digital human and determine the text information of the digital human.
[0023] As Figure 2 shown, before creating the digital human, it is necessary to initialize the materials. In the material initialization stage, the user can obtain the preset image video materials from the local or cloud storage by calling the offline interface, and the image video materials include image video material metadata.
[0024] Optionally, upload the avatar video material to the Object Storage Service (OSS) platform to ensure the security and efficient access of the avatar video material.
[0025] Optionally, persist the avatar video material metadata (such as the action ID, action type, and OSS address of the preset action, etc.) to the database to provide more comprehensive information support for subsequent processing and ensure the reliable storage of the data.
[0026] Further, in response to the user's selection instruction, select a preset avatar for the digital human to be created, and create the digital human according to the preset avatar.
[0027] Optionally, after creating the digital human, the user can, according to the needs, replace or adjust the background image of the digital human through the user interface, and adjust the position of the digital human in the background image to achieve a more personalized design and enhance the richness and attractiveness of the content.
[0028] Optionally, the user interface is any one of a graphical interface, a command-line tool, and an API interface.
[0029] Among them, by adjusting the background image or text content of the digital human through the graphical interface, the adjustment effect can be intuitively reflected to the user; by adjusting the background image or text content of the digital human through other user interfaces (such as command-line tools and API interfaces), although the adjustment effect is not as intuitive as the graphical interface, the same function can still be achieved.
[0030] Further, in response to the user's text input instruction, add the text information of the digital human; the text information includes the image description copywriting of the digital human, the global copywriting related to the digital human, and the product description copywriting to improve the richness and attractiveness of the digital human live broadcast content.
[0031] S120: Display the text information in the multimodal editor according to the speaking speed of the digital human.
[0032] As Figure 2 shown, after the creation of the digital human is completed, it enters the multimodal editing stage. In the multimodal editing stage, the user can display the text information in the multimodal editor according to the speaking speed of the digital human when reading the text information, ensuring the synchronization of the display of the text information with the voice, and improving the naturalness and interactivity of the digital human live broadcast.
[0033] In the interface of the multimodal editor, the abscissa of the multimodal editor represents the time axis, and the ordinate of the multimodal editor represents the modal data type (such as text, action, voice, etc.), which is conducive to the user to intuitively perform data association on the time axis.
[0034] S130: In response to the user's dragging operation, drag multiple preset actions to below the target text corresponding to each preset action respectively, and record the time coordinate of each preset action and the coordinate of the modal data type of each preset action, so as to obtain the speech information and action information of the digital human.
[0035] One preset action corresponds to one target text in the text information.
[0036] Specifically, in a digital human live broadcast, there may be multiple preset actions of the digital human; for each preset action, in response to the user's dragging operation, the preset action can be dragged to below the corresponding target text in the text information. At this time, the multi-modal editor system will automatically record the time coordinate and the coordinate of the modal data type of the preset action to ensure the precise correspondence among the preset action, the target text, and the voice, precisely implement multi-modal arrangement, and improve the naturalness and interactivity of the digital human live broadcast.
[0037] After completing the dragging and arrangement of all preset actions, all the speech information and action information of the digital human are obtained.
[0038] Optionally, the speech information and action information can be stored in the database for persistence respectively to ensure the reliable storage and subsequent processing of the data.
[0039] S140: Generate a digital human media stream based on the speech information and action information.
[0040] For the digital human media stream arrangement method provided by the embodiments of this application, after creating a digital human, display the text information in the multi-modal editor according to the speaking speed of the digital human, and in response to the user's dragging operation, drag multiple preset actions to below the target text corresponding to each preset action respectively, record the time coordinate of each preset action and the coordinate of the modal data type of each preset action, so as to obtain the speech information and action information of the digital human, and then generate a digital human media stream according to the speech information and action information. Through the above method, the function of action dragging and time coordinate recording is introduced in the process of multi-modal data arrangement, and the precise pairing of each action and the corresponding text in the time coordinate is realized by dragging, providing the logic of action-text matching for the arrangement of multi-modal data, which is beneficial to realizing the precise arrangement of multi-modal data, improving the data processing efficiency and the overall processing speed, and further beneficial to ensuring the real-time performance of the digital human live broadcast.
[0041] In some embodiments, generating a digital human media stream based on the speech information and action information includes: obtaining metadata from an object storage service platform based on an asynchronous processing speech arrangement mechanism; the metadata is the metadata of the image video material corresponding to the speech information and action information; converting the metadata into multiple target data with a preset format based on the speech information and action information; generating a digital human media stream based on the multiple target data.
[0042] As Figure 2 shown, after completing multimodal editing, enter the multimodal orchestration processing stage. In the multimodal orchestration processing stage, all the speech information and action information of the digital human can be queried from the database to provide basic data support for subsequent processing.
[0043] Specifically, based on the asynchronous processing speech orchestration mechanism, metadata is obtained from the Object Storage Service (OSS) platform.
[0044] The metadata is the metadata of the image video material corresponding to the speech information and action information. As the original data uploaded in the initialization material stage, the metadata can be directly downloaded from the Object Storage Service (OSS) platform and temporarily saved to the temporary queue List for further processing.
[0045] Furthermore, based on the speech information and action information, the metadata is converted into multiple target data with a preset format, and based on the multiple target data, a digital human media stream is generated.
[0046] In some embodiments, the metadata includes the action type of each preset action and the action ID of each preset action, the speech information includes the time sequence of each target word in the text information, and the action information includes the time coordinates of each preset action and the modal data type coordinates of each preset action; converting the metadata into multiple target data with a preset format based on the speech information and action information includes: performing format conversion on the action type of each preset action and the action ID of each preset action to obtain multiple action data with a preset format; an action data is determined by the action type of a preset action and the action ID of a preset action; determining the text insertion position corresponding to each preset action based on the time coordinates of each preset action; performing copywriting cutting processing on the text information based on the time coordinates of each preset action to obtain the target word corresponding to each preset action; respectively splicing the corresponding action data for each target word based on the text insertion position corresponding to each preset action to obtain multiple target data with a preset format; an action data corresponds to a target word, and a target data is composed of an action data and a target word; storing each target data into the ZSet cache of the Redis database in sequence according to the time sequence of each target word; wherein, the preset format is any one of the VSML format, JSON format, and XML format.
[0047] In this embodiment, the metadata includes the action type of each preset action and the action ID of each preset action, the speech information includes the time sequence of each target word in the text information, and the action information includes the time coordinates of each preset action and the modal data type coordinates of each preset action; the preset format is any one of the VSML format, JSON format, and XML format.
[0048] For ease of understanding, in this embodiment, the preset format is taken as the VSML format for introduction.
[0049] Specifically, in the multi-modal choreography processing stage, for each preset action, the JAXB technology can be used to perform format conversion on the action type of the preset action and the action ID of the preset action, quickly convert the action type of the preset action and the action ID of the preset action into a VSML format language that meets the docking requirements, perform preliminary data encapsulation, and obtain action data in the VSML format, that is, action VSML data. The action VSML data is obtained by combining the action type of the preset action and the action ID of the preset action.
[0050] Furthermore, the system calculates the text insertion position corresponding to the preset action according to the time coordinate of the preset action in the multi-modal editor to ensure the precise matching between the preset action and the target text in the text information.
[0051] Furthermore, based on the time coordinate of the preset action in the multi-modal editor, perform copywriting cutting on the text information to obtain the target text corresponding to the preset action, and according to the text insertion position corresponding to the preset action, splice the corresponding action VSML data for the target text, and combine to obtain the target data in the VSML format corresponding to the preset action (i.e., the target VSML text content) to ensure the integrity and consistency of the data.
[0052] In a digital human live broadcast, there are multiple preset actions of the digital human. The above processing is performed on each preset action, and finally multiple target data in the VSML format can be obtained; among them, one action VSML data corresponds to one target text, and one target data (i.e., the target VSML text content) is composed of one action VSML data and one target text.
[0053] Furthermore, use the ZSet data structure of the Redis database for storage, and store each target data in the ZSet cache of the Redis database in the time order of each target text (i.e., the order from smallest to largest time coordinate) to ensure that the target data is arranged in time order and improve the real-time performance and efficiency of data processing.
[0054] It should be noted that this embodiment only takes the VSML format as an example for introduction. In actual application processes, other forms of languages or API interfaces can also be used to define the behaviors and actions of digital humans. For example, the JSON format or XML format can be used to describe the multi-modal interaction logic, which can avoid directly using the VSML format language but still achieve similar functions.
[0055] The digital human media stream orchestration method provided by the embodiments of the present application realizes precise multi-modal orchestration through a multi-modal editor, improves the naturalness and interactivity of digital human live broadcasts, and realizes the refinement and automation of multi-modal orchestration; at the same time, provides an efficient speech script orchestration slicing processing solution, which improves the real-time performance and efficiency of data processing and reduces data latency through an efficient data processing and caching mechanism.
[0056] In some embodiments, based on multiple target data, a digital human media stream is generated, including: sequentially obtaining each target data from the ZSet cache according to the time sequence of each target text; converting all target data into the form of a data stream based on the Protobuf protocol to generate a digital human media stream; transmitting the digital human media stream based on a target communication protocol; where the target communication protocol is any one of the gRPC communication protocol, the WebSocket communication protocol, and the HTTP / 2 communication protocol.
[0057] As Figure 2 shown, after completing the multi-modal orchestration processing, it enters the digital human media stream generation stage. In the digital human media stream generation stage, each target data can be sequentially obtained from the ZSet cache of the Redis database according to the time sequence of each target text (i.e., the order from smallest to largest time coordinates), ensuring the real-time performance and efficiency of the data.
[0058] Furthermore, all the obtained target data is converted into the form of a data stream based on the Protobuf protocol to generate a digital human media stream, ensuring the high efficiency and compatibility of data transmission.
[0059] Furthermore, the digital human media stream is transmitted based on a target communication protocol.
[0060] Where the target communication protocol is any one of the gRPC communication protocol, the WebSocket communication protocol, and the HTTP / 2 communication protocol.
[0061] Among them, the gRPC communication protocol in the gRPC technology is used for the content transmission of two-way streams, which has the advantages of efficient transmission and multiplexing, can ensure the real-time performance and stability of data during digital human live broadcasts, reduce network load at the same time, and improve the user experience.
[0062] Optionally, in addition to the gRPC communication protocol, other efficient real-time communication protocols can also be adopted, such as the WebSocket communication protocol or other protocols based on the HTTP / 2 communication protocol (such as the Server-SentEvents protocol of HTTP / 2). Although these protocols may not be as efficient as the gRPC communication protocol in some aspects, they can still achieve the content transmission of two-way streams.
[0063] The digital human media stream orchestration method provided by the embodiments of this application adopts real-time and efficient communication technologies, and through an efficient real-time communication protocol, improves the efficiency and real-time performance of data transmission and reduces network load.
[0064] In some embodiments, a digital human is created, and the text information of the digital human is determined, including: in response to a user's selection instruction, determining a preset image of the digital human, and creating the digital human based on the preset image; adjusting the background image of the digital human and the position of the digital human in the background image based on the user interface; in response to the user's text input instruction, determining the text information of the digital human; wherein, the text information includes the image description copywriting and product description copywriting of the digital human; the user interface is any one of a graphical interface, a command-line tool, and an API interface.
[0065] As Figure 2 shown, before creating the digital human, it is necessary to initialize the materials. In the stage of initializing the materials, the user can obtain the preset image video materials from local or cloud storage by calling the offline interface, and the image video materials include image video material metadata.
[0066] Optionally, upload the image video materials to the Object Storage Service (OSS) platform to ensure the security and efficient access of the image video materials.
[0067] Optionally, persist the image video material metadata (such as the action ID, action type, and OSS address of the preset action, etc.) to the database to provide more comprehensive information support for subsequent processing and ensure the reliable storage of data.
[0068] Furthermore, in response to the user's selection instruction, select a preset image for the digital human to be created, and create the digital human according to the preset image.
[0069] Optionally, after creating the digital human, the user can, according to the needs, replace or adjust the background image of the digital human through the user interface, and adjust the position of the digital human in the background image to achieve a more personalized design and enhance the richness and attractiveness of the content.
[0070] Optionally, the user interface is any one of a graphical interface, a command-line tool, and an API interface.
[0071] Among them, by adjusting the background image or text content of the digital human through the graphical interface, the adjustment effect can be intuitively reflected to the user; by adjusting the background image or text content of the digital human through other user interfaces (such as command-line tools and API interfaces), although the adjustment effect is not as intuitive as the graphical interface, the same function can still be achieved.
[0072] Further, in response to a user's text input instruction, text information of the digital human is added; the text information includes a copywriting for describing the image of the digital human, a global copywriting related to the digital human, and a product description copywriting, so as to improve the content richness and attractiveness of the digital human live broadcast.
[0073] The digital human media stream orchestration method provided by the embodiments of this application can achieve dynamic content adjustment and background configuration, and improve the flexibility and diversity of the digital human live broadcast.
[0074] In some embodiments, before creating a digital human and determining the text information of the digital human, it further includes: obtaining metadata of the image video material; uploading the metadata of the image video material to an object storage service platform; storing the metadata of the image video material in a database.
[0075] As Figure 2 shown, before creating a digital human, it is necessary to initialize the materials. In the stage of initializing the materials, the user can obtain preset image video materials from local or cloud storage by calling an offline interface, and the image video materials include metadata of the image video materials.
[0076] Optionally, upload the image video material to an Object Storage Service (OSS) platform to ensure the security and efficient access of the image video material.
[0077] Optionally, persist the metadata of the image video material (such as the action ID, action type, and OSS address of the preset action, etc.) to the database, provide more comprehensive information support for subsequent processing, and ensure the reliable storage of the data.
[0078] The digital human media stream orchestration method provided by the embodiments of this application improves the accuracy of data management and use through detailed metadata recording and persistence, and can provide more comprehensive information support for subsequent processing of multi-modal data.
[0079] The embodiments of this application also provide a digital human media stream orchestration device. Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of the digital human media stream orchestration device provided by the embodiments of this application. In the embodiments of this application, the digital human media stream orchestration device includes a creation module 310, a display module 320, an editing module 330, and a media stream generation module 340.
[0080] The creation module 310 is used to create a digital human and determine the text information of the digital human.
[0081] The display module 320 is used to display the text information in a multi-modal editor according to the speech rate of the digital human.
[0082] The abscissa of the multi-modal editor represents the timeline, and the ordinate of the multi-modal editor represents the modal data type.
[0083] An editing module 330, configured to, in response to a user's dragging operation, drag multiple preset actions to below the target text corresponding to each preset action respectively, and record the time coordinates of each preset action and the modal data type coordinates of each preset action, so as to obtain the speech information and action information of the digital human.
[0084] One preset action corresponds to one target text in the text information.
[0085] A media stream generation module 340, configured to generate a digital human media stream based on the speech information and action information.
[0086] In some embodiments, the media stream generation module 340 is configured to obtain metadata from an object storage service platform based on an asynchronous processing speech arrangement mechanism; the metadata is the image video material metadata corresponding to the speech information and action information; convert the metadata into multiple target data with a preset format based on the speech information and action information; and generate a digital human media stream based on the multiple target data.
[0087] In some embodiments, the metadata includes the action type of each preset action and the action ID of each preset action, the speech information includes the time sequence of each target text in the text information, and the action information includes the time coordinates of each preset action and the modal data type coordinates of each preset action.
[0088] The media stream generation module 340 is configured to perform format conversion on the action type of each preset action and the action ID of each preset action to obtain multiple action data with a preset format; one action data is determined by the action type of one preset action and the action ID of one preset action; determine the text insertion position corresponding to each preset action based on the time coordinates of each preset action; perform copywriting cutting processing on the text information based on the time coordinates of each preset action to obtain the target text corresponding to each preset action; respectively splice the corresponding action data for each target text based on the text insertion position corresponding to each preset action to obtain multiple target data with a preset format; one action data corresponds to one target text, and one target data is composed of one action data and one target text; store each target data into the ZSet cache of the Redis database in sequence according to the time sequence of each target text; wherein, the preset format is any one of the VSML format, JSON format, and XML format.
[0089] In some embodiments, the media stream generation module 340 is configured to sequentially obtain each target data from the ZSet cache according to the chronological order of each target text; convert all the target data into the form of a data stream based on the Protobuf protocol to generate a digital human media stream; and transmit the digital human media stream based on the target communication protocol, where the target communication protocol is any one of the gRPC communication protocol, the WebSocket communication protocol, and the HTTP / 2 communication protocol.
[0090] In some embodiments, the creation module 310 is configured to, in response to a user's selection instruction, determine a preset image of the digital human and create the digital human based on the preset image; adjust the background image of the digital human and the position of the digital human in the background image based on the user interface; and determine the text information of the digital human in response to the user's text input instruction, where the text information includes an image description copywriting and a product description copywriting of the digital human; and the user interface is any one of a graphical interface, a command line tool, and an API interface.
[0091] In some embodiments, the creation module 310 is configured to obtain metadata of the image video material; upload the metadata of the image video material to the object storage service platform; and store the metadata of the image video material in the database.
[0092] The embodiment of the present application further provides an electronic device. Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application. As Figure 4 shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call the logical instructions in the memory 430 to execute the digital human media stream orchestration method.
[0093] In addition, when the logical instructions in the above-mentioned memory 430 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0094] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the digital human media stream choreography method provided by the above-mentioned various methods.
[0095] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the digital human media stream choreography method provided by the above-mentioned various methods.
[0096] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0097] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.
[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for choreographing digital human media streams, characterized in that, Including: Create a digital human and determine the text information of the digital human; Display the text information in a multimodal editor according to the speaking speed of the digital human; The abscissa of the multimodal editor represents the timeline, and the ordinate of the multimodal editor represents the modal data type; In response to the user's dragging operation, drag multiple preset actions to the bottom of each target text corresponding to each preset action, and record the time coordinates of each preset action and the modal data type coordinates of each preset action to obtain the speech information and action information of the digital human; One of the preset actions corresponds to one of the target texts in the text information; Generate a digital human media stream based on the speech information and the action information.
2. The digital human media stream choreography method according to claim 1, wherein The generating the digital human media stream based on the speech information and the action information includes: Based on an asynchronous processing speech arrangement mechanism, obtain metadata from an object storage service platform; the metadata is the metadata of the image video material corresponding to the speech information and the action information; Based on the speech information and the action information, convert the metadata into multiple target data with a preset format; Generate the digital human media stream based on the multiple target data.
3. The digital human media stream choreography method according to claim 2, wherein The metadata includes the action type of each preset action and the action ID of each preset action, the speech information includes the time sequence of each target text in the text information, and the action information includes the time coordinates of each preset action and the modal data type coordinates of each preset action; The converting the metadata into multiple target data with a preset format based on the speech information and the action information includes: Perform format conversion on the action type of each preset action and the action ID of each preset action to obtain multiple action data with the preset format; one action data is determined by the action type of one preset action and the action ID of one preset action; Based on the time coordinates of each preset action, determine the text insertion position corresponding to each preset action; Based on the time coordinates of each preset action, perform copywriting cutting processing on the text information to obtain the target text corresponding to each preset action; Based on the text insertion position corresponding to each preset action, splice the corresponding action data for each target text respectively to obtain multiple target data with a preset format; one action data corresponds to one target text, and one target data is composed of one action data and one target text; Store each target data in the ZSet cache of the Redis database in sequence according to the time sequence of each target text; Wherein, the preset format is any one of the VSML format, the JSON format, and the XML format.
4. The digital human media stream choreography method according to claim 3, wherein The generating the digital human media stream based on the multiple target data includes: Obtain each target data from the ZSet cache in sequence according to the time sequence of each target text; Based on the Protobuf protocol, convert all the target data into the form of a data stream to generate the digital human media stream; Transmit the digital human media stream based on the target communication protocol; wherein, the target communication protocol is any one of the gRPC communication protocol, the WebSocket communication protocol, and the HTTP / 2 communication protocol.
5. The digital human media stream choreography method according to claim 1, characterized in that The creation of the digital human and the determination of the text information of the digital human include: In response to the user's selection instruction, determine the preset image of the digital human, and create the digital human based on the preset image; Based on the user interface, adjust the background image of the digital human and the position of the digital human in the background image; In response to the user's text input instruction, determine the text information of the digital human; wherein, the text information includes the image description copywriting and the product description copywriting of the digital human; the user interface is any one of a graphical interface, a command-line tool, and an API interface.
6. The digital human media stream choreography method according to claim 2, characterized in that, Before creating the digital human and determining the text information of the digital human, it further includes: Obtain the metadata of the image video material; Upload the metadata of the image video material to the object storage service platform; Store the metadata of the image video material in the database.
7. A digital human media stream choreography device, characterized in that It includes: A creation module for creating a digital human and determining the text information of the digital human; A display module for displaying the text information in a multimodal editor according to the speech rate of the digital human; the abscissa of the multimodal editor represents the time axis, and the ordinate of the multimodal editor represents the modal data type; An editing module for, in response to the user's dragging operation, dragging multiple preset actions to the lower part of each target text corresponding to each preset action, and recording the time coordinates of each preset action and the modal data type coordinates of each preset action to obtain the speech information and action information of the digital human; One of the preset actions corresponds to one of the target texts in the text information; A media stream generation module for generating a digital human media stream based on the speech information and the action information.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the digital human media stream orchestration method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the digital human media stream orchestration method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the digital human media stream orchestration method according to any one of claims 1 to 6.