Method for generating digital humans, model training method, device, equipment, and medium
By segmenting material content into scenes and constructing digital humans at the scene granularity, the method addresses the inconsistency issue in existing technologies, ensuring alignment with target content and enhancing user experience in digital human video production.
Patent Information
- Application Number
- JP2024571955
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-15
- Filing Date
- 2022-12-02
- Publication Date
- 2025-07-31
AI Technical Summary
Existing digital human technologies in video production often fail to ensure consistency between the digital human and the content, leading to a poor user experience due to separation of the digital human and broadcast content, and focus on fictional scenes rather than real information.
The method involves segmenting material content into scenes using a pre-trained scene segmentation model, determining target content for each scene, and constructing a digital human specific to the scene label information, ensuring consistency and enhancing user experience.
This approach improves the fusion between material content and digital human, resulting in enhanced user experience by aligning the digital human with the scene and target content, making it suitable for real information broadcast.
Smart Images

Figure 2025524735000001_ABST
Abstract
Description
Cross - reference to related applications
[0001] This application claims the priority of Chinese Patent Application No. 202210681368.3 filed on June 15, 2022, and the entire content thereof is incorporated herein by reference.
Technical Field
[0002] The present disclosure relates to the field of artificial intelligence, specifically to technical fields such as natural language processing, deep learning, computer vision, image processing, augmented reality, and virtual reality, and can be applied to meta Birth and so on Scene and the like, and particularly relates to a method for generating a digital human, a method for training a neural network, a video generation device, a neural network training device, an electronic device, a computer - readable storage medium, and a computer program product.
Background Art
[0003] Artificial intelligence is a subject that studies to simulate some human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) on a computer, including both hardware - side technologies and software - side technologies. The hardware technologies of artificial intelligence generally include technologies such as sensors, artificial - intelligence - dedicated chips, cloud computing, distributed storage, and big - data processing. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big - data processing technology, and knowledge graph technology.
[0004] Digital humans are a technology that uses computer technology to virtually simulate the form and function of the human body. Digital humans can significantly improve the interactivity of applications and enhance the intelligence level of intelligent information services. With the continuous breakthroughs in artificial intelligence technology, the image, expression, and presentation of digital humans are gradually becoming more similar to real people. The application scenarios of digital humans are constantly expanding, and digital humans have gradually become an important service form in the digital world.
[0005] The methods described in this section are not necessarily methods previously contemplated or adopted. Unless otherwise noted, any methods described in this section should not be considered prior art merely because they are included in this section. Similarly, unless otherwise noted, the problems addressed in this section should not be considered to be acknowledged in the prior art. Summary of the Invention
[0006] The present disclosure provides a method for generating a digital human, a method for training a neural network, a video generation apparatus, a neural network training apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
[0007] According to one aspect of the present disclosure, there is provided a method for generating a digital human, the method comprising: Inside the material and determining a plurality of scenes from the material content based on the content of the material and a pre-trained scene segmentation model, wherein each scene of the plurality of scenes is a segment of the material content. 1 One complete Semantic The method includes: corresponding to content segments having information; for each scene of the plurality of scenes, determining target content corresponding to the scene based on the corresponding content segment; determining scene label information for the scene based on the corresponding target content; and constructing a digital human specific to the scene based on the scene label information.
[0008] According to another aspect of the present disclosure, there is provided a method for training a scene segmentation model, which includes obtaining sample material content and a plurality of sample scenes in the sample material content, determining a plurality of predicted scenes from the sample material content based on a preset scene segmentation model, and adjusting parameters of the preset scene segmentation model based on the plurality of sample scenes and the plurality of predicted scenes to obtain a trained scene segmentation model.
[0009] According to another aspect of the present disclosure, there is provided an apparatus for generating a digital human, comprising: a first acquisition unit configured to acquire material content; and determining a plurality of scenes from the material content based on a pre-trained scene segmentation model, wherein each scene of the plurality of scenes comprises: 1 One complete Semantic The system includes a first determination unit configured to correspond to a content segment having information; a second determination unit configured, for each scene of the plurality of scenes, to determine target content corresponding to the scene based on the corresponding content segment; a third determination unit configured to determine scene label information of the scene based on the corresponding target content; and a digital human construction unit configured to construct a digital human specific to the scene based on the scene label information.
[0010] According to another aspect of the present disclosure, there is provided an apparatus for training a scene segmentation model, which includes: a third acquisition unit configured to acquire sample material content and a plurality of sample scenes in the sample material content; a seventh determination unit configured to determine a plurality of predicted scenes from the sample material content based on a preset scene segmentation model; and a training unit configured to adjust parameters of the preset scene segmentation model based on the plurality of sample scenes and the plurality of predicted scenes to obtain a trained scene segmentation model.
[0011] According to another aspect of the present disclosure, there is provided an electronic device, the electronic device comprising at least 1 At least one processor 1 and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method described above.
[0012] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having stored thereon computer instructions for causing a computer to perform the method described above. According to another aspect of the present disclosure, a computer program product including a computer program to offer do The computer program, when executed by a processor, implements the above-described method.
[0013] The present invention 1 According to the above embodiments, by dividing the material content into scenes and constructing a digital human with a scene as the granularity, the consistency between the digital human and the scene and target content is ensured, the integration between the material content and the digital human is improved, and the user's viewing experience of the digital human is enhanced.
[0014] It should be understood that the contents described in this section are not intended to identify key or important features of the embodiments of the present disclosure, and are not intended to limit the scope of protection of the present disclosure. Other features of the present disclosure will be easily understood from the following description.
[0015] The drawings illustratively illustrate examples, constitute a part of the specification, and together with the written description serve to explain exemplary embodiments of the examples. The illustrated examples are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar, but not necessarily identical, elements. [Brief explanation of the drawings]
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
[0017] Hereinafter, exemplary embodiments of the present disclosure will be described in conjunction with the drawings, and various details of the embodiments of the present disclosure included therein are intended to facilitate understanding and should be considered merely exemplary. Therefore, those skilled in the art will be able to easily implement the inventions described herein without departing from the scope and spirit of the present disclosure. to do It should be understood that various changes and modifications can be made to the embodiments. Similarly, for the sake of clarity and conciseness, the following description omits descriptions of well-known functions and structures.
[0018] This disclosure In this specification, unless otherwise specified, the terms "first," "second," and the like, used to describe various elements are not intended to limit the location, timing, or importance of those elements. Such terms are used only to distinguish one element from another. In some instances, a first element and a second element may refer to the same instance of an element, or in some cases, may refer to different instances based on the context.
[0019] The terms used in the description of various examples of the present disclosure are intended only to describe particular examples and are not intended to be limiting. Unless the context clearly indicates otherwise, an element may be one or more, unless the number of elements is specifically limited. It should be noted that the term "and / or" as used in this disclosure refers to any and all possible combinations of the listed items. Method Covers.
[0020] Video is one of the most important information media in the digital world. Naturally, digital humans have an important application space in video production. Currently, digital humans have already begun to be used in video production. For example, news is broadcast through digital humans, and advertising is carried out using the images of digital humans. However, in related technologies, the operation of digital humans in videos is mainly carried out based on templates. For example, fixed digital humans conduct broadcasts. When a digital human broadcasts, the digital human and the content are separated, the broadcast content does not match the image of the digital human, and the user experience is poor. In other related technologies, with the aim of displaying the image of a digital human, the focus is on the precise construction of digital human idols. This method usually targets several fictional and science fiction scenes and is not easy to use for the broadcast of real information. Furthermore, in such scenes, since the main purpose is to display the image, various attributes of the digital human are generally irrelevant to the broadcast content.
[0021] To solve such problems, This disclosure it is to split the material content by scene and construct digital humans at the scene granularity, so as to ensure the consistency between the digital human and the scene and the target content, improve the fusion of the material content and the digital human, and enhance the user experience of watching the digital human.
[0022] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. FIG. 1 shows a schematic diagram of an exemplary system 100 in which various methods and apparatuses described in this specification can be implemented according to an embodiment of the present disclosure. Referring to FIG. 1, the system 100 includes 1 one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and 1 one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 are 1 included.1 The CPU may be configured to run one or more applications.
[0023] In an embodiment of the present disclosure, the server 120 may execute one or more services or software applications that enable the execution of methods for generating digital humans and / or methods for training scene segmentation models.
[0024] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtualized and virtualized environments. In some embodiments, these services may be provided as web-based or cloud services, for example, provided to users of client devices 101, 102, 103, 104, 105, and / or 106 in a Software as a Service (SaaS) model.
[0025] As shown in Figure 1 Configuration In the example, the server 120 implements the functions performed by the server 120. 1 The assembly may include one or more assemblies, each of which may include: 1 The client devices 101, 102, 103, 104, 105, and / or 106 may include software assemblies, hardware assemblies, or a combination thereof, that can run on one or more processors. In order to use the services provided by these assemblies, users operating the client devices 101, 102, 103, 104, 105, and / or 106 may 1 One or more client applications may be used to interact with the server 120. Configuration It should be understood that other systems may be possible and may differ from system 100. Thus, Figure 1 is intended to be an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0026] A user can use client devices 101, 102, 103, 104, 105, and / or 106 to input parameters related to the generation of a digital human. The client devices may also be used as interfaces through which users of the client devices interact with the client devices. dash The client device may provide a user interface through the interface. The client device may also output information to the user, such as outputting the results of the digital human generation to the user. Although only six client devices are shown in Figure 1, one skilled in the art will appreciate that the present disclosure can support any number of client devices.
[0027] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (e.g., personal computers or laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computing devices may run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX operating systems, Linux, or Linux operating systems (e.g., Google Chrome OS), or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices include mobile phones, SmartThe client devices may include phones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (e.g., smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices may run various applications, such as internet-related applications, communication applications (e.g., email applications), short message service (SMS) applications, and may use various communication protocols.
[0028] The network 110 is well known to those skilled in the art. well It may be any type of network known, and it may use any of several available protocols to support data communications. 1 One can use one of the following (including but not limited to TCP / IP, SNA, IPX, etc.). For example: 1 The one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token loop, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0029] The server 120 1 Server 120 may include one or more general-purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, large computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may run a virtual operating system. 1One or more virtual machines, or other computing architectures involving virtualization (e.g., virtualized logical storage devices to maintain virtual storage for a server) 1 In various embodiments, server 120 provides the functionality described below. 1 One or more services or software applications may be running on the network.
[0030] The computing units in the server 120 can include any of the operating systems listed above and any commercial server operating system. 1 Server 120 may run one or more operating systems. Server 120 may also run any of a variety of additional server applications and / or middle tier applications, such as an HTTP server, an FTP server, a CGI server, a JAVA server, a database server, etc. 1 You can also run one.
[0031] In some embodiments, the server 120 may process data feeds and / or event updates received from users of the client devices 101, 102, 103, 104, 105, and / or 106. Analysis and for integration 1 The server 120 may include one or more applications. 1 Displaying data feeds and / or real-time events via one or more display devices 1 The application may include one or more applications.
[0032] In some embodiments, the server 120 may be a server of a distributed system or a server incorporating a blockchain. The server 120 may be a cloud server, or an intelligent cloud computing server or an intelligent cloud host equipped with artificial intelligence technology. A cloud server is a host product in a cloud computing service system, which solves the defects of great management difficulty and weak business scalability existing in conventional physical hosts and virtual private server (VPS) services.
[0033] The system 100 may 1 also include one or more databases 130. In some embodiments, these databases can be used to store data and other information. For example, 1 one or more of the databases 130 can be used to store information such as audio files and video files. The databases 130 can be arranged at various positions. For example, the database used by the server 120 may be local to the server 120, or may communicate with the server 120 via a network or a dedicated connection away from the server 120. The databases 130 can be of various types. In some embodiments, the database used by the server 120 may be a relational database. 1 One or more of these databases can store, update, and retrieve data from the database in response to instructions.
[0034] In some embodiments, 1 one or more of the databases 130 are used by an application and can also store the data of the application. The database used by the application can be various types of databases, such as a key-value repository, an object repository, and a general-purpose repository supported by a file system.
[0035] The system 100 in FIG. 1 can be configured and operated in various ways so as to apply the various methods and apparatuses described based on the present disclosure. According to one aspect of the present disclosure, a method for generating a digital human is provided. As shown in FIG. 2, the method includes a step S201 of acquiring material content, and based on a pre-trained scene segmentation model, determining a plurality of scenes from the material content. Here, each scene of the plurality of scenes is respectively a 1 complete Semantic step S202 corresponding to a content segment having information, a step S203 of determining target content corresponding to each scene of the plurality of scenes based on the corresponding content segment, a step S204 of determining scene label information of the scene based on the corresponding target content, and a step S205 of constructing a digital human specific to the scene based on the scene label information.
[0036] Thereby, by segmenting the material content into scenes and constructing a digital human at the scene granularity, the consistency between the digital human, the scene, and the target content is ensured, the fusion between the material content and the digital human is improved, and the user's experience of viewing the digital human is enhanced.
[0037] According to some embodiments, before starting the generation of the digital human, it is possible to support the user to set basic setting options via an application terminal (for example, any one of the clients 101 to 106 in FIG. 1 1 one). The specific content that can be set is as follows.
[0038] The method of the present disclosure can be applied to various scenes such as a broadcast scene, an explanation scene, and a host scene. In the present disclosure, the various methods of the present disclosure are mainly described by taking the broadcast scene as an example, but it should be understood that the protection scope of the present disclosure is not intended to be limited.
[0039] severalIn the embodiments, it is possible to support the user in selecting and configuring the types of material content. The types of material content and the corresponding files, addresses, or content can include the following: (1) A text document, specifically, a document containing text content or Image text a document containing content, (2) An article URL, that is, the network address corresponding to the content for which digital human generation is desired Image text , (3) Topic keywords and descriptions, which can include, for example, forms such as entity words, search keywords, search questions, etc., and describe the topics for which digital human generation is desired. In some exemplary embodiments, the material content can include at least one of image data and video data, as well as text data, thereby enriching the content of digital human broadcasts, hosting, or commentary.
[0040] In some embodiments, it is possible to support the user in setting the voice broadcast (TTS, Text to Speech) function, including whether to turn on the voice broadcast function, selecting the voice of the voice broadcast (e.g., gender, accent, etc.), voice color, volume, and speech rate, etc.
[0041] In some embodiments, it is possible to support the user in setting background music, including selecting whether to add background music and the type of background music to be added, etc. In some embodiments, it is possible to support the user in setting digital human assets, including selecting the image of the digital human that is desired to appear or be used in a pre-set digital human asset, or generating the image of the digital human by custom means to enrich the digital human asset.
[0042] In some embodiments, it is possible to support the user in setting the digital human background, including selecting whether to add a digital human background and the type of digital human background (e.g., image or video), etc.
[0043] Some embodiments may support user configuration of how the final video is generated, including selection of fully automatic video generation, human-machine interaction assisted video generation, and the like.
[0044] In addition to the above, it should be understood that the input settings may provide the user with further system control, depending on the circumstances, such as (but not limited to) the ratio of copy compression, dynamic material occupancy used in the final presentation, etc.
[0045] According to some embodiments, the step S201 of acquiring the material content may include acquiring the material content based on at least one of a method of acquiring the material content based on a web page address and a method of acquiring the material content based on a search keyword. These different types of material content can be specifically acquired by the following methods:
[0046] For text documents: Directly read the contents of a text document stored locally or remotely. Text URL In the case of : Mainly existing on the Internet Image text It refers to the content, such as the URL of the news article, forum article, Q&A page, official account article, etc., and analyzes the open source solution based on the existing web page to obtain the web page data corresponding to the URL, and then analyzes and obtains the main text and image content in the URL web page, and at the same time, the title, body, paragraph, bold, Image text Important primitive information such as positional relationships, tables, etc. are recorded and used in the subsequent digital human generation process. For example, the title can be used as a query to search for visual materials, while the body text, paragraphs, and bold content can be used to generate broadcast content. Scene keywords and sentence level keywords can also be used. of It can be used to extract keywords, Image textThe positional relationship can provide the correspondence between the image material in the original text and the broadcast content, and the table can enrich the content presentation format in the presentation results of the digital human, which will be introduced in detail below.
[0047] For topic keywords and descriptions: The system also supports generating the final digital human presentation results based on the topic descriptions entered by the user. The topic keywords and descriptions entered by the user are: Encyclopedia words It may be an entity word similar to the topic keyword, multiple keywords, or a format similar to an event description or problem description. Image text The methods for obtaining content are as follows: (1) inputting topic keywords or descriptions into a search engine, (2) obtaining multiple search response results, and (3) selecting from the search results that have a high relevance ranking and richer visual material. Image text (4) Select the URL of the video whose results you want to generate as the body URL, and then use the URL content extraction method to extract the Image text Extract information such as content.
[0048] Once the content of the material is acquired, a script for digital human broadcasting can be generated based on the acquired content. Based on the acquired content of the material, scene division, text conversion and generation are performed according to the needs of digital human broadcasting using semantic understanding and text generation techniques, and the script required for creating digital human video / holographic projection is output, while scene division and Semantic analysis information The above processing of material content determines the digital-human integration method. Important It's the foundation. Specific Next, the target content of digital human broadcasting can be generated as follows.
[0049] In some embodiments, in step S202, a plurality of scenes are determined from the material content based on a pre-trained scene segmentation model, where each scene of the plurality of scenes is a complete segment of the material content. SemanticIt can correspond to a content segment having information. By dividing the material content into a plurality of scenes, it becomes easy to perform different processes on different scenes during subsequent digital human broadcasts. In the present disclosure, digital human fusion is granules fused at the degree.
[0050] In some embodiments According to , as shown in FIG. 3, step S202 of determining a plurality of scenes from the material content based on a pre-trained scene segmentation model includes performing chapter structure analysis and chapter Semantic segmentation on the material content to determine a plurality of subtopics from the material content and determine the structural relationship between the plurality of subtopics, and step S302 of dividing the plurality of subtopics into a plurality of scenes based on the structural relationship. In this way, by using chapter structure analysis and chapter Semantic segmentation Method , it is possible to accurately identify a plurality of semantically is complete subtopics and the structural relationship between them, and utilize the structural relationship between subtopics (for example, the general-part structure, parallel structure, progressive structure, etc.) to further divide the plurality of subtopics into a plurality of scenes more suitable for digital human broadcast tasks.
[0051] In some embodiments, the scene segmentation model can be obtained by pre-training and can further include a chapter structure analysis model, a chapter Semantic segmentation model, and a model for dividing subtopics into scenes. These models may be rule-based models or deep learning-based models, but are not limited here. In an exemplary embodiment, chapter SemanticThe segmentation model can output the semantic relationships between semantically complete content segments and adjacent content segments (e.g., the semantic similarity between two adjacent sentences in each pair). The chapter structure analysis model can output the relationships between content segments (e.g., summation, parallelism, progression, etc.). By combining the two, when the material content is divided into multiple subtopics, the semantic relationships between the content segments can be output. obtained in The relationship between the location of the division boundary and the subtopics can be obtained. Furthermore, the above method divides the material content based only on the text (natural language processing), and the obtained done The segmentation results may not be suitable for use as scenes for broadcasting. For example, the content segments corresponding to some subtopics may be very short and contain only one word. However, if such subtopics are used as scenes, they may have frequent transitions. Therefore, multiple subtopics can be further segmented into multiple scenes. In some embodiments, multiple subtopics can be further segmented into multiple scenes by dividing the subtopics into scenes based on the structure between the subtopics. Structural relationship Further division may be performed using a correlation or other methods (such as, but not limited to, rule-based methods or neural network methods based on sequence labeling).
[0052] After multiple scenes are obtained, the target content (e.g., broadcast words, commentary words, host words) for each scene can be determined. In some embodiments, the complete content within the material can be determined. Semantic The informative content segments can be used directly as target content. In some other embodiments, the source of the material content is complex, content-rich, and usually written, so using the content segments of each scene directly as target content for the digital human is not feasible. The effect is poor The material content and / or content segments can be processed as follows to obtain the corresponding target content.
[0053] According to some embodiments, processes such as text compression, text rewriting, and style conversion may be performed on the material content and / or content segments. In some embodiments, as shown in FIG. 4, for each scene of a plurality of scenes, in step S203 of determining the target content corresponding to the scene based on the corresponding content segment, at least one of text rewriting and text compression may be performed on the corresponding content segment, and step S401 of updating the corresponding content segment may be included. According to the time length set by the user and the characteristics of the content, based on the overall material content and / or the content of each content segment, a concise and information-rich text result can be Target output as the content for each scene. In some embodiments, the target content may be determined using an extractive summarization algorithm, may be determined using a generative summarization algorithm, or may be determined using other methods, but is not limited thereto.
[0054] In some embodiments According to , as shown in FIG. 4, for each scene of a plurality of scenes, in step S203 of determining the target content corresponding to the scene based on the corresponding content segment, step S402 of generating the first content for the scene may further be included based on the relationship Structural relationship between this scene and the previous scene. Therefore, by generating the corresponding first content based on the relationship between scenes, the transition between scenes can be made more Structural relationship smooth and the transition can be made more natural. Smooth
[0055] In some embodiments, the first content may be, for example, a transition word or a transition sentence, and may also be other content that can show the relationship or between scenes by connecting a plurality of scenes, and is not limited in this specification. or something like that
[0056] In some embodiments According to, as shown in FIG. 4, for each scene of a plurality of scenes, the step S203 of determining the target content corresponding to the scene based on the corresponding content segment may further include a step S403 of converting the corresponding content segment into the corresponding target content based on a pre-trained style conversion model. Here, the style conversion model is Prompt obtained by training based on learning. In this way, by performing style conversion, the material content and / or content segment can be converted into a style suitable for the broadcast of the digital human. Also, Prompt by using the method based on learning (prompt learning), Prior the trained style conversion model can automatically output natural target content according to various characteristics of the digital human and Demand (for example, colloquialization requirements, digital human characteristics (gender, accent setting, etc.), content transition requirements). In some embodiments, the above-mentioned various characteristics and requirements are converted into attribute controls, and further, together with the text to be style-converted (and optionally, the context of the text), Prompt the input of the learning-based style conversion model is Construction made, and the model can be trained using samples having the above-mentioned structure (attribute control and text) so that the model can output the desired style conversion result. In some embodiments, a rule-based method (adding spoken language, transition words, etc.) is combined with the learning-based thing to also convert the content segment of each scene into a target content suitable for the digital human broadcast scene.
[0057] In some embodiments According to4, for each scene of the plurality of scenes, step S203 of determining target content corresponding to the scene based on the corresponding content segment may further include step S404 of performing at least one of text rewriting and text compression on the converted target content to update the corresponding target content. Operation The operations of step S401 and step S404 may be performed in the same manner. It should be understood that step S401 and step S404 may be performed alternatively, or both may be performed, and this is not limited here.
[0058] In one exemplary embodiment, for the title "Let's build a castle of museums! Proceed with the construction of six museums in X City," by performing the above steps, multiple scenes and corresponding goal contents can be obtained, for example, as shown in Table 1.
[0059] [Table 1]
[0060] After the target content corresponding to each scene is obtained, further processing is performed on each scene and the corresponding target content. Semantic Analysis can be performed to facilitate video material retrieval and recall, alignment of video material with copy, presentation of key information, and / or the generation of a digital human. To control the system This allows us to obtain a wealth of scene-related information for the following reasons: Semantic Examples of analytical methods include:
[0061] In some embodiments, scene keyword extraction, i.e., Core It allows automatic extraction of keywords that can be used to describe the entire content. Before In the example, for scene 1, keywords for that scene such as “castle in museum” and “X city” can be extracted. These scene keywords are used to construct material search queries such as “X city museum castle” and are used to recall related video material.
[0062] In some embodiments, at the sentence level of keyword extraction, that is, finer-grained keywords can be extracted from the sentence level. In the above example, for the sentence in Scene 4, "The east building of the X City Museum under construction has already achieved the facade and will be almost completed by the end of 2022.", the sentence-level keyword "the east building of the X City Museum" can be automatically extracted. This sentence-level keyword is not only used to construct the material search query, but also used for more accurate alignment of video materials and content (copywriting).
[0063] In some embodiments, in terms of content Table information extraction may be performed to obtain the results in the form of key-value (i.e., Key-Value pairs). In the above example, for Scene 2, a plurality of Key-Value pairs can be automatically extracted. For example, "The number of various registered museums in X City: 20 4」 ", "The number of free museums: 9 4」 ", "The total number of collections in the district museums of X City: 16.25 million Point ", etc. By using these extracted key information, more accurate and rich digital human broadcast materials (for example, the display board in the background of what the digital human broadcasts Scene ) can be automatically generated, thereby significantly improving the overall quality and presentation effect of digital human broadcasts. In some embodiments According to , the method for generating a digital human may further include, for each scene of a plurality of scenes, extracting key-[[]] Value form type information from the target content corresponding to the scene, and generating auxiliary materials for the final digital human presentation result based on the key-value type information. Thereby, more accurate and rich digital human visual materials can be obtained.
[0064] In some embodiments, sentiment Table analysis may be performed to output the sentiment tendency Core expressed by each material Analysis . BeforeIn the example of , for scene 5, the entire scene introduces the situation where the museum business in City X is booming, and the overall tone, sentiment, and emotion are positively improving. Such emotions Analysis As a result, the expressions, tones, and movements of the digital human can be more richly controlled. In some embodiments According to , the scene label information can include semantic labels, and based on the corresponding target content, the step S204 of determining the scene label information of the scene can include performing sentiment Analysis on the corresponding target content to obtain the semantic label. Thus, by performing sentiment Analysis on the target content, information such as the tone and sentiment of the scene can be obtained.
[0065] In some embodiments, the semantic label is used to identify the emotions of positive, neutral, and negative represented by the corresponding target content. It should be understood that the semantic label can identify richer emotions such as tension, joy, anger, etc. This is not limited here. In some embodiments, the semantic label can further include other contents, for example, directly extracting the text semantic features of the target content, the type of the target content (such as narrative type, criticism type, lyric type, etc.), and labels that can represent information about the meaning of the target content corresponding to other scenes, and this is not limited here.
[0066] In some embodiments, in addition to the above method, sentiment Analysis can be performed on the scene and the corresponding target content by other methods to obtain relevant information. For example, the text semantic features of the target content can be directly extracted as input information for subsequent digital human attribute planning.
[0067] After obtaining the target content corresponding to each scene, voice synthesis can also be performed. The purpose of voice synthesis is to generate voice for the digital human scene, that is, to convert the previously obtained target content into voice, and optionally, add background music to the digital human broadcast. The conversion of text to voice can be performed by calling a TTS service. In some embodiments, as shown in FIG. 5, the method of generating a digital human may further include step S505 of converting the target content into voice for digital human broadcast. Note that steps S501 to S504 and step S507 in FIG. 5 are Operation the same as steps S201 to S205 in FIG. 2, so the description is omitted here.
[0068] Regarding the background music, the system can call TTS capabilities with different tones and timbres, and background music of different styles according to the type of the target content (for example, narrative type, critical type, lyrical type, etc.). Further, as described above, the method of the present disclosure supports user specifications such as TTS and background music. The system can provide TTS and background music with various tones, timbres, and speaking speeds for the user to make independent selections, and support the user to customize a dedicated TTS timbre spontaneously.
[0069] In order to create a presentation result of a digital human broadcast rich in visual materials, material expansion is performed, and materials such as videos and images are used for digital broadcast Supplement can be done. The supplementation of video and image materials includes the following.
[0070] In some embodiments According toAs shown in FIG. 5, the method for generating a digital human can further include, for each scene of a plurality of scenes, a step S506 of searching for video materials related to the scene based on the material content and the target content corresponding to the scene, and a step S508 of combining the video materials with the digital human. Thereby, in this way, materials that match the scene and the corresponding target content and are closely related can be searched, the visual content in the presentation result of digital broadcasting can be enriched, and the viewing experience of users can be improved.
[0071] In some embodiments, one or more search keywords can be constructed according to the title of the material content, the aforementioned of scene keyword, sentence-level of keywords, etc., and content-related videos can be obtained through online full-web image / video search and offline image / video libraries. Next, the video content obtained by means such as a video topic segmentation algorithm can be segmented to obtain candidate visual material segments. It should be understood that when implementing the method of the present disclosure, video search and video segmentation can be performed in various ways to obtain candidate visual material segments, and are not limited here.
[0072] In some embodiments According to , for each scene of a plurality of scenes, the step S506 of searching for video materials related to the scene based on the material content and the target content corresponding to the scene can include extracting a scene keyword and searching for video materials related to the scene based on the scene keyword. Thereby, through the above method, video materials related to the entire scene can be obtained, and the available video materials can be enriched.
[0073] In some embodiments According toFor each scene of the plurality of scenes, step S506 of searching for video material related to the scene based on the material content and the target content corresponding to the scene can include extracting sentence-level keywords and searching for video material related to the scene based on the sentence-level keywords, thereby obtaining video material related to sentences in the target content and enriching the available video material.
[0074] In some embodiments, dynamic reports may be generated based on structured data within the target content, and the target content may be generated based on text based on deep learning models. from Images and / or videos Generate The method further enhances the richness of the video material.
[0075] In some embodiments, after target content, audio, and image and video material corresponding to the scene and target content are acquired, the target content is combined in a rendering generation stage. These visual materials to Text, Voice and Alignment do Specifically, the main focus is on pre-training models. Image text Based on the matching, the text and visual material are correlated and ordered, and for each text, the corresponding subtitle, video and image content that matches the TTS audio time period is found, and the audio-subtitle-video timeline is aligned. It is understood that the timeline can be adjusted by the user in this process to achieve manual alignment. want to .
[0076] In some examples According to ,The method for generating a digital human may further include aligning the retrieved video material with the target content based on sentence-level keywords. of Since the video material retrieved based on the keywords may correspond to a sentence in the target content, one segment of the target content can be analyzed at the sentence level. ofIt can accommodate multiple video materials searched based on keywords. this In such an embodiment, multiple sentence-level keywords within the target content are used to map each corresponding video material to each corresponding sentence and Alignment can be performed.
[0077] In one exemplary embodiment, in the example shown in Table 1, the sentence-level of The keywords are "X City Museum East Building," "City Wall Ruins," "XX Museum," "XX Ruins," "A Museum," and "B Natural History Museum." Searches are performed based on the sentence-level keywords corresponding to each sentence. after that , you can obtain the corresponding images and video material. Furthermore, you can also use sentence-level keywords The correspondence relationship between this sentence and multiple sentences, and the sentence-level keywords Based on the correspondence between the images and these video materials, the digital human Sentence reading when The video material is then displayed in the scene, and the video material is aligned with the sentences so that the video material can be switched completely in the gap between sentences. do It is possible.
[0078] It should be understood that if there is sufficient material in the original material content, or if it is determined that the digital human generation result does not need to be combined with the visual material, then material expansion may not be performed.
[0079] This completes all the preliminary work for digital human synthesis. Analysis Optionally, based on the results of the material supplementation, an appropriate digital human presentation method can be planned for each scene, thereby ensuring good interactivity of the final presented results and giving users a good immersive experience.
[0080] In some embodiments, it may be determined whether a digital human is triggered, i.e., whether a digital human is generated in the current scene. The key considerations for triggering are the scene's location in the video and the richness of the material corresponding to the scene. Here, material richness refers to the amount of highly correlated dynamic material present during the entire playback time of the scene. Occupancy rate for In addition to the above, the clarity, consistency, and correlation of the material and the scene are All These are the factors that determine whether a digital human is triggered. Based on these factors, the system can accept user-defined rules or automatically determine whether a digital human is triggered based on machine learning methods. Judgment This disclosure supports: Specific It should be understood that the trigger logic is not limited, and when implementing the method of the present disclosure, the corresponding trigger logic can be set in the manner described above, or a machine learning model can be trained using samples that satisfy the corresponding trigger logic, as needed, and this specification does not limit the scope of the present disclosure.
[0081] In some examples According to The method for generating a digital human further includes determining an occupancy rate of the play time required for a scene corresponding to the video material, and determining whether to trigger the digital human in the corresponding scene based on the occupancy rate, thereby determining whether to trigger the digital human based on the richness of the video material. whether It is possible to determine the following.
[0082] In some embodiments, a digital human determines that it is triggered. After Digital human attribute planning can be carried out based on a series of content features, such as the clothing, posture, Movement The content features include the semantic labels of the scenes (e.g., mood, emotion, semantic features, type of the target content, etc.), key trigger words in the broadcast content, visual elements, etc. Inside the materialIt can include features and the like. In a specific implementation, a rule-based Method is used to determine digital human attributes based on the features of the above content, and a deep learning method that uses content features as input to predict the digital human attribute composition is also supported to do .
[0083] In some embodiments, when a key trigger word is detected, the digital human corresponds to Posture , Movement , or create an expression, or the probability of doing so is a such that a mapping relationship between the key trigger word and the Posture , Movement , or expression of a specific digital human may be established. The rule-based Method may be used such that after the key trigger word is detected, the digital human will definitely to do react, and in order to obtain the relationship between the key trigger word and the Posture , Movement , or expression of a specific digital human, a large number of samples can be learned by the model You can do it like this , it should be understood that it is not limited herein.
[0084] In some embodiments, visual Inside the material content characteristics, such as the Clear of the material, consistency, and the relevance between the material and the scene, can also be used as content characteristics to be considered in digital human attribute planning When considering . Furthermore, specific content within the visual material can trigger specific Posture , Movement , expressions, etc. of the digital human. In some embodiments According toThe method of generating a digital human can further include determining the movement of the digital human based on the display position of the specific material in the video material in response to determining that the specific material is included in the video material. Thereby, by analyzing the content in the video material and combining the specific material of the video material with the movement of the digital human, the consistency among the video material, the target content, and the digital human can be further improved, and the viewing experience of the user can be improved.
[0085] In some embodiments, the special Deterministic element materials can include, for example, tables, drawings, picture-in-pictures, etc. In some embodiments, using the results of the chapter Analysis each scene can further obtain roles such as the summary paragraph and the abstract paragraph in the material content (the original text content or the image text content). From these informations, the lens operation, movement, and specific presentation forms (such as studio, picture-in-picture, etc.) of the digital human can be more richly controlled.
[0086] In some embodiments, the semantic label of the scene determined by the emotion Analysis by Obtained (for example, the mood and emotion of the scene) is the content feature considered in the digital human attribute plan When considering to be considered. According to some embodiments, for each scene of a plurality of scenes, the step S507 of constructing a digital human specific to the scene based on the scene label information can include constructing at least one of the clothing, expression, and movement of the digital human based on the semantic label. Can be considered
[0087] The semantic labels of the scenes may be used to determine the tone of the digital human. According to some embodiments, for each scene of the plurality of scenes, configuring a scene-specific digital human based on the scene label information in step S507 may include configuring a tone of the digital human voice based on the semantic labels. In some embodiments, the tone of the digital human voice may be, for example, Volume , pitch, intonation, etc. Furthermore, these semantic labels can have other uses during the composition of the digital human, for example, to compose a studio background with a more appropriate style for the scene.
[0088] In some embodiments, the digital human attributes include the digital human's clothing, Posture , Movement Attributes such as facial expression, background, etc. Specifically ,For each attribute plan, First Digital Human Attributes each option to to Manually set Can be done .for example" Clothing For the "attribute," there are different options such as "suit," "casual wear," and "shirt." Manual Based on these artificially given classification systems, we use classification algorithms in machine learning to Manually label On the training data, the model is trained by modeling various types of features, such as text features (word features, phrase features, sentence features, etc.), digital human ID, and digital human gender. is convergence do Until the model is artificially generated based on the features. Labeling Fitting the signal to do In the prediction stage, model predictions can be performed on the extracted features to obtain plans for different attributes.
[0089] In some embodiments, after the digital human generation result is obtained, the digital human generation result can be presented. Method for generating a digital human is It can further include presenting the digital human in the form of a holographic image. Thereby, a digital human generation result that can guarantee the consistency between the digital human and the target content is provided.
[0090] In some embodiments, after obtaining the video material and the digital human generation result, step S508 is executed to combine the video material and the digital human and perform video rendering to obtain the final video. As shown in FIG. 5, the method for generating a digital human can further include step S509 of presenting the digital human in the video in the form of Thereby, it is guaranteed that the digital human is consistent with the target content Can be Provide other digital human generation results, improve the vividness of the video generation result, improve the immersion of the user experience, and at the same time fully utilize the digital human Advantages Can make up for the shortage of related materials by utilizing the digital human and the video material. Further, the method of the present disclosure is designed for common scenes, has compatibility with different content types, and has a general type that can be applied in all fields. Interaction
[0091] In some embodiments, the present disclosure supports end-to-end automatic generation and also supports user interaction generation. That is, the user can adjust the generated video result. In the interactive generation scenario in The user can adjust elements of the video, such as text, audio, materials, avatar configuration, etc. The interactive generation method supports user modifications, thereby generating high-quality results. At the same time, the data generated by the user's interaction is also recorded and used as feedback data for system optimization to guide the learning of each step each Model and continuously improve the effectiveness of the system. of
[0092] According to another aspect of the present invention, a method for training a scene segmentation model is provided. As shown in FIG. 6, the training method includes: a step S601 of obtaining sample material content and a plurality of sample scenes in the sample material content; a step S602 of determining a plurality of predicted scenes from the sample material content based on a preset scene segmentation model; and a step S603 of adjusting parameters of the preset scene segmentation model based on the plurality of sample scenes and the plurality of predicted scenes to obtain a trained scene segmentation model. Thus, the scene segmentation model can be trained by the above method, and an accurate scene segmentation result can be output using the trained model. Thereby, a digital human can be generated using the scene segmentation result, and a final presentation result can be obtained, improving the user viewing experience. for The sample material content is the same as the material content obtained in step S201. of generation To perform final Typical to obtain a presentation result to do and the user viewing experience can be improved.
[0093] So that it can be understood , and the sample material content is the same as the material content obtained in step S201. to do The plurality of sample scenes in the sample material content are Manually, or obtained by dividing based on a template or a neural network model, and can be used as the actual result (ground truth) of dividing the sample material content. By performing training using the prediction result and the actual result, the trained scene segmentation model can be made capable of segmenting the material content into scenes. Method, sample material content When implementing the method of the present disclosure, a corresponding neural network can be selected as a scene segmentation model as needed and trained using a corresponding loss function, which is not limited herein. So that it can be understood In some embodiments, the preset scene segmentation model is a chapter can not limited here.
[0094] According to some embodiments, the preset scene segmentation model is a chapter SemanticThe step S602 of determining a plurality of predicted scenes from the content of the sample material based on the preset scene division model may include a chapter structure analysis model. Semantic Processing the sample material content using the segmentation model and the chapter structure analysis model to determine a plurality of predicted subtopics in the material content and predicted structural relationships between the plurality of predicted subtopics; pre and dividing the plurality of predicted subtopics into a plurality of predicted scenes based on their scalar structure relationships. Semantic By training a segmentation model and a chapter structure analysis model, the trained model can accurately extract semantically complete subtopics from the material content and the structural relationships between them. Determination By utilizing the structural relationships between the subtopics (e.g., ensemble structure, parallel structure, incremental structure, etc.), multiple subtopics can be further divided into multiple scenes suitable for digital human broadcasting tasks.
[0095] As can be seen, in addition to the scene segmentation model, other models may be used in the methods of FIG. 1 or FIG. 5, such as models for segmenting subtopics into scenes, models for generating target content (e.g., models for text rewriting, text compression, and / or style transfer), and models for scene emotion. Analysis It is also possible to train models such as models for , scene keyword / sentence level keyword extraction models, digital human attribute planning models, etc. Labeling Corpora or user interaction data can be used to train models, including scene segmentation models.
[0096] According to another aspect of the present invention, an apparatus for generating a digital human is provided. As shown in Figure 7, the apparatus for generating a digital human 700 includes a first acquisition unit 702 configured to acquire material content, and a pre-trained scene segmentation model for determining a plurality of scenes from the material content, where each scene of the plurality of scenes is a complete segment of the material content.Semantic A first determination unit 704 corresponding to a content segment having information, a second determination unit 706 configured to determine target content corresponding to each scene of a plurality of scenes based on the corresponding content segment, and a third determination unit 708 configured to determine scene label information of the scene based on the corresponding target content, and a digital human configuration unit 710 configured to configure a digital human that specifies the scene based on the scene label information. The units 702 to 710 of the apparatus 700 Operation are similar to those in steps S201 to S205 of FIG. 2 Operation and will not be described here for the sake of understanding.
[0097] According to some embodiments, the material content can include at least one of image data and video data and text data. In some embodiments According to , the first acquisition unit 702 can be further configured to acquire the material content based on at least one of a method of acquiring the material content based on a web page address and a method of acquiring the material content based on a search keyword.
[0098] In some embodiments According to , as shown in FIG. 8, the first determination unit 800 includes a first determination subunit 802 configured to perform chaptering on the material content to determine a plurality of subtopics from the material content and determine a structural relationship between the plurality of subtopics, and a first division subunit 804 configured to divide the plurality of subtopics into a plurality of scenes based on the structural relationship.
[0099] According to some embodiments, as shown in FIG. 9 , the second determining unit 900 includes at least one of a first updating subunit 902, which performs at least one of text rewriting and text compression on the corresponding content segments to update the corresponding content segments, and a second updating subunit 908, which performs at least one of text rewriting and text compression on the converted content to update the corresponding content. 1 It may include one.
[0100] In some examples As shown in FIG. 9, the second determination unit 900 determines the structure between the scene and the previous scene. The scene processing unit 902 may include a generating subunit 904 configured to generate first content for the scene based on the relationship.
[0101] In some examples ,As shown in FIG. 9, the second determining unit 900 may include a transforming subunit 906 configured to transform the corresponding content segment into the corresponding target content based on a pre-trained style transfer model, where the style transfer model is: It is obtained through training based on learning.
[0102] In some examples The digital human generating apparatus 700 generates a key image for each of the plurality of scenes from the target content corresponding to the scene. The video processing system may further include an extraction unit configured to extract expression information, and a generation unit configured to generate supplementary material for the video based on the key-value format information.
[0103] According to some embodiments, the third determining unit 708 may perform sentiment analysis on the corresponding target content to obtain semantic labels. Emotions that are structured to It may contain subunits.
[0104] According to some embodiments, the semantic labels are used to identify the positive, neutral, or negative sentiment expressed by the corresponding target content. In some examples 10, the apparatus 1000 for generating a digital human may include a voice conversion unit 1010 configured to convert target content into voice for digital human broadcasting. are the units 702 to 710 of the device 700, respectively. and will not be described here.
[0105] In some examples As shown in FIG. 10, the apparatus 1000 for generating a digital human includes: for each scene of a plurality of scenes, a search unit 1012 configured to search for video material related to the scene based on material content and target content corresponding to the scene; and a combination unit 1016 configured to combine the video material with the digital human.
[0106] In some examples The search unit 1012 may include a first extraction subunit configured to extract scene keywords and a first search subunit configured to search for video material related to the scene based on the scene keywords.
[0107] In some examples , the search unit 1012 may further include a second extraction subunit configured to extract sentence-level keywords, and a second search subunit configured to search for video material related to the scene based on the sentence-level keywords.
[0108] In some examples The digital human generating device 1000 aligns the retrieved video material with the target content based on sentence-level keywords. The alignment unit may include an alignment unit configured to:
[0109] In some examples , the apparatus 1000 for generating a digital human may further include a fifth determination unit configured to determine an occupancy rate of the playback time required for a scene corresponding to the video material, and a sixth determination unit configured to determine whether to trigger the digital human in the corresponding scene based on the occupancy rate.
[0110] In some examples The apparatus 1000 for generating a digital human may further include a fourth determination unit configured to determine a movement of the digital human based on a display position of the specific material in the video material in response to determining that the specific material is included in the video material.
[0111] According to some embodiments, the digital human construction unit 1014 may include a first construction subunit configured to construct at least one of an outfit, an expression, and a movement of the digital human based on the semantic label.
[0112] In some examples , the digital human construction unit 1014 may further include a second construction subunit configured to construct a tone of the digital human voice based on the semantic label.
[0113] In some examples , the apparatus 1000 for generating a digital human may further include a holographic image presenting unit configured to present the digital human in the form of a holographic image.
[0114] In some examples The apparatus 1000 for generating a digital human can further include a video presentation unit 1018 configured to present the digital human in the form of a video.
[0115] According to another aspect of the present invention, a training apparatus for a scene segmentation model is provided. As shown in FIG. 11, the training apparatus 1100 includes a second acquisition unit 1102 configured to acquire sample material content and a plurality of sample scenes in the sample material content, a seventh determination unit 1104 configured to determine a plurality of predicted scenes from the sample material content based on a preset scene segmentation model, and a training unit 1106 configured to adjust the parameters of the preset scene segmentation model based on the plurality of sample scenes and the plurality of predicted scenes to obtain a trained scene segmentation model.
[0116] According to some embodiments, the preset scene segmentation model can include a chapter segmentation model and a chapter structure analysis model. The seventh determination unit 1104 can include a second determination sub-unit configured to process the sample material content using the chapter segmentation model and the chapter structure analysis model to determine a plurality of predicted sub-topics in the material content and a predicted structural relationship between the plurality of predicted sub-topics, and a second segmentation sub-unit configured to segment the plurality of predicted sub-topics into a plurality of predicted scenes based on the predicted structural relationship.
[0117] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision and disclosure of related user personal information all comply with the provisions of relevant laws and regulations and do not violate public order and good customs. According to the embodiments of the present disclosure, an electronic device, a readable storage medium and a computer program product are further provided.
[0118] Next, with reference to FIG. 12, a block configuration diagram of an electronic device 1200 operable as a server or a client of the present disclosure, which is an example of a hardware device applicable to each aspect of the present disclosure, will be described. The electronic device represents various forms of digital electronic computers, such as laptop computers, desktop computers, tablets, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may further represent various forms of mobile devices, such as personal digital processors, mobile phones, smartphones, wearable devices, and other similar computing devices. The components shown in this specification, their connection relationships, and their functions are merely exemplary and do not limit the implementation of the present disclosure described and / or claimed in this specification.
[0119] As shown in FIG. 12, the electronic device 1200 includes a computing unit 1201, which can execute various appropriate operations and processes by a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. In the RAM 1203, various programs and data necessary for operating the electronic device 1200 may be further stored. The computing unit 1201, the ROM 1202, and the RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0120] The components of the electronic device 1200 are connected to an I / O interface 1205, which includes an input unit 1206, an output unit 1207, a storage unit 1208, and a communication unit 1209. The input unit 1206 may be any type of device capable of inputting information into the electronic device 1200. The input unit 1206 can input numeric or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackboard, trackball, joystick, microphone, and / or remote control. The output unit 1207 may be any type of device capable of presenting information, and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 1208 may include, but is not limited to, a magnetic disk or an optical disk. The communications unit 1209 enables the electronic device 1200 to exchange information / data with other devices via a computer network, e.g., the Internet, and / or various telecommunications networks, and may include, but is not limited to, a modem, a network card, an infrared communications device, a wireless communications transceiver, and / or a chipset, e.g., a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, a cellular communications device, and / or the like.
[0121] The computing unit 1201 may be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 1201 may include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes each method and process described above, for example, the method for generating a digital human and the method for training a scene segmentation model. For example, in some embodiments, the method for generating a digital human and the method for training a scene segmentation model may be implemented as a computer software program and tangibly included in a machine-readable medium, such as the storage unit 1208. In some embodiments, part or all of the computer program may be loaded and / or installed into the electronic device 1200 via the ROM 1202 and / or the communication unit 1209. When the computer program is loaded into the RAM 1203 and executed by the computing unit 1201, one or more steps of the method for generating a digital human and the method for training a scene segmentation model described above can be executed. Alternatively, in another embodiment, the computing unit 1201 may be configured to execute the method for generating a digital human and the method for training a scene segmentation model in any other suitable manner (e.g., by firmware).
[0122] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments are 1Implemented in one or more computer programs, the 1 One or more computer programs may be executed and / or interpreted in a programmable system including at least 1 Two programmable processors. The programmable processors may be dedicated or general-purpose programmable processors, and may receive data and instructions from a memory system, at least 1 Two input devices, at least 1 Two output devices, and may send the data and instructions to the memory system, the at least 1 Two input devices, the at least 1 Two output devices.
[0123] The program code for implementing the method of the present disclosure may 1 Be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program codes are executed by the processor or controller, the functions / operations defined in the flowchart and / or block diagram are implemented. The program codes may be executed entirely by a machine, partially by a machine, partially by a machine as an independent software package and partially by a remote machine, or entirely by a remote machine or server.
[0124] In the context of the present disclosure, the machine-readable medium may be a tangible medium, and may comprise or store a program used in or coupled to an instruction execution system, apparatus or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium are 1This includes electrical connections by one or more leads, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.
[0125] To provide for user interaction, a computer may implement the systems and techniques described herein and include a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to a user, and a keyboard and pointing device (e.g., a mouse or trackball) through which a user may provide input to the computer. Other types of devices may also be used to provide for user interaction. For example, feedback provided to a user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or haptic feedback), and input from a user may be received in any form (including sound, speech, or tactile input).
[0126] The systems and techniques described herein may be implemented in a computing system (e.g., as a data server) that includes backend members, or in a computing system (e.g., an application server) that includes middleware members, or in a computing system (e.g., a user computer having a graphical user interface and a web browser through which a user can realize interactions with embodiments of those systems and techniques) that includes frontend members, or in a computing system consisting of any combination of those backend members, middleware members, or frontend members. The members of the system may be interconnected by digital data communication of any form or medium (e.g., a communication network). An example of a communication network includes a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0127] The computer system may include clients and servers. Clients and servers are generally far apart from each other and usually interact via a communication network. By using a computer corresponding to a computer program having a relationship of client and server with each other, a relationship of client and server is generated. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain. It should be understood that the steps may be re-ranked, increased, or deleted using the various forms of flow described above. For example, each step described in this disclosure may be executed in parallel, sequentially, or in a different order, and the text is not limited to this as long as the technical solutions disclosed in this disclosure can achieve the desired results.
[0128]
[0129] Embodiments or examples of the present disclosure have been described with reference to the drawings. However, the above methods, systems, and apparatuses are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples. It should be understood that the scope is limited only by the scope of the claims after authorization and their equivalents. Various elements of the embodiments or examples may be omitted or replaced by their equivalent elements. Note that each step may be executed in an order different from the order described in the present disclosure. Furthermore, various elements of the embodiments or examples may be combined in various ways. Importantly, with the evolution of technology, many of the elements described here can be replaced by equivalent elements that appear after the present disclosure.
Claims
1. A method for generating a digital human, comprising: obtaining the content of the material; determining a plurality of scenes from the material content based on a pre-trained scene segmentation model, where each scene of the plurality of scenes corresponds to a content segment having one complete semantic information of the material content; for each scene of the plurality of scenes, determining the target content corresponding to the scene based on the corresponding content segment; determining the scene label information of the scene based on the corresponding target content; and constructing a digital human specific to the scene based on the scene label information. A method for generating a digital human.
2. Obtaining the content of the material includes obtaining the material content based on at least one of a method for obtaining the material content based on a web page address and a method for obtaining the material content based on a search keyword. The method according to claim 1.
3. The method according to claim 1 or 2, wherein the material content includes at least one of image data and video data and text data.
4. Determining a plurality of predicted scenes from the content of the material based on a pre-trained scene segmentation model includes determining a plurality of sub-topics from the material content by performing chapter structure analysis and chapter semantic segmentation on the material content, and determining the structural relationship between the plurality of sub-topics; and dividing the plurality of sub-topics into the plurality of scenes based on the structural relationship. The method according to any one of claims 1 to 3.
5. For each scene of the plurality of scenes, determining the target content corresponding to the scene based on the corresponding content segment includes generating a first content for the scene based on the structural relationship between the scene and the previous scene. The method according to claim 4.
6. For each scene of the plurality of scenes, determining the target content corresponding to the scene based on the corresponding content segment includes Converting the corresponding content segment into the corresponding target content based on a pre-trained style conversion model, where the style conversion model is obtained by training based on prompt learning. The method according to claim 4 or 5.
7. For each scene of the plurality of scenes, determining the target content corresponding to the scene based on the corresponding content segment includes: Performing at least one of text rewriting and text compression on the corresponding content segment to update the corresponding content segment; The method according to claim 6, further including at least one of performing at least one of text rewriting and text compression on the converted target content to update the corresponding target content.
8. The scene label information includes semantic labels. Here, for each scene in the plurality of scenes, determining the scene label information of the scene based on the corresponding target content includes: The method according to any one of claims 1 to 7, including performing sentiment analysis on the corresponding target content to obtain the semantic label.
9. The method according to claim 8, wherein the semantic label is used to identify the positive, neutral, and negative emotions expressed by the corresponding target content.
10. For each scene of the plurality of scenes, constructing a digital human specific to the scene based on the label information includes: The method according to claim 8 or 9, including constructing at least one of the clothing, expression, and movement of the digital human based on the semantic label.
11. The method according to claim 10, further including converting the target content into audio for digital human broadcasting.
12. For each scene of the plurality of scenes, constructing a digital human specific to the scene based on the scene label information includes: The method according to claim 11, further including constructing the tone of the digital human voice based on the semantic label.
13. The method according to any one of claims 1 to 12, further comprising presenting the digital human in the form of a holographic image.
14. The method according to any one of claims 1 to 12, further comprising presenting the digital human in video format.
15. For each scene of a plurality of scenes, searching for video materials related to the scene based on the material content and the target content corresponding to the scene; The method according to claim 14, further comprising combining the video material and the digital human.
16. For each scene of the plurality of scenes, searching for video materials related to the scene based on the material content and the target content corresponding to the scene includes: extracting scene keywords; The method according to claim 15, further comprising searching for video materials related to the scene based on the scene keywords.
17. For each scene of the plurality of scenes, searching for video materials related to the scene based on the material content and the target content corresponding to the scene includes: extracting keywords at the sentence level; The method according to claim 15 or 16, further comprising searching for video materials related to the scene based on the keywords at the sentence level.
18. The method according to claim 17, further comprising aligning the retrieved video material and the target content based on the keywords at the sentence level.
19. The method according to any one of claims 15 to 18, further comprising determining the movement of the digital human based on the display position of the specific material in the video material in response to determining that the specific material is included in the video material.
20. For each scene of a plurality of scenes, extracting information in key-value format from the target content corresponding to the scene; The method according to any one of claims 14 to 19, further comprising generating auxiliary materials for the video based on the information in key-value format.
21. determining the occupancy rate of the playback time required for the scene corresponding to the video material; The method according to any one of claims 15 to 20, further comprising determining whether to trigger the digital human in the corresponding scene based on the occupancy rate.
22. A method for training a scene segmentation model, comprising: obtaining sample material content and a plurality of sample scenes in the sample material content; determining a plurality of predicted scenes from the sample material content based on a preset scene segmentation model; adjusting parameters of the preset scene segmentation model based on the plurality of sample scenes and the plurality of predicted scenes, and obtaining a trained scene segmentation model.
23. The preset scene segmentation model includes a chapter meaning segmentation model and a chapter structure analysis model. Here, determining a plurality of predicted scenes from the sample material content based on the preset scene segmentation model includes: processing the sample material content using the chapter meaning segmentation model and the chapter structure analysis model to determine a plurality of predicted subtopics in the material content and a predicted structural relationship between the plurality of predicted subtopics; dividing the plurality of predicted subtopics into the plurality of predicted scenes based on the predicted structural relationship. The training method according to claim 22.
24. An apparatus for generating a digital human, the apparatus comprising: a first acquisition unit configured to acquire material content; a first determination unit configured to determine a plurality of scenes from the material content based on a pre-trained scene segmentation model, where each of the plurality of scenes corresponds to a content segment having one complete semantic information of the material content; a second determination unit configured to determine target content corresponding to each scene of the plurality of scenes based on the corresponding content segment; a third determination unit configured to determine scene label information of the scene based on the corresponding target content; a digital human construction unit configured to construct a digital human specific to the scene based on the scene label information. An apparatus for generating a digital human.
25. The first acquisition unit further includes: a method for obtaining the material content based on a web page address; The apparatus according to claim 24, configured to obtain the material content based on at least one of the methods for obtaining the material content based on a search keyword.
26. The apparatus according to claim 24 or 25, wherein the material content includes at least one of image data and video data and text data.
27. The first determination unit is configured to determine a plurality of sub-topics from the material content and determine a structural relationship between the plurality of sub-topics by performing chapter structure analysis and chapter meaning segmentation on the material content, and a first sub-determination unit; The apparatus according to any one of claims 24 to 26, further comprising a first split sub-unit configured to split the plurality of sub-topics into the plurality of scenes based on the structural relationship.
28. The second determination unit The apparatus according to claim 27, further comprising a generation sub-unit configured to generate first content for the scene based on a structural relationship between the scene and a previous scene.
29. The second determination unit The apparatus according to claim 27, further comprising a conversion sub-unit configured to convert the corresponding content segment into the corresponding target content based on a pre-trained style conversion model, where the style conversion model is obtained by training based on prompt learning.
30. The second determination unit a first update sub-unit configured to perform at least one of text rewriting and text compression on the corresponding content segment to update the corresponding content segment; The apparatus according to claim 29, further comprising at least one of a second update sub-unit configured to perform at least one of text rewriting and text compression on the converted target content to update the corresponding target content.
31. The third determination unit The apparatus according to any one of claims 24 to 30, further comprising a sentiment analysis sub-unit configured to perform sentiment analysis on the corresponding target content to obtain the semantic label.
32. The apparatus according to claim 31, wherein the semantic label is used to identify positive, neutral, and negative emotions represented by the corresponding target content.
33. The digital human composition unit is The apparatus according to claim 31 or 32, further comprising a first composition subunit configured to compose at least one of clothing, expression, and movement of the digital human based on the semantic label.
34. The apparatus according to claim 33, further comprising an audio conversion unit configured to convert the target content into audio for the digital human broadcast.
35. The digital human composition unit is The apparatus according to claim 34, further comprising a second composition subunit configured to compose the tone of the digital human voice based on the semantic label.
36. The apparatus according to any one of claims 24 to 35, further comprising a holographic image presentation unit configured to present the digital human in the form of a holographic image.
37. The apparatus according to any one of claims 24 to 35, further comprising a video presentation unit configured to present the digital human in the form of a video.
38. For each scene of the plurality of scenes, a search unit configured to search for video materials related to the scene based on the material content and the target content corresponding to the scene, and The apparatus according to claim 37, further comprising a combining unit configured to combine the video materials with the digital human.
39. The search unit is A first extraction subunit configured to extract scene keywords, and The apparatus according to claim 38, further comprising a first search subunit configured to search for video materials related to the scene based on the scene keywords.
40. The search unit is A second extraction subunit configured to extract keywords at the sentence level, and The apparatus according to claim 38 or 39, further comprising a second search subunit configured to search for video materials related to the scene based on the keywords at the sentence level.
41. The apparatus according to claim 40, further comprising an alignment unit configured to align the searched video materials with the target content based on the keywords at the sentence level.
42. The apparatus according to any one of claims 38 to 41, further comprising a fourth determination unit configured to determine the movement of the digital human based on the display position of the specific material in the video material in response to determining that the specific material is included in the video material.
43. An extraction unit configured to extract information in a key-value format from target content corresponding to each of the plurality of scenes; The apparatus according to any one of claims 37 to 42, further comprising a generation unit configured to generate auxiliary material for the video based on the information in the key-value format.
44. A fifth determination unit configured to determine an occupancy rate of a playback time required for a scene corresponding to the video material; The apparatus according to any one of claims 38 to 43, further comprising a sixth determination unit configured to determine whether to trigger the digital human in the corresponding scene based on the occupancy rate.
45. An apparatus for training a scene segmentation model, comprising: A second acquisition unit configured to acquire sample material content and a plurality of sample scenes in the sample material content; A seventh determination unit configured to determine a plurality of predicted scenes from the sample material content based on a preset scene segmentation model; A training unit configured to adjust parameters of the preset scene segmentation model based on the plurality of sample scenes and the plurality of predicted scenes to obtain a trained scene segmentation model.
46. The preset scene segmentation model includes a chapter meaning segmentation model and a chapter structure analysis model. Here, the seventh determination unit includes: A second determination subunit configured to process the sample material content using the chapter meaning segmentation model and the chapter structure analysis model to determine a plurality of predicted subtopics in the material content and a predicted structural relationship between the plurality of predicted subtopics; The training apparatus according to claim 45, further comprising a second segmentation subunit configured to divide the plurality of predicted subtopics into the plurality of predicted scenes based on the predicted structural relationship.
47. An electronic device, comprising: At least one processor; including a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by at least one processor, the instructions being executed by the at least one processor such that the at least one processor can execute the method according to any one of claims 1 to 23. An electronic device. **Claim 48** A non-transitory computer-readable storage medium storing computer instructions, the computer instructions causing the computer to execute the method according to any one of claims 1 to 23. A non-transitory computer-readable storage medium. **Claim 49** A computer program product including a computer program, the computer program, when executed by a processor, executing the method according to any one of claims 1 to 23. A computer program product.
Citation Information
Patent Citations
Device and method for automatically creating cartoon image based on input sentence
EP3961474A2
Method and device for generating explanatory sentence of video content, method and device for programming digest video, and computer-readable recording medium with program for making computer implement the same methods recorded thereon
JP2001275058A
Speech synthesizer and speech synthesis program
JP2005181840A
Sign language CG generation device and program of the same
JP2015230640A
Dialogue system, dialogue apparatus and computer program therefor
JP2018156272A