Video generation method and device, equipment, storage medium and program product

By adjusting the digital person's image based on user response information and generating personalized target videos, the problem of insufficient digital person's interaction experience is solved, and more natural user interaction and personalized replies are achieved.

CN120343360APending Publication Date: 2025-07-18BAIDU COM TIMES TECH (BEIJING) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510685664.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, digital human interaction experience is insufficient, and it is difficult to personalize the user's reply content, scenes and themes, resulting in unnatural user interaction.

Method used

By determining the target topic based on the user's reply information, the initial digital person is image-adjusted, the target video is generated, the target digital person is used to provide the user with reply information, and the language style and action style adjustments are combined to improve the digital person's interactive experience.

Benefits of technology

It realizes personalized adaptation to the image of digital people, improves the user's interactive experience, and allows digital people to interact with users more naturally, and provides personalized replies and scene adaptation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343360A_ABST
    Figure CN120343360A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, electronic equipment, a computer readable storage medium and a computer program product, and relates to the technical field of artificial intelligence such as computer vision, digital human technology and intelligent questions and answers. A specific embodiment of the method comprises the steps of determining a target theme based on reply information to be provided for a user; performing image adjustment on the initial digital person based on the target theme to obtain a target digital person; and generating a target video for providing the reply information to the user by using the target digital person based on the reply information and the target digital person. Therefore, the image of the digital person can be adjusted according to the reply information provided for the user, so that the image of the digital person can adapt to the reply content, the reply scene and the reply theme in a personalized manner, and the digital person interaction experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, specifically to artificial intelligence technologies such as computer vision, digital human technologies, intelligent question answering, etc., and particularly to methods, devices, electronic devices, computer-readable storage media, and computer program products for generating videos. Background Art

[0002] With the development of computer technologies, in order to provide users with a better interaction experience and enable users to have an experience close to interacting with real humans during human-computer interaction, Digital Human technology has emerged as the times require.

[0003] A digital human can be understood as a virtual character created through digital technologies and similar to the human image. Based on the simulation and reproduction of human appearance, behavior, and interaction, the "digital human" can provide users with an interaction experience close to communicating and interacting with real humans. Thus, in such a context, how to further improve users' interaction and usage experience with digital humans is worthy of attention and an urgent need. Summary of the Invention

[0004] Embodiments of the present disclosure propose a method, device, electronic device, computer-readable storage media, and computer program product for generating videos.

[0005] In a first aspect, embodiments of the present disclosure propose a method for generating a video, including: determining a target theme based on reply information to be provided to a user; adjusting the image of an initial digital human based on the target theme to obtain a target digital human; generating a target video that uses the target digital human to provide the reply information to the user based on the reply information and the target digital human.

[0006] In a second aspect, embodiments of the present disclosure propose a device for generating a video, including: a theme determination unit configured to determine a target theme based on reply information to be provided to a user; a digital human determination unit configured to adjust the image of an initial digital human based on the target theme to obtain a target digital human; a video generation unit configured to generate a target video that uses the target digital human to provide the reply information to the user based on the reply information and the target digital human.

[0007] In a third aspect, embodiments of the present disclosure provide an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the method for generating a video described in any implementation manner of the first aspect.

[0008] Fourthly, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to enable a computer to execute a method for generating a video as described in any implementation manner of the first aspect when executed.

[0009] Fifthly, an embodiment of the present disclosure provides a computer program product including a computer program, and the computer program can implement the method for generating a video as described in any implementation manner of the first aspect when executed by a processor.

[0010] The method, apparatus, electronic device, computer-readable storage medium and computer program product for generating a video provided by the embodiments of the present disclosure determine a target theme based on the reply information to be provided to the user; then, adjust the image of the initial digital human based on the target theme to obtain a target digital human; finally, generate a target video for providing the reply information to the user by using the target digital human based on the reply information and the target digital human.

[0011] The present disclosure can adjust the image of the digital human according to the reply information provided to the user, so that the image of the digital human can be personalized to adapt to the reply content, reply scenario and reply theme, and improve the user's digital human interaction experience.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0013] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects and advantages of the present disclosure will become more obvious:

[0014] Figure 1 is an exemplary system architecture to which the present disclosure can be applied;

[0015] Figure 2 is a flowchart of a process for generating a video provided by an embodiment of the present disclosure;

[0016] Figure 3 is a flowchart of a process for generating reply information provided by an embodiment of the present disclosure;

[0017] Figure 4 is a flowchart of a process for generating a video implemented in a specific application scenario provided by an embodiment of the present disclosure;

[0018] Figure 5 is a structural block diagram of an apparatus for generating a video provided by an embodiment of the present disclosure;

[0019] Figure 6A schematic structural diagram of an electronic device suitable for implementing a method for generating a video provided by an embodiment of the present disclosure. Detailed implementation manners

[0020] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0021] In addition, in the technical solutions involved in the present disclosure, the acquisition, storage, use, processing, transportation, provision, and disclosure of user personal information (such as the "question information" provided by the user involved in the following of the present disclosure, etc.) comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0022] Figure 1 An exemplary system architecture 100 showing embodiments of a method, apparatus, electronic device, and computer-readable storage medium for generating a video to which the present disclosure can be applied is shown.

[0023] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0024] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various applications for implementing information communication between the two may be installed on the terminal devices 101, 102, 103 and the server 105, such as online Q&A applications, live interaction applications, instant messaging applications, etc.

[0025] The terminal devices 101, 102, 103 and the server 105 can be either hardware or software. When the terminal devices 101, 102, 103 are hardware, they can be various electronic devices with a display screen, including but not limited to smartphones, tablet computers, laptop portable computers, desktop computers, etc.; when the terminal devices 101, 102, 103 are software, they can be installed in the above-listed electronic devices, and can be implemented as multiple software or software modules, or can be implemented as a single software or software module, and no specific limitation is made here. When the server 105 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server; when the server is software, it can be implemented as multiple software or software modules, or can be implemented as a single software or software module, and no specific limitation is made here.

[0026] The server 105 can provide various services through various built-in applications. Taking an online Q&A application that can use a "digital human" to provide Q&A services as an example, when the server 105 runs this online Q&A application, the following effects can be achieved: First, the server 105 can interact with the terminal devices 101, 102, 103 used by the user (not shown in the figure) through the network 104 to collect the question information provided by the user. Then, the server 105 can determine the reply information to be provided to the user for this question information, and determine the target topic based on this reply information; next, the server 105 adjusts the image of the initial digital human based on the target topic to obtain the target digital human; finally, the server 105 generates a target video that uses the target digital human to provide the reply information to the user based on the reply information and the target digital human.

[0027] Correspondingly, the server 105 can subsequently use this target video to actually provide a "reply" to the user to implement and complete the "online Q&A" service.

[0028] Since storing the initial image of the digital human, adjusting the image of the digital human, and generating the target video often require a large amount of computing resources and strong computing power, the method for generating a video provided in subsequent embodiments of the present disclosure is generally executed by the server 105 with strong computing power and a large amount of computing resources. Correspondingly, the device for generating a video is generally also set in the server 105. However, it should also be noted that when the terminal devices 101, 102, and 103 also have computing power and computing resources that meet the requirements, the terminal devices 101, 102, and 103 can also complete the above operations originally performed by the server 105 through the online Q&A applications installed thereon, and then output the same results as the server 105. Especially in the case where there are multiple terminal devices with different computing capabilities at the same time, when the online Q&A application determines that the terminal device where it is located has strong computing power and a large amount of remaining computing resources, the terminal device can be allowed to execute the above operations, thereby appropriately reducing the computing pressure on the server 105. Correspondingly, the device for generating a video can also be set in the terminal devices 101, 102, and 103. In this case, the exemplary system architecture 100 may not include the server 105 and the network 104.

[0029] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0030] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 2 , Figure 2 is a flowchart of a process for generating a video provided by an embodiment of the present disclosure, which includes process 200.

[0031] Process 200 specifically includes the following steps:

[0032] Step 201: Determine the target theme based on the reply information to be provided to the user;

[0033] In the embodiments of the present disclosure, this step aims to have the execution entity of the method for generating a video (such as Figure 1 the server 105 shown) first obtain the reply information that will be provided to the user. Or rather, the reply information can also be understood as the "reply content" or "answer" for answering the question information provided by the user.

[0034] For example, when the content of the user's question information is to inquire about the introduction of drug A or knowledge related to drug A, the reply information can be the "introduction" of drug A. For example, after receiving the question information, the execution entity can use the content stored and recorded in the pre-configured expert knowledge base to correspondingly extract the "introduction" of drug A as the "reply" or "answer" to provide the reply information corresponding to the question information.

[0035] Also for example, when the question information requests the execution entity to provide a travel plan, the reply information can be a travel plan reply information. In the travel plan reply information, corresponding and specific travel plans can be provided for the user according to the specific requirements put forward by the user (such as destination, budget, number of days, departure time, etc.) as the "answer". For example, when the reply information is a travel plan reply information, its specific content can be exemplified as: Travel to City B for 7 days. On the first day, take Train No. XX from the departure place to City B. On the second day, visit Scenic Spot c at x1 in the morning, and then after finishing the visit, visit Scenic Spot d at x2... On the seventh day, go to City B Railway Station and take Train No. YY back to the departure place.

[0036] It should be noted that the "reply information" can be generated locally by the above-mentioned execution entity and can be obtained from the local storage device, or can be obtained by the execution entity from a non-local storage device (such as Figure 1 the terminal devices 101, 102, 103 shown). The local storage device can be a data storage module set in the above-mentioned execution entity, such as a server hard disk. In this case, the reply information can be quickly read locally; the non-local storage device can also be any other electronic device set for storing data, such as some user terminals. In this case, the above-mentioned execution entity can send a retrieval command to this electronic device to obtain the required reply information (in such a case, the "reply information" can have been generated in advance for the user's question information, and the execution entity only directly uses such reply information to complete the subsequent video generation process).

[0037] After obtaining the reply information, the execution entity can parse it to determine the corresponding target topic. The target topic is actually associated with the specific content of the reply information and is determined with reference to a certain classification standard. For example, after pre-classifying and determining multiple topics according to the "classification standard", the execution entity can use the "content", "semantics" corresponding to each topic, as well as the "semantics" and "content" of the reply information to correspondingly determine the reply information to the corresponding topic to obtain the target topic.

[0038] For example, the division criteria for the target topic can be determined according to the field involved in the response information. For example, for the response information in the medical field, the target topic can be exemplarily the "medical topic" corresponding to the dimension of "field".

[0039] Accordingly, after determining the target topic, in the next step such as step 202, the execution entity can determine the adjustment direction when adjusting the image of the (initial) digital human based on the "target topic", so that the image of the adjusted target digital human can fit the content and theme of the response information and improve the presentation effect of the digital human. The process of this image adjustment will be described in detail in the discussion related to step 202 below.

[0040] In some embodiments, the above "division criteria" can be configured in a more granular manner, so that the correspondence between the target topic and the content in the response information is higher. For example, when the response information involves movie-related content, the target topic can not only be divided into "movie" and "movie commentary", but the division criteria can actually correspond to, for example, the name and theme of the movie explained, described, and recommended in the response information. For example, when the response information is an explanatory and recommended copy for Movie Z, the target topic can be both the "movie commentary" based on the "field" as discussed above, and can also be more granularly and further determined as "Movie Z".

[0041] After determining the division criteria and strategies for the "target topic" according to actual needs, in some embodiments, a topic (division) model can be encapsulated and trained based on such division criteria, so that the execution entity can determine the target topic of the response information by calling the subject model and using such division criteria and strategies.

[0042] For example, based on natural language processing (Neuro-Linguistic Programming, abbreviated as NLP) technology, topic models (Topic Modeling) such as Latent Dirichlet Allocation (abbreviated as LDA) and Non-negative Matrix Factorization (abbreviated as NMF) can be used to determine the topic of the response information. For example, based on the pre-determined "division criteria" and "granularity", sample response information and corresponding sample topics can be configured to train the topic model to obtain a trained topic model. This trained topic model can have the ability to determine and divide topics with corresponding granularity.

[0043] Then, during the execution of this step, the execution entity can process the response information by calling and using this trained topic model to determine and obtain the corresponding target topic.

[0044] In some alternative implementation manners of this embodiment, in this step, when the execution entity determines the target theme based on the reply information to be provided to the user, it may choose to determine the target theme based on the associated city included in the reply information to be provided to the user. That is, when determining the target theme, the execution entity may choose to refer to and use the associated city included in the reply information as a benchmark to determine the target theme. Correspondingly, such a target theme is actually associated with the associated city.

[0045] For example, when the reply information is travel planning reply information, the tourist city (i.e., the associated city) involved in the travel planning reply information may be used as the target theme. Thus, when providing the travel planning reply information subsequently, the image of the digital human can be adjusted in combination with the regional characteristics of the "associated city".

[0046] Thereby, not only can the efficiency of determining the target theme be improved, but also the "city" and "region" can be used as the target theme, so that the image of the digital human can be adjusted from the dimension of "regional characteristics" subsequently to improve the user's viewing experience.

[0047] Step 202: Adjust the image of the initial digital human based on the target theme to obtain a target digital human;

[0048] In the embodiment of the present disclosure, as discussed above, this step aims to enable the above-mentioned execution entity to use the target theme determined in the above step 201 to adjust the image of the initial digital human on the basis of step 201 to obtain a target digital human whose image conforms to the "target theme".

[0049] In some embodiments, the initial digital human can be simply understood as a pre-constructed digital human template. Subsequently, the execution entity can edit the "digital human template" by changing at least some of the visual style parameters based on this "digital human template" to complete at least partial image adjustment of the initial digital human and obtain a target digital human. Correspondingly, the (target) theme can be configured in advance based on the parameters of each item and each part that are allowed to be adjusted in the digital human template, so that after the target theme is determined, the corresponding parameter included under the target theme can be used to modify the digital human template correspondingly to complete the image adjustment.

[0050] For example, the executing entity can change the hairstyle, clothing, height, etc. of the initial digital human by changing the visual style parameters, so that the initial digital human has different images. Correspondingly, the target theme can correspondingly include parameters such as "hairstyle", "clothing", "height", etc. that should be possessed under the target theme, so that the executing entity can use these "parameters" to complete the image adjustment (for example, adjusting the image of the initial digital human so that its "hairstyle", "clothing", "height" match and adapt to the target theme).

[0051] In practice, the "parameters" corresponding to the target theme can be pre-configured based on the theme and the part of the image expected to be adjusted, or can be generated by the executing entity according to the target theme through, for example, a generative model (for example, after determining the target theme, the executing entity can call the generative model to adaptively determine the "visual style" that the digital human should have according to the content that should be possessed under the target theme and can significantly represent the target theme, and determine the corresponding "parameters" according to the "visual style").

[0052] As discussed above, the adjustment direction of the image is actually corresponding to the target theme. For example, when the executing entity determines that the target theme is the "medical theme" based on the above step 201, the adjustment direction can be to adjust the "clothing" of the initial digital human to a "doctor's outfit" (for example, adjusting the clothing to a "white coat"), and the adjustment direction can also be to add accessories such as a "stethoscope", etc.

[0053] For another example, when the target theme is "Z movie", the "adjustment direction" can actually be to adjust the "clothing" of the initial digital human to the clothing worn by the movie protagonist in the "Z movie".

[0054] For another example, when the tourist city involved in the travel plan reply information (i.e., the associated city) is used as the target theme, the adjustment direction can be to adjust the "clothing" of the initial digital human to clothing with the regional characteristics of the associated city (for example, the characteristic clothing of ethnic minorities, the characteristic clothing associated with the cultural publicity theme of the associated city). For example, for an associated city with "ice and snow tourism" and "winter" tourism as the cultural publicity theme, the executing entity can obtain the target digital human by adjusting the "clothing" of the initial digital human to winter clothing, so as to reflect and present the "ice and snow tourism" that the associated city expects to publicize and is characteristic of through the "winter clothing" of the target digital human.

[0055] Step 203: Generate a target video that uses the target digital human to provide reply information to the user based on the reply information and the target digital human.

[0056] In an embodiment of the present disclosure, this step aims to enable the above-mentioned execution entity to obtain a target digital human based on the image adjustment of the initial digital human in step 202, and generate a target video for feedback and providing reply information. For example, the execution entity can use the reply information as the audio content of the target video and the target digital human as the image content of the target video, and generate the target video by combining the two.

[0057] Correspondingly, in the target video, at least the target digital human can be used as the "image" content (or, the "explanation", "narration" person, etc. in the image), so as to use the target digital human to speak, explain, and provide reply information. Thus, by using the target video, the user can interact with the target digital human to approximately interact with a real human and vividly obtain the "reply information".

[0058] In some embodiments, in order to improve the quality of the target video, interactive and explanatory actions can also be configured for the target digital human in the target video to enhance the interactive effect of the target digital human. For example, when the target digital human "speaks" the reply information, it can correspondingly have mouth and limb movements to simulate and restore the explanation and narration process provided by a real human in the physical space.

[0059] For example, the execution entity can extract one or more actions (such as "mouth" actions that can move adaptively according to the speaking content, "limb actions" related to "explanation", etc.) from a pre-configured action library according to the specific content of the reply information to form an explanation action sequence. Then, when generating the target video, the execution entity can control the target digital human to perform the corresponding "actions" in the action sequence in the corresponding video frames and images based on the playback position of the reply information corresponding to each explanation action sequence, so that the target digital human can perform these "actions" correspondingly during the process of speaking, explaining, and providing reply information in the target video, and provide the "reply information" to the user more vividly.

[0060] The method for generating a video provided by the embodiment of the present disclosure determines a target theme based on the reply information to be provided to the user; adjusts the image of the initial digital human based on the target theme to obtain a target digital human; generates a target video that uses the target digital human to provide reply information to the user based on the reply information and the target digital human. Thus, the image of the digital human can be adjusted according to the reply information provided to the user, so that the image of the digital human can be personalized to adapt to the reply content, reply scenario, and reply theme, and improve the user's digital human interaction experience.

[0061] Considering that the source of the reply information provided to the user may often be the content in the online knowledge base or the expert knowledge base, after matching and determining the reply information based on the user's question information, the obtained reply information may be too homogeneous in form. Moreover, for such reply information, it may not be convenient for the user to directly read due to differences in the user's language and expression habits. Therefore, in some embodiments, the execution entity may choose to first generate an initial reply information based on the user's question information. That is, the execution entity may first provide matching and searching for the question information through, for example, the expert knowledge base discussed above, and use the matching and searched results as the initial reply information.

[0062] Then, the execution entity uses the language style parameter determined based on the question information to rewrite the style of the initial reply information to obtain the reply information. For example, the execution entity may analyze the question information through, for example, the NLP technology discussed above to determine the language style parameter used to represent and reflect the user's expression habits.

[0063] Then, the execution entity may rewrite the initial reply information with this language style parameter to make the language form and style of the reply information closer to the user's expression habits by adjusting the language style. Correspondingly, the execution entity may use the rewritten "reply information" as the "reply information" used to determine the target theme and generate the target video instead of using the "initial reply information" to determine the target theme and generate the video. Thus, through the rewritten reply information, not only can the over-homogeneity of the reply information in form be avoided, but also its form and style can fit the user's expression habits, making it convenient for the user to read and improving the user's reading experience.

[0064] In some embodiments, when it is allowed to add the "actions" of the target digital human to the target video, if the execution entity uses the language style parameter to adjust the (initial) reply information to make the reply information have the corresponding language style, the execution entity may also choose to determine the "action style" that the user may prefer and tend to based on the "language style" corresponding to the language style parameter, and adjust the actions of the target digital human based on this "action style".

[0065] For example, by pre-maintaining the action relationship between the language style and the action style, after the execution entity determines the language style parameter, it can use the "language style" induced and corresponding to the language style parameter to correspondingly determine the "action style" (for example, based on the pre-maintained relationship between the language style and the action style, corresponding the language style to the action style).

[0066] Then, the execution subject can use the action style parameters corresponding to the "action style" to adjust the target digital human. For example, the action style parameters may include the "amplitude" of the actions made by the target digital human, such as raising hands, turning around, etc., and the actions made when expressing specific words (for example, corresponding to the "regional characteristics" actions of the words included in the "dialect").

[0067] Next, as discussed above, the execution subject may generate the above-mentioned explanation action sequence of the target digital human based on the action style parameters.

[0068] Accordingly, in the process of generating a target video using a target digital human to provide reply information to a user based on the reply information and the target digital human, the execution subject may select as an alternative or substitute to generate a target video using a target digital human to provide reply information to a user based on the reply information, the target digital human and the narration action sequence. Thus, the target digital human in the target video can provide reply information to the user in a more personalized and vivid manner according to the language style and action style preferred and accustomed by the user, thereby improving the user's viewing and obtaining experience of the reply information.

[0069] In some embodiments, after acquiring and obtaining the reply information, the executing entity can also manage the contextual integrity and coherence of the reply information by analyzing the completeness and fluency of the semantic analysis results of the reply information, and adjust and polish the incoherent parts therein to improve the quality of the reply information.

[0070] In some embodiments, the execution subject may also detect whether the pronouns used in the reply information for the (same) user are the same. If at least two different pronouns are used for the user in the reply information, the execution subject may choose to adjust the target pronoun that is different from the preset pronoun among the at least two different pronouns to the preset pronoun (for example, the preset pronoun may be a pre-configured "pronoun" expression form that is expected to be used by the execution subject).

[0071] For example, if the reply message to the same user uses both "you" and "you" at the same time, in this case, the execution entity can adjust the "you" which is not the "preset person" to "you" which is the "preset person". In this way, the person used in the reply message is consistent to avoid causing reading troubles to users.

[0072] For ease of understanding, in the embodiments provided in the present disclosure, a specific process for generating a reply message when the reply message is a travel planning reply message is provided. For example, such a travel planning reply message may be generated by the execution subject in response to the request for the execution subject to provide a travel plan.

[0073] For this process, please refer to Figure 3 . Figure 3 FIG. is a flowchart of a process for generating response information provided by an embodiment of the present disclosure, which includes process 300.

[0074] Specifically, process 300 includes the following steps:

[0075] Step 301: Based on the user's question information, determine the travel destination and the number of travel days;

[0076] Specifically, for the purpose of easy understanding in this embodiment, still taking, for example, server 105 as the "execution subject" for exemplary discussion. It should be understood that, as discussed above, in different scenarios, the execution subject of this embodiment may actually be different from, for example, the execution subject of the process of "generating a video" discussed above, and the present disclosure is not intended to limit this.

[0077] In this step, after obtaining the question information provided by the user (i.e., the relevant information requesting to provide "travel planning response information"), the execution subject may choose to read the question information to determine the travel destination and the number of travel days.

[0078] In practice, for different question information, the reading results may be of various situations. First, when the information provided by the user is relatively complete, the execution subject may directly read based on the question information and correspondingly determine the travel destination and the number of travel days.

[0079] In some scenarios, the question information provided by the user may be incomplete. In such a case, the execution subject may, for example, "pop up a window" to the user to further interact with the user to obtain the missing information. For example, when the user only provides the travel destination (e.g., City B) but does not provide the "number of days", the execution subject may "pop up a window" to request the user to provide the "number of days".

[0080] In some embodiments, in order to improve the user's providing efficiency, for the missing information, the execution subject may also provide some reference items for the user to refer to based on the interaction results with other historical users. For example, if the user only provides the destination of "Country K", the execution subject may select the top preset number of cities as references for the user to provide the "travel destination" based on the popularity of historical users traveling to each city in "Country K" within a historical time period (for example, provide the "top preset number of cities" in the "pop-up window" interface for the user to select and refer to).

[0081] Step 302: Compare the number of travel days with the number-of-days threshold;

[0082] Specifically, after the execution entity determines the "travel days" based on the above step 301, it can compare the travel days with the day threshold. The day threshold can generally be determined according to the criteria for users to distinguish and classify "long-term travel" and "short-term travel".

[0083] Correspondingly, using the day threshold enables the execution entity to distinguish "long-term travel" and "short-term travel", and based on two different strategies, make travel plans for users.

[0084] Next, if the travel days are less than the day threshold, the execution entity can respond to this and execute step 303 to execute the travel plan for "short-term travel". If the travel days are greater than or equal to the day threshold, step 304 is executed to execute the travel plan for "long-term travel".

[0085] Step 303: Generate travel plan reply information based on the first type of recommended places in the recommended place database associated with the travel destination.

[0086] Specifically, for "short-term travel", the execution entity can generate travel plan reply information based on the first type of recommended places in the recommended place database associated with the travel destination.

[0087] The recommended place database usually includes the first type of recommended places and the second type of recommended places located in the travel destination. The recommended index and priority of the first type of recommended places are higher than those of the second type of recommended places. For example, the first type of recommended places can be the current popular scenic spots in the travel destination (such as cities), 5A scenic spots, restaurants that are strongly concerned by tourists, etc., which are more likely to be the scenic spots, restaurants, etc. that users expect to "go to first".

[0088] In some embodiments, the classification of the first type of recommended places and the second type of recommended places can also be realized according to the user portrait and historical behavior of the user. For example, the execution entity can determine the first type of recommended places and the second type of recommended places in the travel destination according to the preference parameters of the user for scenic spots and restaurants.

[0089] Correspondingly, in this step, if it is "short-term travel", the execution entity will only use the first type of recommended places for travel planning, so that the travel plan reply information can more tend to guide users to those "high-value" scenic spots and restaurants that may better meet the user's needs and are more popular. For example, the execution entity can use the scenic spot density algorithm to generate a density strategy, and then fill in the corresponding first type of recommended places based on the density strategy to complete the planning.

[0090] Step 304: Generate travel plan reply information based on the first type of recommended places and the second type of recommended places in the recommended place database.

[0091] Specifically, in this step, if it is a "long-term trip", the executing entity will combine the first type of recommended locations and the second type of recommended locations for trip planning, so that the trip planning reply information can provide more substantial and comprehensive planning and guidance for the user.

[0092] In some embodiments, even when combining the first type of recommended locations and the second type of recommended locations, when the executing entity specifically generates a plan, it can still preferentially choose to plan based on the "first type of recommended locations". For example, the executing entity can first generate an initial plan based on each first type of recommended location, and then complete the final and overall plan by adding the second type of recommended locations.

[0093] Thus, by distinguishing between "long-term trips" and "short-term trips", when making trip plans for users, it is possible to adaptively make plans in combination with different situations, improving the quality of the plans.

[0094] In some alternative implementation manners of this embodiment, if at least two recommended locations are simultaneously recommended in the trip planning reply information (for example, two first type of recommended locations are simultaneously recommended, or a first type of recommended location and a second type of recommended location are simultaneously recommended, etc.), the executing entity can also choose to detect the distance between two adjacent recommended locations. Correspondingly, if the distance is greater than or equal to a distance threshold (for example, this "distance threshold" can be determined based on the criterion that the distance between two continuously recommended locations is considered "too far"), the executing entity can, in response to this, choose to add traffic information between the two adjacent recommended locations in the trip planning reply information. For example, when the distance between the first type of recommended location A and the second type of recommended location B is greater than or equal to the distance threshold, the executing entity can add "traffic information" (for example, it can take bus route T to travel from the first type of recommended location A to the second type of recommended location B, etc.) from the first type of recommended location A to the second type of recommended location B in the trip planning reply information.

[0095] Thus, when the distance between two consecutive recommended locations is relatively far, the executing entity can actively provide "traffic information" in the trip planning reply information for the user to refer to, reducing the user's planning and operation costs while improving the quality of the trip plan.

[0096] In some alternative implementation manners of this embodiment, when the executing entity makes a plan based on recommended locations, it may actually plan at least two combinations of recommended locations. The combination of recommended locations can include at least two first type of recommended locations, or include a first type of recommended location and a second type of recommended location. For example, although the number of days is the same, different combinations of recommended locations may be planned based on different routes, scenic spots, dining options, and arrival sequences.

[0097] In such a case, if the travel plan reply information includes at least two combinations of recommended locations, the execution entity can respond thereto and obtain the user's recommended location preference parameters. For example, the recommended location preference parameters may include the user's preferences for scenic spots, dining (in some embodiments, may also include the user's preferences for locations included in the travel plan reply information, such as hotel preferences, mall preferences, etc.).

[0098] Then, the execution entity can use the recommended location preference parameters to determine the preference matching value of each recommended location combination with the user based on the user's preferences for these "types" such as scenic spots and dining (in some embodiments, can further select based on specific preferences under each type. For example, for "scenic spots", specific preferences may include "natural scenery", "architectural scenery", etc., and for "dining", specific preferences may include, for example, "cuisine", "queue index", etc.). For example, in the case where the user has a stronger preference for "scenic spots", the recommended location combination that includes more "scenic spot" type recommended locations may have a higher preference matching value.

[0099] Subsequently, the execution entity can update the travel plan reply information based on the sorting result of the preference matching values to select, for example, the recommended location combination with the highest preference matching value (or, according to different requirements, can also be to select the recommended location combinations corresponding to the top preset number of preference matching values after sorting the preference matching values from high to low) as the "plan" finally and actually recommended to the user in the travel plan reply information.

[0100] Thus, in the case where multiple recommended location combinations can be initially determined, according to the different needs of the user, to adaptively and personalized provide a plan for the user, improving the planning quality of the travel plan.

[0101] In some alternative implementation manners of this embodiment, if the user's "recommended location preference parameters" are used in the process of determining the recommended location combination, the execution entity can further use the "recommended location preference parameters" to "optimize" the recommended location combination and the plan. For example, even for a "recommended location combination", the execution entity can further select to bring forward the order of going to and arriving at the recommended locations that the user prefers and inclines to, so that the user can go to these recommended locations that the user may prefer and incline to earlier.

[0102] In some alternative implementation manners of this embodiment, for the target video, the execution entity can also choose to add "auxiliary information" thereto so that the user can obtain more information and content in the target video.

[0103] For example, in the case where the reply information is travel plan reply information, the target video actually shows and presents the travel plan explained by the target digital human. Therefore, in such a case, the executing entity can also add the first display video of the first type of recommended destination as auxiliary information to the part corresponding to the first type of recommended destination in the travel plan reply information in the target video. For example, the executing entity can add the explanatory video and promotional video for the first type of recommended destination to the video content part corresponding to the explanation of the first type of recommended destination in the target video, so that the target digital human can use the auxiliary information such as the first display video in the target video to introduce the first type of recommended destination to the user in a more detailed and rich manner.

[0104] For example, the executing entity can add the explanatory video and promotional video (video frames) to each video frame (for example, within a preset area of the video frame) corresponding to the video content part explaining the first type of recommended destination in the target video, so as to add the explanatory video and promotional video to the target video.

[0105] Similarly, if the travel plan reply information also includes a second type of recommended destination, add the second display video of the second type of recommended destination as auxiliary information to the part corresponding to the second type of recommended destination in the travel plan reply information in the target video, so that the target digital human can introduce the second type of recommended destination in the target video in a more detailed manner.

[0106] In some embodiments, for the first display video and the second display video, there can also be differences in video attributes and parameters, so as to enable different strategies to guide users. For example, for the first display video, it can have a higher video pixel, so that users can understand the first type of recommended destination through the real and original content of the first type of recommended destination. For the second display video, the attractiveness of the second type of recommended destination to users can be enhanced by adding, for example, interactive special effects.

[0107] In some embodiments, if the executing entity determines the target theme based on the associated city included in the reply information, the executing entity can also choose to determine the digital human name of the target digital human based on the associated city, and add the "city name", "introduction information of the city", and "digital human name" as auxiliary information. For example, for the case where the associated city is "City E", the digital human name can correspondingly be "Xiaoe of City E".

[0108] Then, the executing entity can add the digital human name and the introduction information of the associated city at the target position in the target video (for example, the position or area where the digital human name is pre-determined to be presented). Thus, the association between the target digital human and the city can be enhanced with the digital human name associated with the city and having the cultural color of the city.

[0109] In some embodiments, if the target topic is "associated city", the executing entity may similarly select the weather data of the associated city as auxiliary information and add it to the target video to provide users with more information for reference.

[0110] In some embodiments, the executing entity may similarly select links that can be included in the reply information (such as document acquisition links, service acquisition links, etc.) as auxiliary information and insert and present them in the target video so that users can conveniently obtain this content through the target video.

[0111] In some embodiments, to facilitate users' acquisition and understanding of the reply information, the executing entity may also present subtitle information for the reply information in the generated target video so that users can assist in understanding and obtaining the reply information through the subtitle information during the process of watching the target video.

[0112] In such a case, if subtitle information for the reply information is presented in the target video, the executing entity may also mark the key information in the reply information (such as the names of the first and second types of recommended locations) based on a pre-configured key information library.

[0113] Then, the executing entity can further highlight the key information in the subtitle information so that users can more intuitively and efficiently obtain this key information through the highlighted content in the subtitle, enhancing the "exposure effect" of the key information.

[0114] In some embodiments, as discussed above, the executing entity may also enable users to directly obtain these services in the target video through auxiliary information in the form of links. For example, in the target video, the executing entity may also add a method such as a ticket purchase link associated with the reply information to reduce the user operation cost.

[0115] In such a case, the executing entity may also choose to continuously check whether these links are still valid and whether there are errors to ensure the availability of the links. For example, for the ticket purchase link, the executing entity may choose to continuously detect the content corresponding to the ticket purchase link in the reply information, the destination pointed to by the ticket purchase link, and the product content provided at the destination. If the executing entity detects that the content corresponding to the ticket purchase link in the reply information does not match the destination pointed to by the ticket purchase link and the product content provided at the destination, the executing entity may choose to block the ticket purchase link to avoid misleading users.

[0116] For better understanding, the present disclosure also gives an implementation solution for a specific process of generating a video in combination with a specific application scenario. For this, please refer to Figure 4 , Figure 4This is a flowchart of the process of generating a video implemented in a specific application scenario provided by the embodiments of the present disclosure, which includes process 400.

[0117] In process 400, for example, the above-mentioned server 105 can also be used as the "execution entity" to implement the process of generating the video.

[0118] In process 400, for the reply information 410 to be provided to the user (not shown in the figure), it can be exemplified as "travel plan reply information" including "travel plan". For example, in form, the reply information 410 can include a five-day itinerary to the "Ice City".

[0119] For the reply information 410, the execution entity can determine the target theme 415 based on the reply information 410 by executing S401. For example, the target theme can be exemplified as the "Ice City", which is the associated city included in the reply information, in form.

[0120] Next, the execution entity can adjust the image of the initial digital human 420 based on the target theme 415 ("Ice City") by executing S402 to obtain the target digital human 425. For example, the execution entity can obtain the target digital human 425 by adjusting the "clothing" of the initial digital human 420 to "winter clothing" so that the "image" of the target digital human 425 corresponds to the target theme 415 such as the "Ice City".

[0121] Then, the execution entity can continue to execute S403 to generate the target video 430 based on the reply information 410 and the target digital human 425. For example, in the generated target video 430, the target digital human 425 can be used as the "image", and the content of the reply information 410 can be used as the audio to provide the reply information 410 to the user in the way that the target digital human 425 explains the specific content of the reply information 410.

[0122] In addition, as discussed above, the execution entity can also provide more information to the user and enhance the user's interactive experience by adding "auxiliary information" to the target video 430.

[0123] For example, S403 may also be included in process 400, so that the execution entity can further execute S404 to add auxiliary information 431 to 434 to the target video 430, obtaining the target video 430'. Among them, in the target video 430', the auxiliary information 431 may be introduction information about "Ice City" in terms of content form, the auxiliary information 432 may be the name of the target digital human 425 generated based on "Ice City" (for example, "Xiaobing of Ice City"), the auxiliary information 433 may be, for example, the first display video of the first type of recommended place currently being explained by the target digital human 425 included in the reply information 410 (for example, the first type of recommended place may be reached on the "second day" in the "plan"), and the auxiliary information 434 may be the subtitle information of the content of the reply information 410 that the target digital human 425 is currently explaining and speaking.

[0124] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for generating a video. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0125] As Figure 5 shown, the video generation device 500 in this embodiment may include: a theme determination unit 501, a digital human determination unit 502, and a video generation unit 503. Among them, the theme determination unit 501 is configured to determine a target theme based on the reply information to be provided to the user; the digital human determination unit 502 is configured to perform an image adjustment on the initial digital human based on the target theme to obtain the target digital human; the video generation unit 503 is configured to generate a target video that uses the target digital human to provide the reply information to the user based on the reply information and the target digital human.

[0126] In this embodiment, in the video generation device 500: the specific processing of the theme determination unit 501, the digital human determination unit 502, and the video generation unit 503 and the technical effects brought by them can respectively refer to Figure 2 the relevant descriptions of steps 201 - 203 in the corresponding embodiments, which will not be elaborated here.

[0127] In some optional implementation manners of this embodiment, the device 500 further includes: an initial reply generation unit configured to generate initial reply information based on the user's question information; an initial reply rewriting unit configured to perform style rewriting on the initial reply information by using the language style parameters determined based on the question information to obtain the reply information.

[0128] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: an action style determination unit configured to determine an action style parameter of a target digital human based on a language style parameter; an action sequence generation unit configured to generate an explanation action sequence of the target digital human based on the action style parameter; and a video generation unit 503 further configured to generate a target video that uses the target digital human to provide a reply message to the user based on the reply message, the target digital human, and the explanation action sequence.

[0129] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: a person adjustment unit configured to, in response to at least two different persons being used for the user in the reply message, adjust a target person different from a preset person among the at least two different persons to the preset person.

[0130] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: a key information marking unit configured to, in response to caption information for the reply message being presented in the target video, mark key information in the reply message based on a pre-configured key information library; and a key information highlighting unit configured to highlight the key information in the caption information.

[0131] In some alternative implementation manners of this embodiment, the reply message is a travel plan reply message, and the travel plan reply message is generated based on the following method: based on the user's question information, determine a travel destination and the number of travel days; in response to the number of travel days being less than a day threshold, generate a travel plan reply message based on a first type of recommended locations in a recommended location database associated with the travel destination; or in response to the number of travel days being greater than or equal to the day threshold, generate a travel plan reply message based on the first type of recommended locations and a second type of recommended locations in the recommended location database.

[0132] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: a traffic recommendation information adding unit configured to, in response to the distance between two adjacent recommended locations in the travel plan reply message being greater than or equal to a distance threshold, add traffic information between the two adjacent recommended locations to the travel plan reply message.

[0133] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: a location preference acquisition unit configured to, in response to at least two types of recommended location combinations being included in the travel plan reply message, acquire a recommended location preference parameter of the user; wherein, the recommended location combination includes at least two first type of recommended locations, or includes a first type of recommended location and a second type of recommended location; a preference matching value determination unit configured to determine a preference matching value of each recommended location combination with the user based on the recommended location preference parameter; and a reply message update unit configured to update the travel plan reply message based on the sorting result of the preference matching values.

[0134] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: a first display video adding unit, configured to add a first display video of the first type of recommended places to a part corresponding to the first type of recommended places in the travel planning reply information in the target video.

[0135] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: a second display video adding unit, configured to, in response to the travel planning reply information further including a second type of recommended places, add a second display video of the second type of recommended places to a part corresponding to the second type of recommended places in the travel planning reply information in the target video.

[0136] In some alternative implementation manners of this embodiment, the theme determining unit 501 is further configured to determine a target theme based on the associated cities included in the reply information to be provided to the user.

[0137] In some alternative implementation manners of this embodiment, the apparatus 500 further includes: a name determining unit, configured to determine a digital human name of the target digital human based on the associated cities; an auxiliary information adding unit, configured to add the digital human name and the introduction information of the associated cities at a target position in the target video.

[0138] This embodiment exists as a device embodiment corresponding to the above method embodiment. The device for generating a video provided in this embodiment can adjust the image of the digital human according to the reply information provided to the user, so that the image of the digital human can be personalized to adapt to the reply content, reply scenario, and reply theme, and improve the user's digital human interaction experience.

[0139] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0140] Figure 6 A schematic block diagram of an exemplary electronic device 600 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0141] As Figure 6As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0142] Multiple components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0143] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the method of generating a video. For example, in some embodiments, the method of generating a video can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the method of generating a video described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the method of generating a video by any other appropriate means (e.g., by means of firmware).

[0144] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0145] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0146] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0147] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input received from the user can be in any form (including acoustic input, speech input, or tactile input).

[0148] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0149] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the problems of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS, Virtual Private Server) services. The server can also be a server of a distributed system, or a server combined with a blockchain.

[0150] According to the technical solution of the embodiment of the present disclosure, the image of the digital human can be adjusted according to the reply information provided to the user, so that the image of the digital human can be personalized to adapt to the reply content, reply scenario, and reply theme, and improve the user's digital human interaction experience.

[0151] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired results of the technical solution provided by this disclosure can be achieved, and no limitations are imposed herein.

[0152] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for generating a video, comprising: Determining a target theme based on reply information to be provided to a user; Adjusting the image of an initial digital human based on the target theme to obtain a target digital human; Generating a target video for providing the reply information to the user by using the target digital human based on the reply information and the target digital human.

2. The method according to claim 1, further comprising: Generating initial reply information based on the user's question information; Rewriting the style of the initial reply information by using a language style parameter determined based on the question information to obtain the reply information.

3. The method according to claim 2, further comprising: Determining an action style parameter of the target digital human based on the language style parameter; Generating an explanatory action sequence of the target digital human based on the action style parameter; And The generating a target video for providing the reply information to the user by using the target digital human based on the reply information and the target digital human comprises: Generating a target video for providing the reply information to the user by using the target digital human based on the reply information, the target digital human and the explanatory action sequence.

4. The method according to claim 2, further comprising: In response to at least two different personal pronouns being used for the user in the reply information, adjusting a target personal pronoun different from a preset personal pronoun among the at least two different personal pronouns to the preset personal pronoun.

5. The method according to claim 1, further comprising: In response to subtitle information for the reply information being presented in the target video, marking key information in the reply information based on a pre-configured key information library; Highlighting the key information in the subtitle information.

6. The method according to any one of claims 1-5, wherein The reply information is travel planning reply information, and the travel planning reply information is generated based on the following method: Determining a travel destination and the number of travel days based on the user's question information; In response to the number of travel days being less than a days threshold, generating the travel planning reply information based on first-category recommended places in a recommended place database associated with the travel destination; or In response to the number of travel days being greater than or equal to the days threshold, generating the travel planning reply information based on the first-category recommended places and second-category recommended places in the recommended place database.

7. The method according to claim 6, further comprising: In response to the distance between two adjacent recommended places in the travel planning reply information being greater than or equal to a distance threshold, adding traffic information between the two adjacent recommended places in the travel planning reply information.

8. The method according to claim 6, further comprising: In response to the travel planning reply information including at least two recommended place combinations, obtaining the user's recommended place preference parameter; wherein, the recommended place combination includes at least two of the first-category recommended places, or includes the first-category recommended places and the second-category recommended places; Determining a preference matching value of each recommended place combination with the user based on the recommended place preference parameter; Update the travel planning reply information according to the sorting result based on the preference matching value.

9. The method according to claim 6 further includes: Add the first display video of the first type of recommended location to the part corresponding to the first type of recommended location in the travel planning reply information in the target video.

10. The method according to claim 9 further includes: In response to the travel planning reply information further including a second type of recommended location, add the second display video of the second type of recommended location to the part corresponding to the second type of recommended location in the travel planning reply information in the target video.

11. The method according to claim 1, wherein The determining of the target theme based on the reply information to be provided to the user includes: Determine the target theme based on the associated cities included in the reply information to be provided to the user.

12. The method according to claim 11 further includes: Determine the digital human name of the target digital human based on the associated city; Add the digital human name and the introduction information of the associated city at the target position in the target video.

13. A device for generating a video, comprising: A theme determination unit configured to determine a target theme based on the reply information to be provided to the user; A digital human determination unit configured to perform an image adjustment on an initial digital human based on the target theme to obtain a target digital human; A video generation unit configured to generate a target video that uses the target digital human to provide the reply information to the user based on the reply information and the target digital human.

14. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method for generating a video according to any one of claims 1-12.

15. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method for generating a video according to any one of claims 1-12.

16. A computer program product comprising a computer program which, when executed by a processor, implements the method for generating a video according to any one of claims 1-12.

Citation Information

Cited By

  • Method and program product for providing question and answer service based on video content

    CN121262422A

  • Interaction method and device based on virtual character, electronic equipment and storage medium

    CN121531149A