Resource recommendation video generation method and device, electronic equipment and storage medium
By generating target carrier information and character image packages through intelligent agents, the problem of low quality of single-person broadcast videos is solved, and the accuracy and diversity of resource-recommended videos are achieved, making it suitable for multiple application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-21
AI Technical Summary
Existing single-person live-streaming videos use fixed templates for virtual anchors and rigid scripts that combine selling points with promotional slogans, resulting in low video quality and user resistance.
The system employs intelligent agents to generate resource recommendation videos. By acquiring resource information, it generates target carrier information, including target role information and script information. Based on the role image package and script information, it generates resource recommendation videos. It utilizes multiple large models and evaluation agents to optimize candidate carrier information, generates optimal target carrier information, and generates role image packages and final videos through text-to-image and image-to-video intelligent agents.
It improves the accuracy and quality of resource-recommended videos, generates character images and scripts that match the resources, avoids visual fatigue, and is suitable for multiple application scenarios such as corporate publicity, medical, education, and finance industries.
Smart Images

Figure CN121908077A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, specifically to the fields of resource recommendation and artificial intelligence technology such as intelligent agents, and particularly to a method, apparatus, electronic device and storage medium for generating resource recommendation videos. Background Technology
[0002] In the current e-commerce advertising ecosystem, single-person headshot ads remain the most widely used creative format. Their advantages lie in their low production threshold, fast deployment speed, and wide platform compatibility, making them popular among many advertisers.
[0003] Existing solo narration videos mostly use fixed templates for virtual anchors, such as "a young woman in her twenties, with a white background and professional attire," with minimal differences in the anchor's appearance between different advertisements and brands. Existing solo narration scripts typically consist of a simple list of product selling points plus promotional slogans. Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for generating resource recommendation videos.
[0005] According to one aspect of this disclosure, a method for generating resource recommendation videos is provided, comprising:
[0006] Obtain information about resources;
[0007] An intelligent agent is used to generate target carrier information for recommending the resources based on the information of the resources. The target carrier information includes target role information and corresponding target script information.
[0008] Based on the target character information, a character image package is generated;
[0009] Based on the character image package and the target script information, a resource recommendation video is generated.
[0010] According to another aspect of this disclosure, an apparatus for generating resource recommendation videos is provided, comprising:
[0011] The acquisition module is used to obtain information about resources;
[0012] The text generation module is used to generate target carrier information for recommending the resource based on the information of the resource using an intelligent agent. The target carrier information includes target role information and corresponding target script information.
[0013] The image generation module is used to generate a character image package based on the target character information;
[0014] The video generation module is used to generate resource recommendation videos based on the character image package and the target script information.
[0015] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0016] At least one processor; and
[0017] A memory communicatively connected to the at least one processor; wherein,
[0018] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods described above and any possible implementations.
[0019] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform the methods described above and any possible implementation thereof.
[0020] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the aspects and any possible implementations described above.
[0021] According to the technology disclosed herein, the accuracy of generated resource recommendation videos can be effectively improved.
[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0023] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0024] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure;
[0025] Figure 2 This is a schematic diagram according to the second embodiment of the present disclosure;
[0026] Figure 3 This is a schematic diagram according to the third embodiment of the present disclosure;
[0027] Figure 4 This is a schematic diagram according to the fourth embodiment of the present disclosure;
[0028] Figure 5 This is a block diagram of an electronic device used to implement the methods of the embodiments of this disclosure. Detailed Implementation
[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0030] Obviously, the described embodiments are only some, not all, of the embodiments disclosed herein. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0031] It should be noted that the terminal devices involved in the embodiments of this disclosure may include, but are not limited to, smart devices such as mobile phones, personal digital assistants (PDAs), wireless handheld devices, and tablet computers; the display devices may include, but are not limited to, personal computers, televisions, and other devices with display functions.
[0032] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0033] Existing single-person live-streaming videos use virtual anchors with fixed templates and scripts consisting of "listing product selling points + promotional slogans," resulting in rigid and low-quality videos that are met with strong resistance from users.
[0034] Figure 1 This is a schematic diagram based on the first embodiment of the present disclosure; as shown Figure 1 As shown in the figure, this embodiment provides a method for generating resource recommendation videos, which may specifically include the following steps:
[0035] S101, Obtain resource information;
[0036] The resources in this embodiment can be advertising resources or other types of resources to be recommended.
[0037] The resource in this embodiment can be in the form of a video. For example, the resource information may include an advertising video for a product, or a promotional video for a scenic spot or tourist city, etc. Optionally, the resource information may also include the name of the resource, such as the name of a product, or the name of a scenic spot or tourist city, etc.
[0038] S102. Using an intelligent agent, based on resource information, generate target carrier information for recommended resources. The target carrier information includes target role information and corresponding target script information.
[0039] In this embodiment, the objective is to generate recommended videos for resources. Considering the characteristics of the videos themselves, this embodiment employs an intelligent agent to intelligently generate target carrier information for the recommended resources, which may include target role information and corresponding target script information. Target role information identifies the role of the recommended resource; target script information identifies the specific content of the recommended resource.
[0040] S103. Generate a character image package based on the target character information;
[0041] The character image package in this embodiment may refer to a character image video, but this image video is a semi-finished product, does not include any sound, and is only used to identify the target character's information.
[0042] S104. Based on the character image package and target script information, generate resource recommendation videos.
[0043] Specifically, based on the character image package and target script information, a resource recommendation video is generated that uses the target character to verbally recite the target script information.
[0044] The resource recommendation video generated in this embodiment can be used to recommend resources. For example, when the resource is an advertisement video for a product, the resource recommendation video can specifically be a recommended video for that product advertisement video. In practical use, the resource recommendation video can be used as the intro video of an advertisement, loaded before the advertisement video, to effectively recommend the product, grab the user's attention, and make the user more willing to watch the subsequent advertisement video.
[0045] The execution entity of the resource recommendation video generation method in this embodiment can be a resource recommendation video generation device. This device can be an electronic device or a software-integrated application. When in use, user information is input into the device, which can call an intelligent agent to generate target carrier information of the recommended resources based on the resource information, thereby generating an image package and a resource recommendation video.
[0046] The resource recommendation video generation method of this embodiment, by employing an intelligent agent, generates target carrier information for recommended resources based on resource information. This ensures the compatibility between the generated target carrier information and the resources, as well as the rationality and accuracy of the target role information and corresponding target script information within the generated target carrier information. Consequently, it effectively improves the accuracy of the generated target role information and target script information, further enhancing the rationality and accuracy of the generated character image package, and ultimately improving the accuracy of the generated resource recommendation video.
[0047] The technical solution of this embodiment can avoid using fixed character templates and rigid scripts that generate resource recommendation videos with selling points and slogans. It can intelligently generate resource recommendation videos that match the information of the resources, effectively improving the quality of the generated resource recommendation videos.
[0048] Figure 2 This is a schematic diagram based on the second embodiment of this disclosure; the method for generating resource recommendation videos in this embodiment, as described above... Figure 1 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be further described in more detail. For example... Figure 2 As shown, the method for generating resource recommendation videos in this embodiment may specifically include the following steps:
[0049] S201. Obtain the name and video of the resource;
[0050] For example, the video in this embodiment can be an advertisement video for a product, and the name of the resource can be the name of the product. In this embodiment, the technical solution of this disclosure is described using the example of generating a recommended video for a product advertisement video.
[0051] S202. A text generation agent is used to call multiple large models. Based on resource information, multiple sets of candidate carrier information for recommended resources are generated. Each set of candidate carrier information includes candidate role information and corresponding candidate script information. The candidate role information includes the basic attribute information of the candidate role and the suitable scenario information of the candidate role.
[0052] Similarly, taking a video advertisement for a product as an example, in order to obtain the most accurate and suitable target carrier information, this embodiment can use a text generation agent to call multiple large models in parallel to generate multiple sets of candidate carrier information for recommended advertisement videos based on the product name and the product's advertisement video.
[0053] For example, if we call three large models, such as model a, model b, and model c, and each model can generate three sets of candidate carrier information, then we can obtain nine sets of candidate carrier information. Each set of candidate carrier information includes candidate role information and corresponding candidate script information; each candidate role information includes the candidate role's basic attribute information and the candidate role's suitable scenario information.
[0054] The basic attribute information of candidate characters can include age, identity, and gender, which are used to limit the basic information of the characters in the resource recommendation video. Identity is used to identify the target character's profession and rank within that profession. The suitable scene information for candidate characters can include the environment in which the candidate character is located, their clothing, and lighting, which are used to limit the environmental information and suitable clothing and lighting information for the characters in the resource recommendation video.
[0055] In this embodiment, by employing a text generation agent to invoke multiple large models and generate multiple sets of candidate carrier information, it is possible to effectively ensure that the candidate role information and corresponding candidate script information in each set of candidate carrier information are well-matched with the resource information. For example, if the resource information is a food product, the generated candidate role information could be a nutrition PhD, and the suitable scenario information for the candidate role could be a laboratory, making the role information more authoritative and convincing to users. As another example, if the resource information is a medicine, the generated candidate role information could be a doctor, and the suitable scenario information for the candidate role could be a hospital. If the resource information is a training course, the generated candidate role information could be a company founder, and the suitable scenario information for the candidate role could be a company office. The candidate script information is a script that recommends the resource information using the corresponding candidate role as the recommender. Therefore, the candidate script information includes at least the candidate role information, the product's selling points, and may also include pain points of existing technologies and calls to action. The above are just a few examples of embodiments of this disclosure. In actual application scenarios, the text generation agent, by invoking various models and based on the resource information, can generate highly matching candidate role information and candidate script information.
[0056] The resource recommendation video to be generated in this embodiment is used to recommend resources. Its duration can be 10-15 seconds, and this duration can be preset. Additionally, the text length included in the script information can also be preset, for example, it can include 100 characters. In this embodiment, candidate script information can be generated based on the duration of the recommended resource video and the limited text length of the script information. Each candidate script information can have a reasonable structure. For example, one example of a candidate script information structure is as follows: First, 0-2 seconds are used for the character to introduce themselves, establishing authority and making the user more trusting; then 0-3 seconds are used to point out a typical pain point; then 0-3 seconds are used to propose a solution and integrate it into the product; finally, 0-2 seconds are used to give a call to action. Of course, the above script information structure is not unique; the two middle sections can also be combined to directly highlight the product's selling points. Other structures can also be used, which will not be elaborated upon here.
[0057] In this embodiment, the language description used in the generated candidate script information needs to be consistent with the identity of the corresponding candidate role; if the candidate script information includes pain points, the pain points should closely correspond to the selling points of the product; moreover, the text length of the candidate script information should match the specified duration based on the preset speaking speed, and the sentence structure of the candidate script information should be suitable for solo broadcasting in a conversational style.
[0058] S203. A text evaluation agent calls multiple large models to evaluate candidate carrier information that is not generated by itself, and obtains the corresponding evaluation scores.
[0059] For example, consider the following: large model a generates candidate carrier information 1, candidate carrier information 2, and candidate carrier information 3; large model b generates candidate carrier information 4, candidate carrier information 5, and candidate carrier information 6; and large model c generates candidate carrier information 7, candidate carrier information 8, and candidate carrier information 9.
[0060] During the evaluation, a text evaluation agent can be used to call the large model a to evaluate all or part of the candidate carrier information 4 to 9.
[0061] A text evaluation agent is used to call the large model b to evaluate all or part of the candidate carrier information 1-3 and candidate carrier information 7-9.
[0062] A text evaluation agent is used to call the large model c to evaluate all or part of the candidate carrier information 1 to candidate carrier information 6.
[0063] In short, cross-evaluation is essential to ensure a more objective and accurate assessment score.
[0064] For example, in the specific implementation of this step, the evaluation of any set of candidate carrier information may include an evaluation of at least one of the following dimensions:
[0065] (1) The text evaluation agent calls the first model, and based on the information of resources and the basic attribute information of candidate roles generated by non-first model and the corresponding candidate script information, the role and identity matching evaluation is carried out to obtain the corresponding role and identity matching degree;
[0066] For example, role-identity matching assessment includes checking whether the script matches the tone and professional depth of the candidate role.
[0067] The first major model here can be any of the multiple major models.
[0068] (2) The text evaluation agent calls the first model and evaluates the rationality of the information structure based on the candidate script information generated by the non-first model to obtain the rationality of the corresponding information structure.
[0069] For example, assessing the rationality of information structure includes checking whether candidate script information has a complete closed loop of "identity-pain point-solution-call to action", or whether it conforms to other reasonable information structures.
[0070] (3) The text evaluation agent calls the first model and evaluates the fit between resources and recommendations based on the information of resources and candidate scripts generated by non-first models to obtain the corresponding fit.
[0071] For example, the fit assessment between resources and recommendations includes detecting whether candidate script information accurately expresses the selling points of resources such as products.
[0072] (4) The text evaluation agent calls the first-largest model and evaluates the naturalness of the spoken script based on the candidate script information generated by non-first-largest models to obtain the corresponding naturalness; and
[0073] For example, the assessment of the naturalness of spoken delivery includes detecting whether the sentence length and pause positions of the candidate script information are suitable for single-person spoken delivery.
[0074] (5) The text evaluation agent calls the first model and performs compliance evaluation based on the candidate script information generated by the non-first model to obtain the corresponding compliance degree.
[0075] For example, compliance assessment includes checking whether candidate script information contains non-compliant content such as sensitive words or extreme words; if non-compliant content is found, the compliance score is directly 0.
[0076] S204. Based on the evaluation scores of each group of candidate vector information and the preset evaluation threshold, obtain the target vector information;
[0077] This step can be implemented in the following ways:
[0078] (a) Calculate the target evaluation score for each candidate carrier information based on at least one of the following: role identity matching degree, information structure rationality, adaptability, naturalness, and compliance; then proceed to step (b).
[0079] For example, taking into account the matching degree of role identity, the rationality, adaptability, naturalness and compliance of information structure, for each candidate carrier information, the evaluation structure obtained from the above steps (1)-(5) can be weighted and summed to obtain the target evaluation score of the candidate carrier information. Specifically, the weight of the evaluation value of each dimension can be configured based on the requirements. For the evaluation value of the dimension that you do not want to refer to, you can also directly set the weight to 0. In actual application scenarios, you can also select only some of the evaluation values of the dimensions for weighted summation according to the requirements to obtain the final target evaluation score.
[0080] (b) Detect whether there are candidate carrier information whose target evaluation scores in each group of candidate carrier information are greater than or equal to the preset evaluation threshold; if so, proceed to step (c); if not, proceed to step (d).
[0081] (c) Obtain the candidate carrier information with the highest target evaluation score among the candidate carrier information of each group, and use it as the target carrier information, then end.
[0082] (d) Employ a reflective agent to generate optimization suggestions for each candidate carrier based on the information of each candidate carrier; execute step (e).
[0083] In this embodiment, the preset evaluation threshold can be set to a relatively high value. For example, if the maximum score is 1, it can be set to a large value such as 0.9, 0.92, or 0.88. If there are candidate carrier information with a target evaluation score greater than or equal to the preset evaluation threshold, the candidate carrier information with the highest target evaluation score is directly obtained and used as the target carrier information. Otherwise, if there are no candidate carrier information with a target evaluation score greater than or equal to the preset evaluation threshold, a reflective agent is directly used to reflect and generate optimization suggestions for the candidate carrier information, so as to further obtain the target carrier information after optimization.
[0084] Further optionally, in one embodiment of this disclosure, during the implementation of step S203, a text evaluation agent calls multiple large models to evaluate candidate carrier information that is not generated by itself, and while obtaining the corresponding evaluation scores, it can also generate a corresponding problem list to identify the problems detected during the evaluation.
[0085] At this point, when implementing this step, a reflective agent can be used to generate optimization suggestions for each candidate carrier based on the information of each candidate carrier and the corresponding problem list, which can further improve the accuracy of the generated optimization suggestions.
[0086] (e) Based on the optimization suggestions of each candidate vector information, a genetic algorithm is used to optimize each candidate vector information; return to step S203 to continue to evaluate the optimized candidate vector information until the target vector information is obtained.
[0087] The target carrier information obtained in this embodiment includes target character information and corresponding target script information. Specifically, the target character information may include the target character's basic attribute information and the target character's suitable scene information. Correspondingly, the target character's basic attribute information may include age, identity, and gender, etc., used to limit the basic information of the characters in the resource recommendation video, where identity is used to identify the target character's profession and level within that profession. The target character's suitable scene information may include the target character's environment information, clothing, and lighting, etc., used to limit the environmental information and suitable clothing, lighting, etc., of the characters in the resource recommendation video.
[0088] In this embodiment, the optimization of each candidate carrier information mainly involves optimizing the candidate script information of each candidate carrier information. This can effectively improve at least one of the following aspects of the optimized candidate script information: role identity matching degree, information structure rationality, resource and recommendation adaptability, naturalness of verbal delivery, and compliance.
[0089] When optimizing candidate vector information using a genetic algorithm, the structural features, key sentences, and generation parameters of the corresponding candidate script information are encoded and used as individual genetic algorithms. Then, based on the optimization suggestions, the script structure and wording are modified in a targeted manner during the crossover and mutation stages of the genetic algorithm. For example, more specific scenarios are added, business operation details are supplemented, and overly marketing expressions are removed, in order to obtain more reasonable and accurate optimized script information.
[0090] S205. Generate initial character prompts based on target character information; target character information includes the target character's basic attribute information and the target character's suitable scene information;
[0091] For example, in one scenario of this embodiment, the generated character prompt could be: a male company founder around 50 years old, with short hair, wearing a dark suit, standing in front of the floor-to-ceiling window in his office, looking confident and composed, etc.
[0092] For example, in a practical implementation, a text-generating agent can be used to generate accurate initial role prompts based on the target role information.
[0093] S206. Employ a Video Prompt Optimization (VPO) agent to optimize the initial role prompts a preset number of times, thereby obtaining a preset number of optimized role prompts.
[0094] The preset number in this embodiment can also be a positive integer preset according to actual needs. For example, it can be 4 times, 6 times, or 8 times. For example, in this embodiment, the VPO model can optimize character prompts. For example, during the optimization process, the facial features of the character prompts can be adjusted, such as adjusting facial features or face shape, eyebrows, etc., and the character's hairstyle and clothing can also be adjusted.
[0095] In practice, role prompts can be input into the VPO agent, which can then output a preset number of optimized role prompts based on the input prompts.
[0096] S207. Using a text-based image-based intelligent agent, a preset number of character images are generated based on a preset number of optimized character prompt words; the character images include images of various poses;
[0097] In other words, the character image generated in this step can be considered not as a single image, but as a set of images including multiple poses, facilitating subsequent video generation. It's important to note that because the optimized character prompts include the target character's suitable scene information, the generated character images also carry this scene information, such as the background.
[0098] In this embodiment, multiple poses that need to be included can be preset. For example, it can include the poses of the character in the four directions of front, back, left and right, or it can further include the poses in four diagonal directions, that is, a total of 8 poses.
[0099] For any optimized character prompt, the Wenshengtu AI agent can generate a corresponding character image based on that optimized character prompt.
[0100] Alternatively, in this embodiment, the same step can be used for the initial character prompts: a text-based image agent can be used to generate corresponding initial character images, which are then used together with a preset number of character images.
[0101] S208. Using image-generated video intelligence, based on a preset number of character images, generate a preset number of character image packages respectively;
[0102] The character image package generated in this embodiment can be considered as a silent video clip that includes character information, mainly used to represent the character and the environment in which the character is located.
[0103] It should be noted that although the generated character image package has no sound, its mouth still needs to simulate the movements of speaking, so that when generating videos later, audio and lip-syncing technology can be combined to drive the script dialogue to synchronize with the character's lip movements.
[0104] Similarly, the same method can be used to generate a character image pack for the initial character prompt words and their corresponding initial character images.
[0105] In this embodiment, the generated preset number of character image packs can also be stored and used as source material for generating resource recommendation videos. If it is necessary to generate recommendation videos for other resources without specified characters in the future, the preset number of character image packs can be used directly.
[0106] S209. Based on the target script information and the pre-selected timbre, synthesize the target audio;
[0107] S210. Using a video synthesis intelligent agent, based on the target audio and a preset number of character image packages, synthesize resource recommendation videos corresponding to each character image package to obtain a preset number of resource recommendation videos.
[0108] In this embodiment, for each character image package, when the video synthesis agent synthesizes the corresponding resource recommendation video based on the target audio and the character image package, it can use lip-sync technology to synchronize the script spoken by the character with the character's lip movements, making the resource recommendation video more natural.
[0109] Optionally, in this embodiment, a video synthesis agent can be directly used to generate resource recommendation videos corresponding to each character image package based on the target script information, the pre-selected timbre, and a preset number of character image packages, thereby obtaining a preset number of resource recommendation videos; the timbre of each resource recommendation video is the selected timbre.
[0110] In this embodiment, a predetermined number of resource recommendation videos need to be generated. These videos share the same target script information but differ slightly in character appearance. This allows for subsequent resource recommendation based on a random selection of a video. After a certain recommendation period, the video with the highest play count or longest playtime can be statistically analyzed and identified as the optimal resource recommendation video. This allows for optimization of subsequent agent generation tasks based on the character image package and optimized character prompts corresponding to the optimal resource recommendation video, thus favoring the generation of similar resource recommendation videos.
[0111] The resource recommendation video generated in this embodiment can serve as an "authoritative introducer" or "host" at the beginning of the resource video, quickly explaining the settings and lowering the threshold for understanding.
[0112] The method for generating resource recommendation videos in this embodiment can be applied to the field of corporate publicity and training to generate virtual executives and virtual lecturers for use in scenarios such as welcome speeches and opening remarks.
[0113] The method for generating resource recommendation videos in this embodiment can also be applied to industry scenarios that emphasize trust and compliance, such as healthcare, education, and finance, to generate more professional and credible virtual expert avatars.
[0114] In these scenarios, the character image package generated in the embodiments of this disclosure can also be directly reused, and then virtual characters with a unified style and long-term reusability can be automatically output according to the scripts of different projects.
[0115] The resource recommendation video generation method in this embodiment uses a text generation agent to call multiple large models, generates multiple sets of candidate carrier information based on resource information, and obtains the optimal target carrier information based on the evaluation scores of multiple sets of candidate carrier information and preset evaluation scores. This can effectively improve the accuracy of the target carrier information required to generate resource recommendation videos and provide necessary data support for the generation of resource recommendation videos.
[0116] The resource recommendation video generation method in this embodiment generates multiple candidate carrier information and obtains target carrier information through evaluation. This results in target carrier information that is more matched with roles and identities, has a more reasonable information structure, better adapts to recommended resources, is more natural, and more in line with standards. As a result, more accurate and higher-quality resource recommendation videos can be generated.
[0117] The resource recommendation video generation method in this embodiment generates a preset number of optimized character prompts through optimization. Then, it uses text-based image intelligence and image-based video intelligence to generate a preset number of character image packages. Multiple character image packages can be generated, which in turn can generate multiple resource recommendation videos. This not only effectively enriches the number of resource recommendation videos, but also avoids visual fatigue caused by only one type of resource recommendation video during random recommendations, thereby further improving the efficiency of resource recommendation.
[0118] Figure 3 This is a schematic diagram based on the third embodiment of this disclosure; as shown Figure 3 As shown, this embodiment provides a resource recommendation video generation device 300, including:
[0119] Module 301 is used to obtain resource information;
[0120] The text generation module 302 is used to generate target carrier information for recommending the resource based on the information of the resource using an intelligent agent. The target carrier information includes target role information and corresponding target script information.
[0121] Image generation module 303 is used to generate a character image package based on the target character information;
[0122] The video generation module 304 is used to generate resource recommendation videos based on the character image package and the target script information.
[0123] The resource recommendation video generation device 300 in this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0124] Figure 4 This is a schematic diagram based on the fourth embodiment of the present disclosure; as shown Figure 4 As shown, the resource recommendation video generation device 400 of this embodiment, in the above-described... Figure 3 Based on the technical solutions of the illustrated embodiments, the technical solutions of this disclosure will be further described in more detail. For example... Figure 4 As shown, the resource recommendation video generation device 400 of this embodiment includes the above-described... Figure 3 The modules with the same name and function shown are: acquisition module 401, text generation module 402, image generation module 403, and video generation module 404.
[0125] like Figure 4 As shown, the text generation module 402 in this embodiment includes:
[0126] The generation unit 4021 uses a text generation agent to call multiple large models and, based on the information of the resources, generates multiple sets of candidate carrier information to recommend the resources. Each set of candidate carrier information includes candidate role information and corresponding candidate script information. The candidate role information includes the basic attribute information of the candidate role and the suitable scenario information of the candidate role.
[0127] The evaluation unit 4022 is used to use a text evaluation agent to call multiple large models to evaluate the candidate carrier information that is not generated by itself, and obtain the corresponding evaluation score.
[0128] The acquisition unit 4023 is used to acquire the target carrier information based on the evaluation score of each group of candidate carrier information and a preset evaluation threshold.
[0129] Further, optionally, in one embodiment of this disclosure, the evaluation unit 4022 is configured to perform at least one of the following steps:
[0130] The text evaluation agent invokes the first major model, and based on the information of the resource and the basic attribute information of the candidate role not generated by the first major model and the corresponding candidate script information, performs a role-identity matching evaluation to obtain the corresponding role-identity matching degree.
[0131] The text evaluation agent calls the first major model and evaluates the rationality of the information structure based on the candidate script information not generated by the first major model, thereby obtaining the rationality of the corresponding information structure.
[0132] The text evaluation agent calls the first major model to evaluate the suitability between the resource and the recommendation based on the information of the resource and the candidate script information not generated by the first major model, and obtains the corresponding suitability.
[0133] The text evaluation agent invokes the first major model, and based on the candidate script information not generated by the first major model, evaluates the naturalness of the spoken delivery to obtain the corresponding naturalness score; and
[0134] The text evaluation agent invokes the first major model, and performs a compliance evaluation based on the candidate script information in the candidate carrier information not generated by the first major model, to obtain the corresponding compliance degree.
[0135] Further optionally, in one embodiment of this disclosure, the acquisition unit 4023 is used for:
[0136] Based on at least one of the following factors for each group of candidate carrier information: role identity matching degree, information structure rationality, adaptability, naturalness, and compliance, the target evaluation score for each group of candidate carrier information is calculated.
[0137] Detect whether there are any candidate carrier information in each group whose target evaluation score is greater than or equal to the preset evaluation threshold;
[0138] If it exists, obtain the candidate carrier information with the largest target evaluation score among the candidate carrier information in each group, and use it as the target carrier information.
[0139] Further optional, such as Figure 4 As shown, in one embodiment of this disclosure, the text generation module 402 further includes:
[0140] The generation unit 4024 is used to generate optimization suggestion information for each candidate carrier information if the target evaluation score of each group of candidate carrier information is less than the preset evaluation threshold.
[0141] The carrier optimization unit 4025 is used to optimize the candidate carrier information based on the optimization suggestion information of each candidate carrier information using a genetic algorithm.
[0142] Further optional, such as Figure 4 As shown, in one embodiment of this disclosure, the evaluation unit 4022 is used to call multiple large models using the text evaluation agent to evaluate the candidate carrier information that is not generated by itself, and at the same time obtain the corresponding evaluation score, it also generates a list of questions.
[0143] The suggestion generation unit 4024 is used to generate optimization suggestion information for each candidate carrier information based on the candidate carrier information and the corresponding problem list, using the reflexive agent.
[0144] Further optional, such as Figure 4 As shown, in one embodiment of this disclosure, the image generation module 403 includes:
[0145] The prompt word generation unit 4031 is used to generate initial character prompt words based on the target character information; the target character information includes the basic attribute information of the target character and the suitable scene information of the target character;
[0146] The prompt word optimization unit 4032 is used to optimize the initial role prompt words a preset number of times using a video prompt optimization agent to obtain a preset number of optimized role prompt words.
[0147] The text-based image unit 4033 is used to generate a preset number of character images based on the preset number of optimized character prompt words using a text-based image agent; the character images include images of various poses;
[0148] Image-generated video unit 4034 is used to generate a preset number of character image packages based on the preset number of character images using image-generated video intelligence.
[0149] Further optional, such as Figure 4 As shown, in one embodiment of this disclosure, the video generation module 404 is used for:
[0150] Based on the target script information and the pre-selected timbre, the target audio is synthesized;
[0151] A video synthesis agent is used to synthesize resource recommendation videos corresponding to each of the target audio and a preset number of character image packages, thereby obtaining a preset number of resource recommendation videos.
[0152] The resource recommendation video generation device 400 of this embodiment achieves the same implementation principle and technical effect as the above-mentioned related method embodiments by using the above-mentioned modules. For details, please refer to the description of the above-mentioned related method embodiments, which will not be repeated here.
[0153] The acquisition, storage, and application of any type of information, such as user personal information, involved in the technical solutions disclosed herein comply with relevant laws and regulations and do not violate public order and good morals.
[0154] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0155] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0156] like Figure 5 As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 502 or a computer program loaded from storage unit 508 into random access memory (RAM) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0157] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0158] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the methods of this disclosure. For example, in some embodiments, the methods of this disclosure may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods of this disclosure described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the methods of this disclosure by any other suitable means (e.g., by means of firmware).
[0159] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0160] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0161] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0162] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0163] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0164] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0165] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0166] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating resource recommendation videos, comprising: Obtain information about resources; An intelligent agent is used to generate target carrier information for recommending the resources based on the information of the resources. The target carrier information includes target role information and corresponding target script information. Based on the target character information, a character image package is generated; Based on the character image package and the target script information, a resource recommendation video is generated.
2. The method according to claim 1, wherein, Using an intelligent agent, based on the information of the resources, target carrier information for recommending the resources is generated, including: A text generation agent is used to invoke multiple large models. Based on the information of the resources, multiple sets of candidate carrier information are generated to recommend the resources. Each set of candidate carrier information includes candidate role information and corresponding candidate script information. The candidate role information includes the basic attribute information of the candidate role and the suitable scenario information of the candidate role. A text evaluation agent is used to call multiple large models to evaluate the candidate carrier information that is not generated by itself, and obtain the corresponding evaluation scores. Based on the evaluation scores of the candidate carrier information in each group and the preset evaluation threshold, the target carrier information is obtained.
3. The method according to claim 2, wherein, A text evaluation agent invokes multiple large models to evaluate the candidate carrier information that is not generated by the agent itself, and obtains corresponding evaluation scores, including at least one of the following: The text evaluation agent invokes the first major model, and based on the information of the resource and the basic attribute information of the candidate role not generated by the first major model and the corresponding candidate script information, performs a role-identity matching evaluation to obtain the corresponding role-identity matching degree. The text evaluation agent calls the first major model and evaluates the rationality of the information structure based on the candidate script information not generated by the first major model, thereby obtaining the rationality of the corresponding information structure. The text evaluation agent calls the first major model to evaluate the suitability between the resource and the recommendation based on the information of the resource and the candidate script information not generated by the first major model, and obtains the corresponding suitability. The text evaluation agent calls the first major model and evaluates the naturalness of the spoken script based on the candidate script information not generated by the first major model, thereby obtaining the corresponding naturalness. as well as The text evaluation agent invokes the first major model and performs a compliance evaluation based on the candidate script information not generated by the first major model to obtain the corresponding compliance score.
4. The method according to claim 3, wherein, Based on the evaluation scores of the candidate carrier information in each group and a preset evaluation threshold, the target carrier information is obtained, including: Based on at least one of the following factors for each group of candidate carrier information: role identity matching degree, information structure rationality, adaptability, naturalness, and compliance, the target evaluation score for each group of candidate carrier information is calculated. Detect whether there are any candidate carrier information in each group whose target evaluation score is greater than or equal to the preset evaluation threshold; If it exists, obtain the candidate carrier information with the largest target evaluation score among the candidate carrier information in each group, and use it as the target carrier information.
5. The method according to claim 4, wherein, The method further includes: If the target evaluation score of each group of candidate carrier information is less than the preset evaluation threshold, a reflective agent is used to generate optimization suggestion information for each candidate carrier information based on each candidate carrier information. Based on the optimization suggestions for each candidate vector, a genetic algorithm is used to optimize the candidate vector information.
6. The method according to claim 5, wherein, The method further includes: The text evaluation agent invokes multiple large models to evaluate the candidate carrier information that is not generated by itself, obtaining the corresponding evaluation scores and generating a list of questions. A reflective agent is employed to generate optimization suggestions for each candidate carrier based on the candidate carrier information, including: Using the aforementioned reflective agent, optimization suggestions are generated for each candidate carrier based on the candidate carrier information and the corresponding problem list.
7. The method according to any one of claims 1-6, wherein, Based on the target character information, a character image package is generated, including: Based on the target character information, initial character prompts are generated; the target character information includes the target character's basic attribute information and the target character's suitable scene information; The intelligent agent is optimized using video prompts. The initial role prompts are optimized a preset number of times to obtain a preset number of optimized role prompts. Using a text-based image-based intelligent agent, a preset number of character images are generated based on the preset number of optimized character prompts; the character images include images of various poses; Using image-to-video intelligence, a preset number of character image packages are generated based on the preset number of character images.
8. The method according to any one of claims 1-6, wherein, Based on the character image package and the target script information, a resource recommendation video is generated, including: Based on the target script information and the pre-selected timbre, the target audio is synthesized; A video synthesis agent is used to synthesize resource recommendation videos corresponding to each of the target audio and a preset number of character image packages, thereby obtaining a preset number of resource recommendation videos.
9. An apparatus for generating resource recommendation videos, comprising: The acquisition module is used to obtain information about resources; The text generation module is used to generate target carrier information for recommending the resource based on the information of the resource using an intelligent agent. The target carrier information includes target role information and corresponding target script information. The image generation module is used to generate a character image package based on the target character information; The video generation module is used to generate resource recommendation videos based on the character image package and the target script information.
10. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-8.
11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.
12. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Virtual spokesman generation method and device, equipment and storage medium
CN116932798A
Role visual data generation method and device, equipment and medium
CN119205489A
Video generation method and device based on digital human, storage medium and program product
CN119277168A
Video generation method and device based on multi-agent cooperation and agents
CN120151560A
Digital human video generation method and device based on large model, intelligent agent, electronic equipment and storage medium
CN120302122A