Target video generation method and device, electronic equipment and medium
By performing digital human verification and product correspondence determination on multiple video materials, and using a big model to generate target videos, the problem of inconsistency in the display of multiple products in digital human live broadcasts is solved, and the audience experience and product display effect are improved.
Patent Information
- Application Number
- CN202510400120.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
The display of multiple products in digital people's live broadcast often needs to be switched in different shots or scenes, resulting in inconsistent displays and affecting the audience experience and product display effect.
By verifying the multi-segment video material, we determine whether the digital person is the same digital person, and based on product information and correspondence, we use a big model to generate target videos to adapt to multi-product display scenarios and adjust the image characteristics of the digital person.
The generated target video can maintain coherence and adaptability in multi-product display scenarios, improving audience experience and product display effects.
Smart Images

Figure CN120264070A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, in particular to the fields of digital humans, large models, data transmission, and information recommendation technology, and specifically to a method, device, electronic device, computer-readable storage medium, and computer program product for generating a wooden target video. Background Art
[0002] Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0003] Digital human live broadcast is a virtual anchor technology driven by AI, which combines digital images with voice and motion generation technology to interact with the audience in real time in the live broadcast room. Compared with traditional real-life anchors, digital human live broadcast can not only work around the clock, but also avoid camera fatigue, which is especially suitable for e-commerce sales and self-media content creation. However, the visual expression of digital human live broadcast videos is currently poor. For example, the connection between the display screen of product details and the overall appearance is relatively stiff, and the user experience is poor. Summary of the invention
[0004] The present disclosure provides a target video generation method, device, electronic device, computer-readable storage medium and computer program product.
[0005] According to one aspect of the present disclosure, a target video generation method is provided, comprising: in response to obtaining multiple video materials, verifying the multiple video materials to determine whether digital humans in the multiple video materials are the same digital human, wherein each of the multiple video materials includes a digital human, and wherein the digital humans in the multiple video materials have different image features; in response to determining that the digital humans in the multiple video materials are the same digital human, determining product information of one or more commodities, and a correspondence between the one or more commodities and the multiple video materials; and based on the multiple video materials, the product information and the correspondence, generating a target video through a large model, wherein the target video is used to recommend the one or more commodities through the corresponding digital humans.
[0006] According to another aspect of the present disclosure, a method for generating a target video is provided, including: receiving multiple video materials, where each of the multiple video materials includes a digital human, and the digital humans in the multiple video materials have different image features; in response to determining that the digital humans in the multiple video materials are the same digital human image, obtaining product information of one or more products and the corresponding relationship between the one or more products and the multiple video materials; and obtaining a target video, where the target video is generated by a large model based on the multiple video materials, the product information, and the corresponding relationship, and the target video is used to recommend the one or more products through the corresponding digital human.
[0007] According to another aspect of the present disclosure, a target video generation device is provided, including: a digital human verification module configured to verify the multiple video materials in response to obtaining the multiple video materials to determine whether the digital humans in the multiple video materials are the same digital human, where each of the multiple video materials includes a digital human and the digital humans in the multiple video materials have different image features; a first determination module configured to determine product information of one or more products and the corresponding relationship between the one or more products and the multiple video materials in response to determining that the digital humans in the multiple video materials are the same digital human; and a video generation module configured to generate a target video based on the multiple video materials, the product information, and the corresponding relationship by a large model, where the target video is used to recommend the one or more products through the corresponding digital human.
[0008] According to another aspect of the present disclosure, a target video generation device is provided, including: a video receiving module configured to receive multiple video materials, where each of the multiple video materials includes a digital human and the digital humans in the multiple video materials have different image features; a second determination module configured to obtain product information of one or more products and the corresponding relationship between the one or more products and the multiple video materials in response to determining that the digital humans in the multiple video materials are the same digital human; and a video obtaining module configured to obtain a target video, where the target video is generated by a large model based on the multiple video materials, the product information, and the corresponding relationship, and the target video is used to recommend the one or more products through the corresponding digital human.
[0009] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method described in the present disclosure.
[0010] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in the present disclosure.
[0011] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program which, when executed by a processor, implements the method described in the present disclosure.
[0012] According to one or more embodiments of the present disclosure, a corresponding target video is generated based on multiple video footages. Since the digital humans in the multiple video footages have different image features, by determining the correspondence between the products to be recommended and the video footages, the generated target recommended video can achieve the effect of adjusting the image features of the digital humans based on product adaptability to meet the requirements of the multi-product display scenario in the live broadcast.
[0013] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The drawings exemplarily illustrate embodiments and form a part of the description, and are used together with the written description of the description to explain the exemplary embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0015] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein can be implemented according to an embodiment of the present disclosure is shown;
[0016] Figure 2 A flowchart of a target video generation method for a server according to an embodiment of the present disclosure is shown;
[0017] Figure 3 A flowchart of a target video generation method for a client according to an embodiment of the present disclosure is shown;
[0018] Figure 4 A schematic diagram of a client display page according to an embodiment of the present disclosure is shown;
[0019] Figure 5 Shows an enlarged schematic view of area 404 in Figure 4 ;
[0020] Figure 6 Shows a block diagram of a target video generation device for a server according to an embodiment of the present disclosure;
[0021] Figure 7 Shows a block diagram of a target video generation device for a client according to an embodiment of the present disclosure; and
[0022] Figure 8 Shows a block diagram of an exemplary electronic device capable of implementing the embodiments of the present disclosure. Detailed implementation manners
[0023] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.
[0024] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.
[0025] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.
[0026] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0027] Figure 1 Shows a schematic diagram of an exemplary system 100 in which the various methods and devices described herein can be implemented according to an embodiment of the present disclosure. Refer to Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.
[0028] In embodiments of the present disclosure, the server 120 may run one or more services or software applications that enable the execution of corresponding target video generation methods.
[0029] In certain embodiments, the server 120 may also provide other services or software applications, which may include non-virtual environments and virtual environments. In certain embodiments, these services may be provided as web-based services or cloud services, such as provided to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software-as-a-service (SaaS) model.
[0030] In Figure 1 the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which may differ from the system 100. Thus, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0031] Users may use the client devices 101, 102, 103, 104, 105, and / or 106 to implement corresponding target video generation methods, input instructions, etc. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via this interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.
[0032] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.
[0033] Network 110 may be any type of network known to those skilled in the art, which may support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0034] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.
[0035] The computing units in server 120 can run one or more operating systems including any of the above operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0036] In some embodiments, server 120 can include one or more applications to analyze and combine data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0037] In some embodiments, server 120 can be a server of a distributed system, or a server incorporating a blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to address the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.
[0038] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as merchandise, video files. Databases 130 can reside in various locations. For example, the databases used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be relational databases, for example. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.
[0039] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.
[0040] Figure 1The system 100 may be configured and operated in various ways to enable the application of various methods and apparatuses described in the present disclosure.
[0041] With the development of technology, the application of digital human live broadcast in the field of e-commerce is becoming more and more extensive. Digital humans are not restricted by time and place and can continuously provide stable live broadcast services. Digital humans refer to virtual characters that exist in digital space in digital form and have the appearance, behavior and characteristics of anthropomorphic or real people. They are also called virtual images, digital virtual humans, virtual digital humans, etc.
[0042] E-commerce live broadcasts with multiple products can provide more choices and meet the diverse needs of different customers. This strategy helps to increase the chances of sales and sales. Due to the technical limitations of digital live broadcasts, the display of multiple products often needs to be switched between different shots or scenes. This may appear stiff or incoherent, which not only affects the audience's viewing experience, but may also reduce the display effect and sales conversion rate of the products.
[0043] Therefore, according to an embodiment of the present disclosure, a target video generating method for a server is provided. Figure 2 A flow chart of a target video generation method according to an embodiment of the present disclosure is shown. Figure 2 As shown, method 200 includes: in response to obtaining multiple video materials, verifying the multiple video materials to determine whether the digital people in the multiple video materials are the same digital person, wherein each of the multiple video materials includes a digital person, and the digital people in the multiple video materials have different image characteristics (step 210); in response to determining that the digital people in the multiple video materials are the same digital person, determining product information of one or more commodities, and the corresponding relationship between the one or more commodities and the multiple video materials (step 220); based on the multiple video materials, the product information and the corresponding relationship, generating a target video through a large model, wherein the target video is used to recommend the one or more commodities through the corresponding digital person (step 230).
[0044] According to the embodiments of the present disclosure, corresponding target videos are generated based on multiple video materials. Since the digital humans in the multiple video materials have different image features, by determining the corresponding relationship between the recommended products and the video materials, the generated target recommendation video can achieve the effect of adaptively adjusting the image features of the digital humans based on the products, so as to meet the needs of multiple product display scenarios in live broadcasts.
[0045] In the present disclosure, the image features of the digital human mainly refer to the appearance features of the digital human, such as actions (e.g., standing), facial expressions, clothing, and expressions, etc. These features can intuitively display the appearance image of the digital human, so as to adapt to the corresponding products. In some examples, the image features include action features. For example, when recommending products such as outdoor equipment and sports clothing, the host in a standing state may be more suitable for the current recommendation scenario; when recommending products such as office supplies, coffee machines, and cosmetics, the host in a sitting state may be more suitable for the current recommendation scenario.
[0046] Therefore, the target videos generated by digital humans with different image features can effectively adapt to the recommendation scenarios for a variety of products or in a variety of scenarios, bringing a more diverse live broadcast experience to the audience.
[0047] According to some embodiments, the multiple video materials are videos obtained by shooting based on the same scenario.
[0048] For example, the scenario information may include the background environment information where the corresponding digital human is located. For example, the multiple video materials can all be obtained by shooting in a green screen scenario or the same actual scenario. In some examples, the background environment information may further include ambient light information to ensure that the light of the multiple video materials is consistent.
[0049] In some examples, it is possible to perform image similarity calculation on the background area of the video frames corresponding to the corresponding video materials to determine whether the video materials are of the same scenario. Further, among the multiple video materials, it is possible to perform image similarity calculation on the background areas of their corresponding video frames to determine whether the multiple video materials are of the same scenario.
[0050] Through the above embodiments, the multiple video materials obtained by shooting based on the same scenario can ensure the coherence of the subsequent generated target videos and enhance the recommendation effect.
[0051] According to some embodiments, at least one of the multiple video materials is obtained based on the following operations: determining target video materials in a preset video material library, where the preset video material library includes multiple video materials corresponding to multiple digital humans; and generating new video materials corresponding to the target human image based on the target video materials and a preset target human image.
[0052] In the example, the preset video material library can be a collection of sorted and classified video resources. In this library, there are a large number of and rich video materials. These video materials can correspond to multiple different digital humans, covering a variety of scenarios, actions, expressions, and styles. For example, the video materials can show digital humans carrying out daily activities in different virtual environments, or can be clips of digital humans introducing different products. Each digital human has its unique image and characteristics.
[0053] In the above embodiment, when selecting target video materials from the preset video material library, actual requirements are considered, and factors such as the scene of the materials, the actions and expressions of the digital human are comprehensively considered. For example, if it is for the promotion of beauty products, materials with a bright scene, rich facial actions of the digital human, and capable of showing the product usage process will be selected. After determining the target video materials, they are processed in combination with the target person image. For example, specific technical means can be used to integrate the target person image into the target video materials, replacing the original digital human image in the target video materials with the target person image, so as to generate a video material that is consistent with the actions and expressions of the original digital human but matches the target person image.
[0054] In the above embodiment, by reusing the target video materials in the preset library and combining with the target person image to generate new materials, the workload and cost of material creation are greatly reduced. At the same time, video materials that meet specific requirements can be quickly produced, improving the flexibility and efficiency of material generation and better meeting the needs of diverse application scenarios.
[0055] According to an embodiment of the present disclosure, verifying the multiple video materials to determine whether the digital humans in the multiple video materials are the same digital human includes: performing a frame-by-frame operation on each of the multiple video materials to obtain multiple video frames respectively corresponding to each of the multiple video materials, where each video frame in the multiple video frames includes a digital human; and performing a face similarity verification based on the multiple video frames respectively corresponding to each video material to determine whether the digital humans in the multiple video materials are of the same digital human image.
[0056] In the example, after the frame-by-frame operation, each video material can obtain a corresponding series of video frames, and each video frame contains a digital human. Then, face similarity verification work can be carried out based on the video frames obtained by frame-by-frame. During the verification process, the face features of the digital human in each frame can be extracted, such as the shape, position of facial features, facial contour and other information, and the face features in the video frames of different video materials are compared and analyzed to calculate the face similarity between them.
[0057] In some examples, if the face similarity between at least a preset proportion (e.g., 60%) of the video frames extracted from the first video material is greater than a second threshold, it can be determined that the digital humans in the first video material are the same digital human; thereafter, one or more video frames in the first video material can be used as reference frames, and the face similarity between the video frames extracted from multiple other video materials different from the first video material and the reference frames is checked respectively. If the face similarity between at least a preset proportion of the video frames in the multiple video materials and the reference frames is greater than a third threshold, it can be determined that the digital humans in the multiple video materials are the same digital human. Among them, the second threshold and the third threshold can be the same or different.
[0058] In the above embodiment, through frame division and face similarity verification, the identity of digital humans in multiple video materials can be accurately identified, avoiding content chaos caused by the confusion of digital human images, and ensuring the coherence and professionalism of video display.
[0059] According to some embodiments, before performing frame division operations on the multiple video materials respectively, it further includes: determining the video duration corresponding to the corresponding video material in the multiple video materials; and in response to determining that the video duration is less than a first threshold, determining that the corresponding video material verification fails and generating a first prompt message.
[0060] In the example, first, it is necessary to determine the video duration of each corresponding video material in the multiple video materials to clarify its playing time span, because if the video is too short, it will lead to incomplete content or insufficient video frames for face verification. After obtaining the duration of each video material, it can be compared with the set first threshold. The first threshold is a standard duration value set according to actual needs and experience. If the duration of a certain video material is less than the first threshold, it indicates that this video material does not meet the requirements and will be determined to fail the verification. Once it is determined that the verification fails, the system will immediately generate a first prompt message to inform the user of the video material with insufficient duration in time for processing.
[0061] According to some embodiments, in response to determining that the digital humans in the multiple video materials are not the same digital human, it is determined that the verification of the multiple video materials fails and a second prompt message is generated.
[0062] In the present disclosure, correspondingly to the previous embodiment, if the digital humans in the multiple video materials are not the same digital human, at this time, the system will determine that the verification of the multiple video materials fails and issue a second prompt message so that the user can notice the problem in time and make a response.
[0063] Therefore, through duration verification, video materials that do not meet the duration requirements can be quickly screened out, avoiding wasting resources in subsequent frame segmentation and other complex processing links, and improving the overall processing efficiency. At the same time, the generated prompt information facilitates users to accurately locate the problematic materials and take timely measures to ensure the quality of video materials.
[0064] According to an embodiment of the present disclosure, in response to determining that the digital humans in the multiple video materials are the same digital human, copyright verification is performed on the multiple video materials through face recognition operations, and wherein generating the target video through the large model includes: in response to determining that the copyright verification of the multiple video materials passes, based on the multiple video materials, the commodity information, and the corresponding relationship, generating the target video through the large model.
[0065] In some examples, copyright verification can be achieved by means of face recognition operations. That is, by analyzing the facial features of the digital human and comparing them with the corresponding white list or black list, it is confirmed whether there are copyright issues with these materials, such as whether there is unauthorized use of others' images. When it is determined that the multiple video materials pass the verification, it will enter the stage of generating the target video. In this stage, the multiple video materials, commodity information, and the corresponding relationship between the commodity and the video materials can be combined, and the large model can be used to generate the target video. For example, in the e-commerce live broadcast scenario, the commodity information can be descriptive information including the characteristics and advantages of the commodity, and the corresponding relationship clearly defines which video materials can be used to display each commodity. Based on this information, the large model reasonably edits, combines, and adds special effects to the video materials to generate the target video for live streaming with goods.
[0066] According to some embodiments, in response to determining that the copyright verification of the multiple video materials fails, a third prompt information is generated.
[0067] In the example, correspondingly, once it is found that there are copyright problems with the materials, such as using unauthorized digital human images or involving infringement of others' intellectual property rights, the system will send a third prompt information, clearly informing the relevant personnel that there are copyright problems with the video materials.
[0068] In the above embodiments, copyright verification is performed through face recognition, providing guarantee for the legal dissemination and use of videos; on the other hand, using the large model to generate the target video based on various information can improve the efficiency and quality of video generation, make the generated video more in line with the commodity promotion needs, and enhance the commodity display effect.
[0069] According to some embodiments, determining the correspondence between the one or more products and the multiple video materials includes: determining one or more storyboard information respectively corresponding to each of the one or more products; and determining the video materials respectively corresponding to each storyboard information to obtain the correspondence between the one or more products and the multiple video materials.
[0070] In some examples, the storyboard information can specify in detail the specific picture content that each product needs to present in the video. For example, when showing a mobile phone, the storyboard information may include a full view of the mobile phone appearance display, a detailed display of the buttons, a screen operation display, etc. After determining the storyboard information, the corresponding video materials can be found for each storyboard information. For example, for the storyboard of the full view of the mobile phone appearance display, the video material of a digital human holding the mobile phone and showing the appearance from all directions can be selected; for the storyboard of the screen operation display, the video material of the digital human demonstrating the functions of the mobile phone can be matched. In this way, the correspondence between one or more products and multiple video materials can be accurately obtained.
[0071] In the above embodiments, the accurate determination of the correspondence based on the storyboard granularity enables each storyboard to be matched with the most suitable video material, ensuring that the products are comprehensively and prominently presented in the video, improving the product display effect, and attracting the attention of the audience.
[0072] According to some embodiments, generating a target video through a large model based on the multiple video materials, the product information, and the correspondence includes: determining whether the front and rear storyboards in the target video to be generated correspond to the same video material; in response to determining that the front and rear storyboards correspond to different video materials, generating first instruction information; and generating a target video through the large model based on the multiple video materials, the product information, the correspondence, and the first instruction information, where the first instruction information is used to guide the large model to achieve a video frame transition between the video materials corresponding to the front and rear storyboards when generating the target video.
[0073] In some examples, in order to make the generated target video more smooth and natural, it can be first determined whether the front and rear storyboards correspond to the same video material. Then, when it is determined that there are front and rear storyboards corresponding to different video materials, the corresponding first instruction information is generated based on the determined information to be used to guide the large model to achieve a video frame transition between the video materials corresponding to the front and rear storyboards when generating the target video.
[0074] Since the digital humans in the multiple video materials have different image characteristics, through the above embodiments, a smooth transition of the video frames between the front and rear video materials in the target video can be achieved, improving the quality of the generated target video.
[0075] Additionally or alternatively, according to some embodiments, generating a target video based on the multi-segment video material, the product information, and the corresponding relationship includes: obtaining the configuration information corresponding to the corresponding video material in the multi-segment video material; and generating a target video based on the multi-segment video material, the configuration information corresponding to the corresponding video material, the product information, and the corresponding relationship through a large model.
[0076] In some embodiments, the configuration information may include video parameter information, video frame configuration information, etc. For example, the configuration information may be information such as the size and / or position of the digital human in the video frame. Through this configuration information, before generating the target video, the video material can be effectively adjusted to improve the product display effect, so that the subsequent generated target video can better meet the user's requirements.
[0077] In the above embodiments of determining the corresponding relationship between the product and the video material based on the storyboard information, the configuration information may be the configuration information of the video material corresponding to the corresponding storyboard. That is, the video material can be configured based on the storyboard granularity.
[0078] In some examples, the user can generate the configuration information through the display interface. For example, the configuration information corresponding to a single storyboard or a single segment of video material can be generated, and it can also be further applied globally (i.e., the multi-segment video material).
[0079] According to an embodiment of the present disclosure, there is also provided a target video generation method for a client. Figure 3 A flowchart of the target video generation method according to an embodiment of the present disclosure is shown, as Figure 3 shown, method 300 includes: receiving multi-segment video material, wherein each segment of the multi-segment video material includes a digital human, and the digital humans in the multi-segment video material have different image characteristics (step 310); in response to determining that the digital humans in the multi-segment video material are the same digital human, obtaining the product information of one or more products and the corresponding relationship between the one or more products and the multi-segment video material (step 320); obtaining a target video, wherein the target video is generated based on the multi-segment video material, the product information, and the corresponding relationship through a large model, and the target video is used to recommend the one or more products through the corresponding digital human (step 330).
[0080] In the present disclosure, the terms and features in the target video generation method for the client have the same or similar meanings as the related terms and features in the target video generation method for the server. And, in some embodiments, they have the same or similar implementation manners as those described in the above embodiments, which will not be elaborated here.
[0081] According to some embodiments, receiving multiple segments of video material includes: in response to detecting a first operation for uploading video material, displaying a first window page for uploading video material on a display interface; and in response to detecting a second operation for uploading video material on the first window page, receiving one or more segments of video material corresponding to the second operation.
[0082] In some examples, when a user performs a video material upload operation, the system will react accordingly based on the user's behavior. For example, Figure 4 as shown, if a first operation for uploading video material is detected, such as a user clicks a button 402 for uploading video material on a display interface 401, or performs a specific shortcut key operation, a first window page 403 for uploading video material can be popped up on the display interface 401. Through the first window page 403, the user can be effectively guided and prompted to perform subsequent operations.
[0083] Continuing to refer to Figure 4 , the user can perform further operations through the first window page 403. In response to detecting a second operation for uploading video material on the first window page, such as clicking a video upload button "+" in an area 404 of the first window page 403 to upload corresponding video material. For example, the user can trigger the second operation by clicking a "Select File" button and selecting a local video file; or, the video file can be directly dragged to a specified area to trigger the second operation. In some examples, the first window page 403 may also include some production instructions for the target video, including but not limited to information such as the required video material format, memory size, shooting scene, etc., so that the user can know the upload specifications in advance, as Figure 4 shown in an area 405 of the first window page 403 in
[0084] In the above embodiments, this method of triggering the display of a window and receiving materials based on operations provides a clear and convenient upload path for users, improves the user experience, and enables users to efficiently upload video material.
[0085] According to some embodiments, during the process of receiving one or more segments of video material corresponding to the second operation, loading state information corresponding to each of the one or more segments of video material is displayed on the first window page.
[0086] In some examples, after the user finishes selecting video materials and performs an upload operation, the system starts to receive one or more segments of video materials corresponding to a second operation. During this process, the first window page can display the upload dynamics in real time, that is, the loading state information. For example, in the first window page, each segment of video material can have an exclusive loading progress identifier. For example, it can be presented in the form of a progress bar, and the progress bar is gradually filled as the video upload progresses. At the same time, text descriptions can also be provided, such as "Video 1 is uploading, 30% has been uploaded".
[0087] In some examples, multiple segments of video materials can be uploaded serially. During the serial upload of multiple segments of video materials, the loading state information of each segment of video material can be displayed separately. Figure 5 Shows the Figure 4 enlarged schematic diagram of area 404 in the first window page 403 according to an embodiment of the present disclosure, as Figure 5 shown, three segments of video materials are shown uploaded in area 404, and each video material corresponds to corresponding loading state information 406 - 408 to identify the current upload state of the corresponding video material. Through these loading state information, users can intuitively know the real-time progress of the video material upload. In the above embodiment, the display of the loading state information can enable users to have a clear understanding of the video upload process, which can not only eliminate the anxiety of users during waiting, but also this real-time feedback mechanism helps users to timely discover abnormal situations during the upload process, such as upload stagnation or upload failure, facilitating users to handle them in a timely manner and improving the overall efficiency of the video material upload work.
[0088] According to some embodiments, in response to determining that the loading state information is used to identify that the corresponding video material in the one or more segments of video materials has been received completely and detecting a preview operation for the corresponding video material, play the corresponding video material on the first window page.
[0089] Continue to refer to Figure 5 , the loading state information 407 identifies that the video material has been uploaded completely (that is, has been received completely), and at the same time identifies that the duration of the video material is 12 seconds. At this time, if a preview operation for the corresponding video material is detected, for example, the user clicks the play button of the current video material in the interface or performs a specific shortcut key operation, the video material can be played on the current first window page. Through the above operations, users can further view the content of the video material, further ensuring the video quality.
[0090] According to some embodiments, in response to determining that the loading state information is used to identify that the corresponding video material in the one or more video materials is being received, and detecting a new second operation for video material uploading on the first window page, after waiting for the one or more video materials to be received, one or more video materials corresponding to the new second operation are received.
[0091] Continuing to refer Figure 5 , the loading state information 406 indicates that the video material is being uploaded (i.e., being received). In some examples, when the video material is being received (i.e., during the video material uploading process), the upload button may be set to be non-clickable again, and after waiting for the current upload task to complete, a new upload task can be continued. In this way, excessive consumption of computing resources at the same time is prevented, resulting in the failure of receiving the video material.
[0092] According to some embodiments, the loading state information includes information for identifying that the corresponding video material in the one or more video materials fails to be received. The loading state information for identifying the failure of receiving the corresponding video material is determined by performing a video verification operation on the corresponding video material.
[0093] Continuing to refer Figure 5 , the loading state information 408 indicates that the video material upload fails (i.e., the reception fails). This may be caused by various reasons, such as current network problems, video format problems, etc. Through real-time feedback of the loading state information, it is convenient for users to process in a timely manner, improving the overall efficiency of the video material upload work.
[0094] In some embodiments, during the video material upload (i.e., reception) process, for example, the client can perform a video verification operation on the video material to ensure that all received video materials meet the preset requirements.
[0095] Therefore, according to some embodiments, receiving multiple video materials includes: in response to determining that the target video material in the multiple video materials passes the video verification operation, receiving the target video material. The video verification operation includes video parameter information verification, and the video parameter information includes at least one of the following items: video memory size, resolution, bit rate, video duration.
[0096] In some examples, video parameter information can be verified during the video upload process. For example, the video memory size can be set to be verified to be less than a preset value (such as 30M). Additionally or alternatively, the video resolution can be set to be verified to ensure that the uploaded video resolutions are the same size, and if they are inconsistent, the upload process is blocked. Additionally or alternatively, the bit rate, duration, and other parameters of the video can also be set to be verified.
[0097] In some examples, when the verification fails, corresponding prompt information can be generated and displayed on the display interface. For example, a pop-up window can be used to display the prompt information, and it can be indicated which video material fails to be uploaded successfully for what reason.
[0098] According to some embodiments, the video verification operation further includes video quantity verification, and the video verification operation is used to determine that the quantity of the received multiple video materials is within a preset quantity range.
[0099] In some examples, the quantity of the video materials corresponding to the previous second operation can be added to the quantity of the video materials corresponding to the current second operation to determine whether, after all the video materials corresponding to the current second operation are successfully uploaded, it will exceed a preset quantity threshold (such as 30 segments). In this way, through the video quantity verification, both the richness of the video materials is ensured, thereby improving the quality of the subsequent generated target video, and the consumption of computing resources by excessive video materials is prevented, thereby improving the user experience.
[0100] According to some embodiments, it further includes: in response to receiving the prompt information, displaying the prompt information on the display interface, where the prompt information is used to identify at least one of the following items: the digital humans in the multiple video materials are not the same digital human, the copyright verification of the multiple video materials fails, and the video duration corresponding to the corresponding video material in the multiple video materials is less than a first threshold.
[0101] Through this embodiment, when the corresponding prompt information is received, it can be displayed on the display interface in real time, so that the user can intuitively know the production progress of the target video, which is convenient for the user to process in time and improves the overall efficiency of the video material uploading work.
[0102] According to an embodiment of the present disclosure, as Figure 6 shown, there is also provided a target video generation device 600, including: a digital human verification module 610 configured to, in response to obtaining multiple video materials, verify the multiple video materials to determine whether the digital humans in the multiple video materials are the same digital human, where each video material in the multiple video materials includes a digital human, and the digital humans in the multiple video materials have different image characteristics; a first determination module 620 configured to, in response to determining that the digital humans in the multiple video materials are the same digital human, determine the product information of one or more products and the corresponding relationship between the one or more products and the multiple video materials; and a video generation module 630 configured to generate a target video based on the multiple video materials, the product information, and the corresponding relationship through a large model, where the target video is used to recommend the one or more products through the corresponding digital human.
[0103] Here, the operations of the above-mentioned respective units 610 to 630 of the target video generation device 600 are respectively similar to the operations of steps 210 to 230 described above, and will not be elaborated here.
[0104] According to an embodiment of the present disclosure, as Figure 7 shown, there is also provided a target video generation device 700, including: a video receiving module 710 configured to receive multiple video materials, wherein each of the multiple video materials includes a digital human, and the digital humans in the multiple video materials have different image features; a second determination module 720 configured to, in response to determining that the digital humans in the multiple video materials are the same digital human, obtain product information of one or more products and the corresponding relationship between the one or more products and the multiple video materials; and a video acquisition module 730 configured to acquire a target video, wherein the target video is generated by a large model based on the multiple video materials, the product information, and the corresponding relationship, and the target video is used to recommend the one or more products through the corresponding digital human.
[0105] Here, the operations of the above-mentioned respective units 710 to 730 of the target video generation device 700 are respectively similar to the operations of steps 310 to 330 described above, and will not be elaborated here.
[0106] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0107] According to an embodiment of the present disclosure, there is also provided an electronic device, a readable storage medium, and a computer program product.
[0108] Referring to Figure 8 , the block diagram of an electronic device 800 that can be used as a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0109] As Figure 8As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0110] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device capable of inputting information into the electronic device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 807 can be any type of device capable of presenting information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include but are not limited to a magnetic disk, an optical disk. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0111] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as method 200 or 300. For example, in some embodiments, method 200 or 300 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of method 200 or 300 described above can be executed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute method 200 or 300 by any other suitable means (e.g., by means of firmware).
[0112] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0113] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0114] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0115] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0116] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0117] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0118] It should be understood that the various forms of processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0119] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after this disclosure.
Claims
1. A method for generating a target video, comprising: In response to obtaining multiple video materials, verifying the multiple video materials to determine whether the digital humans in the multiple video materials are the same digital human, where each of the multiple video materials includes a digital human, and the digital humans in the multiple video materials have different image features; In response to determining that the digital humans in the multiple video materials are the same digital human, determining the product information of one or more products and the corresponding relationship between the one or more products and the multiple video materials; and Based on the multiple video materials, the product information, and the corresponding relationship, generating a target video through a large model, where the target video is used to recommend the one or more products through the corresponding digital human.
2. The method according to claim 1, wherein, The multiple video materials are videos obtained by shooting based on the same scene.
3. The method according to claim 1, wherein At least one of the multiple video materials is obtained based on the following operations: Determining a target video material in a preset video material library, where the preset video material library includes multiple video materials corresponding to multiple digital humans; And Based on the target video material and a preset target person image, generating a new video material corresponding to the target person image.
4. The method according to claim 1, wherein, Verifying the multiple video materials to determine whether the digital humans in the multiple video materials are the same digital human includes: Performing frame splitting operations on the multiple video materials respectively to obtain multiple video frames corresponding to each of the multiple video materials, where each video frame in the multiple video frames includes a digital human; and Based on the multiple video frames corresponding to each video material respectively, performing face similarity verification to determine whether the digital humans in the multiple video materials have the same image.
5. The method according to claim 4, wherein, Before performing the frame splitting operations on the multiple video materials respectively, it further includes: Determining the video duration corresponding to the corresponding video material in the multiple video materials; and In response to determining that the video duration is less than a first threshold, determining that the corresponding video material verification fails and generating a first prompt message.
6. The method according to claim 1, further comprising: In response to determining that the digital humans in the multiple video materials are not the same digital human, determining that the verification of the multiple video materials fails and generating a second prompt message.
7. The method according to any one of claims 1-6, further comprising: In response to determining that the digital humans in the multiple video materials are the same digital human, performing copyright verification on the multiple video materials through face recognition operations, and Generating a target video through a large model includes: in response to determining that the copyright verification of the multiple video materials passes, generating the target video through a large model based on the multiple video materials, the product information, and the corresponding relationship.
8. The method according to claim 7, further comprising: In response to determining that the copyright verification of the multiple video materials fails, generating a third prompt message.
9. The method according to claim 1, wherein, Determining the corresponding relationship between the one or more products and the multiple video materials includes: Determine one or more storyboard information respectively corresponding to each of the one or more commodities; and Determine the video materials respectively corresponding to each storyboard information to obtain the corresponding relationship between the one or more commodities and the multiple video materials.
10. The method according to claim 9, wherein, Generating a target video through a large model based on the multiple video materials, the commodity information, and the corresponding relationship includes: Determine whether the two consecutive storyboards in the target video to be generated correspond to the same video material; In response to determining that the two consecutive storyboards correspond to different video materials, generate a first instruction message; and Generate a target video through a large model based on the multiple video materials, the commodity information, the corresponding relationship, and the first instruction message, where the first instruction message is used to guide the large model to achieve the video frame transition between the video materials corresponding to the two consecutive storyboards when generating the target video.
11. The method according to claim 1 or 9 or 10, wherein Generating a target video through a large model based on the multiple video materials, the commodity information, and the corresponding relationship includes: Obtain the configuration information corresponding to the corresponding video material in the multiple video materials; and Generate a target video through a large model based on the multiple video materials, the configuration information corresponding to the corresponding video material, the commodity information, and the corresponding relationship.
12. A method for generating a target video, including: Receive multiple video materials, where each of the multiple video materials includes a digital human, and the digital humans in the multiple video materials have different image features; In response to determining that the digital humans in the multiple video materials are the same digital human, obtain the commodity information of one or more commodities, and the corresponding relationship between the one or more commodities and the multiple video materials; and Obtain a target video, where the target video is generated through a large model based on the multiple video materials, the commodity information, and the corresponding relationship, and the target video is used to recommend the one or more commodities through the corresponding digital human.
13. The method according to claim 12, wherein, Receiving multiple video materials includes: In response to detecting a first operation for uploading video materials, display a first window page for uploading video materials on the display interface; and In response to detecting a second operation for uploading video materials on the first window page, receive one or more video materials corresponding to the second operation.
14. The method according to claim 13, further including: During the process of receiving one or more video materials corresponding to the second operation, display the loading state information corresponding to each of the one or more video materials on the first window page.
15. The method according to claim 14, further including: In response to determining that the loading state information is used to indicate that the corresponding video material in the one or more video materials has been received and detecting a preview operation for the corresponding video material, play the corresponding video material on the first window page.
16. The method according to claim 14 or 15, where In response to determining that the loading state information is used to identify that the corresponding video material in the one or more video materials is being received, and detecting a new second operation for video material uploading on the first window page, after waiting for the one or more video materials to be received, receive the one or more video materials corresponding to the new second operation.
17. The method according to claim 14, wherein The loading state information includes information for identifying that the corresponding video material in the one or more video materials fails to be received, and wherein the loading state information for identifying the failure of the corresponding video material to be received is determined by performing a video verification operation on the corresponding video material.
18. The method according to claim 12 or 17, wherein Receiving multiple video materials includes: In response to determining that the target video material in the multiple video materials passes the video verification operation, receive the target video material, wherein, The video verification operation includes video parameter information verification, and the video parameter information includes at least one of the following items: video memory size, resolution, bit rate, video duration.
19. The method according to claim 17 or 18, wherein, The video verification operation further includes video quantity verification, and the video verification operation is used to determine that the quantity of the multiple video materials received is within a preset quantity range.
20. The method according to claim 12, further comprising: In response to receiving a prompt message, display the prompt message on the display interface, wherein the prompt message is used to identify at least one of the following items: The digital humans in the multiple video materials are not the same digital human, the copyright verification of the multiple video materials fails, and the video duration corresponding to the corresponding video material in the multiple video materials is less than a first threshold.
21. A target video generation device, comprising: A digital human verification module, configured to, in response to obtaining multiple video materials, verify the multiple video materials to determine whether the digital humans in the multiple video materials are the same digital human, wherein each of the multiple video materials includes a digital human, and wherein the digital humans in the multiple video materials have different image characteristics; A first determination module, configured to, in response to determining that the digital humans in the multiple video materials are the same digital human, determine the product information of one or more products and the corresponding relationship between the one or more products and the multiple video materials; and A video generation module, configured to generate a target video based on the multiple video materials, the product information, and the corresponding relationship through a large model, wherein the target video is used to recommend the one or more products through the corresponding digital human.
22. A target video generation device, comprising: A video receiving module, configured to receive multiple video materials, wherein each of the multiple video materials includes a digital human, and wherein the digital humans in the multiple video materials have different image characteristics; A second determination module, configured to, in response to determining that the digital humans in the multiple video materials are the same digital human, obtain the product information of one or more products and the corresponding relationship between the one or more products and the multiple video materials; and A video acquisition module, configured to acquire a target video, wherein the target video is generated by a large model based on the multiple video materials, the commodity information, and the corresponding relationship, and the target video is used to recommend the one or more commodities through a corresponding digital human.
23. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; Wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 1-20.
24. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-20.
25. A computer program product, comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-20.
Citation Information
Cited By
Digital human interaction method and system based on multi-mode sensing intelligent action switching
CN121050590A