Generation method of introduction video and generation device of introduction video

Automatically generate APP introduction videos through multimodal large models and digital human technology, solving the inefficiency problems caused by manual operations in the existing technology, and realizing intelligent video production and high-quality content generation.

CN120409434APending Publication Date: 2025-08-01CHINA UNICOM (GUANGDONG) IND INTERNET CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510470814.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, the production process of APP product function introduction relies on manual operations, resulting in fragmented workflows and inefficient efficiency, lack of intelligent support, making it difficult to achieve efficient content iteration and innovation.

Method used

Multimodal large model is used to analyze multiple interfaces of the target application, generate introduction copy including text and pictures, and generate introduction videos through digital human technology to reduce manual intervention and achieve automated generation.

Benefits of technology

It improves the intelligence level of introduction video production, improves efficiency and quality, generates more exquisite and high-quality introduction videos, reduces human errors, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409434A_ABST
    Figure CN120409434A_ABST
Patent Text Reader

Abstract

According to the introduction video generation method and the introduction video generation device disclosed by the invention, each of a plurality of interfaces of a target application is analyzed mainly by adopting a multi-modal large model so as to generate an introduction document of the target application, and the introduction document comprises a character part and a picture part; and according to the introduction copywriting, an introduction video of the target application is generated, and the introduction video comprises an interface picture of each interface and a voice introduction corresponding to each interface. According to the method, the introduction video of the application product is generated by introducing the multi-modal large model technology, and the intelligent level of introduction video production is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information processing, and particularly relates to a method for generating an introduction video and a generating device for the introduction video. Background Art

[0002] The function introduction of APP products is not only a tool for demonstrating product functions, but also a core link for attracting users, improving user experience, shaping brand image, and enhancing market competitiveness. It is directly related to the market performance and user loyalty of APPs, so the quality of its design and content is crucial.

[0003] In the prior art, during the production process of APP product function introductions, various tasks are usually completed by relying on manual operations, such as interface recording, copywriting creation, and digital human driving. And different tools are required for each link, which leads to fragmentation and low efficiency of the work process, resulting in insufficient intelligence in the overall production process of APP product function introductions. Summary of the Invention

[0004] Embodiments of the present application disclose a method for generating an introduction video and a generating device for the introduction video, which improve the intelligence level of introduction video production by introducing multi-modal large model technology.

[0005] In a first aspect of the embodiments of the present application, a method for generating an introduction video is disclosed. A multi-modal large model is used to parse each interface of multiple interfaces of a target application to generate an introduction copy for the target application. The introduction copy includes a text part and a picture part. The text part is used to describe each interface and each functional component of each interface, and the picture part includes interface pictures of each interface.

[0006] According to the introduction copy, an introduction video for the target application is generated. The introduction video includes interface pictures of each interface and a voice introduction corresponding to each interface.

[0007] As an optional implementation manner, in the first aspect of this embodiment, the using a multi-modal large model to parse each interface of multiple interfaces of a target application to generate an introduction copy for the target application includes:

[0008] Using a multi-modal large model to parse each interface of multiple interfaces of a target application to obtain a description text, where the description text is used to describe each interface and each functional component of each interface;

[0009] According to the prompt word template corresponding to the description text, guiding the multi-modal large model to generate the introduction copy based on the description text.

[0010] As an alternative implementation, in the first aspect of this embodiment, a multi-modal large model is used to parse each interface of the target application to obtain a description text, where the description text is used to describe each interface and each functional component of each interface, including:

[0011] Using the multi-modal large model to obtain the interface picture of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface, where the functional characteristics of each interface are used to represent the functions that each interface is used to implement, and the interaction logic of each functional component is used to represent the interaction situation of each functional component with other interfaces;

[0012] Based on the interface picture of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface, the description text is obtained.

[0013] As an alternative implementation, in the first aspect of this embodiment, the step of using the multi-modal large model to obtain the interface picture of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface includes:

[0014] Using the multi-modal large model to obtain the interface picture of each interface;

[0015] Analyzing the interface picture to obtain each functional component of each interface;

[0016] Performing simulation operations on each functional component of each interface;

[0017] In response to the simulation operations, obtaining the functional characteristics of each interface and the interaction logic of each functional component of each interface.

[0018] As an alternative implementation, in the first aspect of this embodiment, the text part includes the description content of each interface, where the description content of each interface includes the function introduction of each interface and the function introduction of each functional component of each interface, and the description content of each interface corresponds to the interface picture of each interface.

[0019] As an alternative implementation, in the first aspect of this embodiment, the text part further includes: the highlights of the usage experience of the target application, and / or, the introduction of the operation logic of the target application.

[0020] As an alternative implementation, in the first aspect of this embodiment, the step of generating the introduction video of the target application based on the introduction copy includes:

[0021] Generate video content according to the described introduction copy;

[0022] Generate a voice introduction corresponding to each interface according to the text part;

[0023] Generate the introduction video according to the video content and the voice introduction corresponding to each interface.

[0024] As an optional implementation manner, in the first aspect of this embodiment, the generating video content according to the described introduction copy includes:

[0025] Generate a picture video by arranging the picture part in the interface order of the multiple interfaces and the operation logic of the target application;

[0026] Fuse the text part with the picture video to generate the video content.

[0027] As an optional implementation manner, in the first aspect of this embodiment, the generating a voice introduction corresponding to each interface according to the text part includes:

[0028] Generate a voice introduction corresponding to each interface according to the text part through the digital human technology, and the voice introduction is a voice introduction with a virtual digital human.

[0029] The second aspect of the embodiment of the present application discloses a generating device for an introduction video, including a copywriting generation module and a video generation module, where:

[0030] The copywriting generation module is used to parse each interface of the target application by using a multimodal large model to generate an introduction copy of the target application. The introduction copy includes a text part and a picture part. The text part is used to describe each interface and each functional component of each interface. The picture part includes the interface pictures of each interface;

[0031] The video generation module is used to generate an introduction video of the target application according to the introduction copy. The introduction video includes the interface pictures of each interface and the corresponding voice introductions.

[0032] Compared with the related art, the embodiment of the present application at least includes the following beneficial effects:

[0033] A method for generating an introduction video disclosed in an embodiment of the present application mainly parses each interface of a target application by using a multimodal large model to generate an introduction text for the target application. The introduction text includes a text part and a picture part. The text part is used to describe each interface and each functional component of each interface, and the picture part includes interface pictures of each interface. Then, according to the introduction text, an introduction video for the target application is generated, and the introduction video includes the interface pictures of each interface and the voice introduction corresponding to each interface. By introducing the multimodal large model technology to parse each interface of the target application and automatically generate the introduction text of the application, and then generate the introduction video of the application. In this way, the manual intervention link is reduced, and the automatic generation of the introduction video can be realized, that is, the automatic process from the interface parsing of the application to the output of the introduction video, which improves the intelligent level of the production of the introduction video. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0035] Figure 1 It is a schematic structural diagram of a device for generating an introduction video provided by an embodiment of the present application;

[0036] Figure 2 It is a schematic structural diagram of another method for generating an introduction video provided by an embodiment of the present application;

[0037] Figure 3 It is a schematic structural diagram of another method for generating an introduction video provided by an embodiment of the present application;

[0038] Figure 4 It is a schematic structural diagram of another method for generating an introduction video provided by an embodiment of the present application;

[0039] Figure 5 It is a schematic flowchart of a method for generating a description text for each interface provided by an embodiment of the present application;

[0040] Figure 6 It is a schematic structural diagram of another method for generating an introduction video provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts belong to the scope of protection of the present application. <{} <{}

[0042] It should be noted that the terms "first / second / third" involved in the embodiments of the present application are used to distinguish similar or different objects, and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here. <{} <{}

[0043] It should be noted that the terms "include" and "have" and any variations thereof in the embodiments of the present application and the accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. <{} <{}

[0044] The function introduction of APP products plays a crucial role in modern application development and promotion. It is not just a simple function display tool, but also an important bridge to attract user attention, convey brand value, and enhance user experience. Through clear and concise function introduction, users can quickly understand the core value and actual use of the APP, which is crucial for attracting the interest of potential users. At the same time, a good function introduction can help users quickly master the operation method of the APP, improve their usage experience, reduce the learning cost, and thus increase their stickiness and usage frequency. <{} <{}

[0045] In the current technological environment, during the production process of APP product function introductions, most steps still rely on manual operations and lack sufficient intelligent support. For example, in interface recording, usually professional personnel need to use independent screen recording tools for operation. This process is not only cumbersome but also error-prone. Especially when multiple functions or complex interactions need to be demonstrated, the recording personnel must constantly adjust the screen and scene switching, increasing the complexity of the operation. In copywriting creation, it is still necessary to manually write accurate text descriptions to explain each function point. This process requires creative personnel to have strong language expression abilities and a profound understanding of the product, and often requires multiple revisions and optimizations. In addition, in each production link, technicians need to use different tools to complete, and these tools often do not have good compatibility, resulting in an increase in the switching and communication costs between tools. It can be seen that due to the lack of intelligent and automated support in the current technology, not only does it take a lot of time to complete various tasks, affecting the production speed and quality, but it is also very difficult to achieve efficient content iteration and innovation, restricting the exertion of the team's creative potential.

[0046] Based on this, the embodiments of the present application disclose a method for generating an introduction video, which improves the intelligent level of the production of the introduction video of the application product by introducing multi-modal large model technology.

[0047] The method for generating an introduction video disclosed in the embodiments of the present application can be applied to a variety of fields, and can help various institutions and enterprises improve production efficiency, reduce costs, enhance creative expressiveness, and provide users with a more interactive and personalized experience. For example, the e-commerce industry, education and online learning, fitness and health, finance and investment, social media and communication, tourism and travel, the gaming field, enterprise management and productivity tools, and shopping and coupons, etc., are not specifically limited here.

[0048] For the above-mentioned introduction video generation methods in different fields, the following will be described in detail. For example, in the e-commerce industry, this introduction video generation method can be applied to e-commerce platform APPs, shopping process guides, and promotional activity publicity, etc.; in the field of education and online learning, this introduction video generation method can be applied to course function introductions, learning method demonstrations, and learning feedback and success case displays in educational APPs, etc.; in the fitness and health field, this introduction video generation method can be applied to fitness plan displays and health data tracking displays in fitness APPs, etc.; in the finance and investment field, this introduction video generation method can be applied to function introductions, investment skills, and education displays for investment APPs, etc.; in the field of social media and communication, this introduction video generation method can be applied to the display of social functions and interactive displays in APPs, etc.; in the field of tourism and travel, this introduction video generation method can be applied to itinerary planning function displays, travel guides, and transportation bookings in travel APPs, etc.; in the gaming field, this introduction video generation method can be applied to gameplay demonstrations and social interaction function demonstrations in gaming APPs, etc.; in the field of enterprise management and productivity tools, this introduction video generation method can be applied to the display of collaboration functions in enterprise-level APPs, etc.; in the shopping and coupon field, this introduction video generation method can be applied to the display of discount functions and shopping guides and recommendations in discount APPs, etc.

[0049] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a device for generating an introduction video provided by an embodiment of the present application. The introduction video generation device 10 includes a copywriting generation module 11 and a video generation module 21. The copywriting generation module 11 is connected to the video generation module 21. The copywriting generation module 11 is mainly used to generate an introduction copy of the target application by using a multimodal large model. The video generation module 21 is mainly used to generate an introduction video of the target application according to the introduction copy. To understand the above-mentioned introduction video generation method more clearly, the embodiment of the present application is described with the introduction video generation device as the execution subject. It should be understood that the execution subject of the embodiment of the present application can be a chip or a processor in an electronic device including the introduction video generation device, etc., which is not specifically limited here.

[0050] Please refer to Figure 2 , Figure 2 which is a flowchart of an introduction video generation method provided by an embodiment of the present application. The flowchart includes at least steps S101 - S102.

[0051] Step S101: Parse each interface of multiple interfaces of the target application by using a multimodal large model to generate an introduction copy of the target application.

[0052] Among them, the introduction copy includes a text part and a picture part. The text part is used to describe each interface and each functional component of each interface. The picture part includes the interface pictures of each interface, and the interface pictures can be screenshots of the interface during the running of the target application.

[0053] In some embodiments, the text part of the introduction copy includes the description content of each interface. The description content of each interface includes the function introduction of each interface and the function introduction of each functional component of each interface. The description content of each interface corresponds to the interface picture of each interface, or in other words, the description content of each interface matches the interface picture of each interface, so as to provide a visual contrast for the description content of each interface.

[0054] Take as an example to make an exemplary description of the function introductions of different interfaces, the function introductions of each functional component of each interface, the operation logic, and the highlights of the usage experience.

[0055] Exemplarily, taking the chat interface of as an example, the function introduction of the interface may include: The main interface of

[0056] is used to display the user's chat records, including functions such as private chats with friends, group chats, and other interactions. This is the main interface when most users use daily.

[0057] Exemplarily, the functional components of the chat interface of may include a chat list and a bottom menu bar. The bottom menu bar may include functional components such as WeChat, Contacts, Discover, and Me. Below, the function introductions of these functional components will be described. Chat list: Used to display all conversation records with friends and groups, sorted by the latest message. The number of unread messages will be displayed next to each message to remind the user to view new information. Top search box: Used to allow the user to quickly search for contacts, group chats, chat records, etc., improving the usage efficiency. Message reminder: Used to display the number of new unread messages at the top or bottom of the interface, allowing the user to quickly identify unviewed messages.

[0057] WeChat: Used to click to enter the social function page to view friends' dynamics, moments, etc.

[0058] Contacts: Used to view the contact list, add new friends or group chats.

[0059] Discover: Used to view functions such as moments, scan QR code, shake, etc.

[0060] Me: Used to enter the personal center to set personal information, view payment, wallet, emojis, etc.

[0061] Optionally, the text part of the introduction copy can also include highlights of the usage experience of the target application and / or an introduction to the operation logic of the target application.

[0062] Among them, the highlights of the usage experience of the target application refer to the key elements or functional features that can significantly enhance the user experience during the design and development of the application. These highlights usually include those designs and functions that can make users feel convenient, pleasant, efficient, and meet their needs. For example, a simple and intuitive interface design, a fast response and smooth operation experience, intelligent and personalized recommendations, seamless cross-device synchronization, etc., are not specifically limited here.

[0063] Exemplarily, The highlights of the usage experience of the chat interface can include the following.

[0064] Quick switching: Users can quickly switch to other pages, such as the Moments and the contact list, through the bottom menu bar, improving the convenience of operation.

[0065] Search function: The search box at the top helps users quickly locate the required information in the large amount of chat records, saving time.

[0066] Message reminder: Unread messages are marked by numbers and push notifications to ensure that users do not miss important information.

[0067] In the embodiments of the present application, the overall operation logic of the target application refers to the framework and process of how the internal functional modules, interfaces, and processes of the application cooperate with each other when the user interacts with the application to provide the services or functions required by the user. Its goal is to help users complete tasks efficiently and improve the user experience through a smooth and intuitive process.

[0068] The overall operation logic of the target application can include user startup of the application, user registration and login, main interface display and function navigation, operation process and task execution, real-time interaction and feedback, and user logout and cancellation, etc., which are not specifically limited here.

[0069] Exemplarily, The overall operation process of [application name] can include: The overall operation process of [application name] starts from entering the chat interface, with chatting as the core, and expands around functions such as communication, payment, and social interaction. After operating all the interfaces and all the functional components of the interfaces in [application name], the overall process is displayed. After operating all the interfaces and all the functional components of the interfaces in [application name], the overall process is displayed.

[0070] The text part of the introduction copy in this method includes the description content of each interface. The description content of each interface includes the function introduction of each interface and the function introduction of each functional component of each interface. It can also include the highlights of the usage experience of the target application or the introduction of the operation logic of the target application. The text part of this introduction copy includes so much information that it can comprehensively and systematically display the functions and usage experience of the target application. At the same time, through the combination of interface pictures and text, it enhances the user's understanding and attraction to the target application, and improves the clarity and effect of information transmission.

[0071] In the embodiments of this application, multiple interfaces refer to all the interfaces that can be displayed in the target application, which can include multi-level interfaces, such as first-level interfaces, second-level interfaces, third-level interfaces, and more levels of interfaces. Among them, the lower-level interface is the sub-interface of the upper-level interface, and the upper-level interface is the parent interface of the lower-level interface. For example, the first-level interface is the parent interface of the second-level interface. Correspondingly, the second-level interface is the sub-interface of the first-level interface. For another example, the second-level interface is the parent interface of the third-level interface. Correspondingly, the third-level interface is the sub-interface of the second-level interface. Any lower-level interface is generated based on a functional component of the upper-level interface. Specifically, the lower-level interface is the interface displayed after the functional component of the upper-level interface is triggered.

[0072] For example, taking the chat interface as the first-level interface as an example, the second-level interface of can be entering the chat interface of a certain friend by clicking on the chat list, and the third-level interface of

[0073] can be the collection interface jumped to by clicking on the collection control at the bottom.

[0074] Optionally, the multimodal large model is arbitrary. For example, it can include CLIP, VisualBERT, Flamingo, BERT, METER, and UNITER, etc. There is no specific limitation here, and it can be selected according to actual needs.

[0075] A multimodal large model can include aspects such as data preprocessing, modality fusion, and multi-task learning. There is no specific limitation here. Among them:

[0076] Preprocessing refers to the need for appropriate preprocessing of data in different modalities to extract their effective information. For example, image data is usually subjected to feature extraction through a convolutional neural network, while text data is processed through natural language processing techniques such as Transformer or BERT for vocabulary embedding and semantic representation. Audio data may extract spectral features through an acoustic model for subsequent analysis.

[0077] Modal fusion refers to fusing data from different modalities together for joint learning and reasoning. There are usually two ways of modal fusion: early fusion and late fusion. Early fusion is to merge the features of different modalities together for unified processing before the data is input into the model, while late fusion is to first process the features of each modality separately and then merge the processing results for decision-making. The way of modal fusion can be selected according to actual needs and is not specifically restricted here.

[0078] Multi-task learning means that the model processes different tasks by sharing parameters and feature representations, which can not only improve efficiency but also help the model acquire more knowledge from multiple tasks, thereby enhancing the generalization ability of the model to improve its performance in specific tasks.

[0079] Step S102: Generate an introduction video for the target application according to the introduction copy.

[0080] Among them, the introduction video includes interface pictures of each interface and the corresponding voice introductions for each interface. During the playback of the introduction video, each interface picture is played synchronously with the corresponding voice introduction.

[0081] In this method, a multi-modal large model is used to automatically generate the introduction copy and automatically generate the introduction video according to the introduction copy. This method can parse through the multi-modal large model interface, thus reducing manual operation intervention, automatically generating the introduction video, and outputting the introduction video. The manual intervention link is reduced, thereby improving the efficiency of the introduction video production method and realizing a more intelligent introduction video generation method; moreover, using a multi-modal large model to generate the introduction video, since this multi-modal large model can process and fuse information from various types, it can provide a more comprehensive and accurate understanding and decision-making, generating a more exquisite and high-quality introduction video, avoiding the uneven quality of the introduction video caused by the limited level of people during manual operation.

[0082] In some embodiments, before parsing each interface of a target application using a multimodal large model to generate an introduction text for the target application, it is also necessary to connect the multimodal large model to the application programming interface of the target application so that the generation device for the introduction video can perform the operation of generating an introduction video for the target application. By connecting the multimodal large model to the application programming interface of the target application, the production of the introduction video for the target application can be achieved without complex operations, making the production of the introduction video for the target application more efficient and convenient.

[0083] In some embodiments, in order to make the generated introduction text more in line with the expected effect, the introduction text can be generated according to a prompt template, so that by using clear and structured prompts, deviations and misunderstandings can be reduced, enabling the model to accurately identify the key parts of the problem, thereby helping the model better understand the task requirements and generating more vivid and fluent introduction texts of different types and styles according to the production requirements.

[0084] To further understand this method, please refer to Figure 3 , Figure 3 which is a schematic flowchart of another method for generating an introduction video provided by an embodiment of the present application. The flowchart includes at least steps S201 - S203:

[0085] Step S201: Parse each interface of the target application using a multimodal large model to obtain a description text;

[0086] wherein the description text is used to describe each interface and each functional component of each interface;

[0087] Step S202: Guide the multimodal large model to generate an introduction text based on the description text according to the prompt template corresponding to the description text.

[0088] Step S203: Generate an introduction video for the target application according to the introduction text;

[0089] The introduction video includes an interface picture of each interface and a voice introduction corresponding to each interface.

[0090] Exemplarily, when the multimodal large model analyzes the description text to generate a corresponding prompt template according to the characteristic type and style. For APP products in the field of leisure and entertainment, a prompt template with entertainment properties in this field can be used to make the generated introduction style language more vivid, lively, humorous and interesting, improving people's usage pleasure; for another example, for the field of shopping and consumption, a prompt template highlighting aspects such as rich commodities, price discounts, shopping convenience, and user experience in this field can be used to attract consumers to download and use the application.

[0091] Optionally, the entertaining prompt templates can include interest-arousing type, function demonstration type, or user experience type prompt templates, without specific restrictions here, and can be selected according to actual needs.

[0092] Among them, the purpose of the interest-arousing type template is to directly arouse users' interest and stimulate their curiosity through concise language.

[0093] For example: "Want to relax every day? [APP name] will take you into the world of relaxation and entertainment!"

[0094] "Turn boring time into wonderful time, [APP name] brings unlimited entertainment experience!"

[0095] The function demonstration type template helps users understand what kind of entertainment experience it can provide.

[0096] For example: "Whether it's games, movies, or music, [APP name] can meet all your needs!"

[0097] "Updated in real time, providing the hottest film and television content, games, and music! At [APP name], entertainment is everywhere!"

[0098] The user experience type template focuses on enhancing the user experience and conveying the pleasure after using the APP.

[0099] For example: "Easy to click, immediately enjoy the unparalleled entertainment experience, and start relieving stress from [APP name]!"

[0100] "Immersive experience, as if being on the scene, giving you a different sensory enjoyment! Come and try it!"

[0101] Optionally, the templates for the consumption field can include discount promotion type, rich commodity type templates, and user evaluation and trust type, etc., without specific restrictions here, and can be selected according to actual needs.

[0102] Among them, the discount promotion type prompt word layout mainly highlights discounts, preferential activities, and limited-time promotions, aiming to make users feel that the cost performance during the purchase process is very high.

[0103] For example: "Global good products, ultra-low prices! At [APP name], enjoy discounts and buy all over the network!"

[0104] "First order with a direct discount, and get a lot of coupons for free! Come to [APP name] and enjoy continuous shopping!"

[0105] The prompt word templates of the rich commodity type mainly emphasize that the APP has a rich variety of commodities and diverse choices, meeting various needs of consumers.

[0106] For example: "From home appliances to fashion and beauty, [APP name] provides you with a full-category shopping experience!"

[0107] "Whatever you want to buy, [APP name] has it all! Hundreds of brands for you to choose from!"

[0108] The user evaluation and trust type template can display users' evaluations, word-of-mouth, and trust levels to enhance the trust of potential users.

[0109] For example: "Real evaluations, transparent shopping, [APP name] makes your shopping more reassuring!"

[0110] "Highly recommended by millions of users, a trustworthy shopping platform! Come and join us on [APP name]!"

[0111] In the above step S201, that is, in the method step of parsing each interface of the target application using a multimodal large model to obtain the description text, in some embodiments, the multimodal large model is used to obtain the interface picture of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface. The functional characteristics of each interface are used to represent the functions that each interface is used to implement, and the interaction logic of each functional component is used to represent the interaction situation of each functional component with other interfaces; according to the interface picture of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface, a description text is generated.

[0112] To understand this method step more clearly, please refer to Figure 4 , Figure 4 which is a flowchart of another method for generating an introduction video provided by an embodiment of the present application. The flowchart includes at least steps S301 - S304:

[0113] Step S301: Use a multimodal large model to obtain the interface picture of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface;

[0114] Among them, the functional characteristics of each interface are used to represent the functions that each interface is used to implement, and the interaction logic of each functional component is used to represent the interaction situation of each functional component with other interfaces.

[0115] Step S302: Generate a description text according to the interface picture of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface.

[0116] Step S303: According to the prompt word template corresponding to the description text, guide the multimodal large model to generate an introduction copy based on the description text.

[0117] Step S304: Generate an introduction video of the target application according to the introduction copy.

[0118] Among them, the introduction video includes the interface pictures of each interface and the voice introduction corresponding to each interface.

[0119] In this method, not only can more abundant relevant information about the target application be obtained by acquiring the interface pictures of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface, so as to generate more accurate and high-quality description text, but also the generated introduction copy can be further made to better meet the expected effect through the prompt template, thereby generating a higher-quality introduction video.

[0120] In some other embodiments, the description text can also be used as the introduction copy to generate an introduction video. This method may include: using a multimodal large model to obtain the interface pictures of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface; where the functional characteristics of each interface are used to represent the functions implemented by each interface, and the interaction logic of each functional component is used to represent the interaction situation of each functional component with other interfaces; generating an introduction copy according to the interface pictures of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface, where the introduction copy is the description text above; generating an introduction video of the target application according to the introduction copy.

[0121] In this method, the interface pictures of each interface of the target application, the functional characteristics of each interface, and the interaction logic of each functional component of each interface are obtained, so as to obtain more abundant relevant information about the target application, so as to generate more accurate and high-quality introduction copy, and thus the generated introduction video can also contain rich and high-quality information.

[0122] In the above step S301, that is, in the step of using a multimodal large model to obtain the interface pictures of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface, in some embodiments, use a multimodal large model to obtain the interface pictures of each interface; analyze the interface pictures to obtain each functional component of each interface; perform simulation operations on each functional component of each interface; in response to the simulation operations, obtain the functional characteristics of each interface and the interaction logic of each functional component of each interface.

[0123] To understand the above steps more clearly, please refer to Figure 5 , Figure 5 which is a schematic flowchart of a method for generating a description text for each interface provided by an embodiment of the present application. The flowchart at least includes steps S401 - S404:

[0124] Step S401: Use a multimodal large model to obtain the interface pictures of each interface;

[0125] Step S402: Analyze the interface pictures to obtain various functional components of each interface;

[0126] Step S403: Perform simulated operations on various functional components of each interface;

[0127] Among them, the simulated operation is to simulate the operation of a real user using the target application.

[0128] Optionally, the simulated operation may include operations such as swiping, dragging, and clicking. There is no specific limitation here.

[0129] Step S404: In response to the simulated operation, obtain the functional characteristics of each interface and the interaction logic of various functional components of each interface.

[0130] Step S405: Generate a description text based on the interface pictures of each interface, the functional characteristics of each interface, and the interaction logic of various functional components of each interface.

[0131] Optionally, in steps S402, S403, and S404, the operations are not strictly carried out in accordance with the above process steps. Exemplarily, the steps of analyzing the interface pictures to obtain the components in each interface and performing simulated operations on various functional components in the interface to obtain the interaction logic of various functional components of each interface can be carried out simultaneously.

[0132] In this method, by simulating user operations, the functional characteristics of each interface and the interaction logic of various functional components of each interface are automatically obtained, so as to generate an application introduction efficiently and accurately, reduce manual intervention, and improve the video production efficiency.

[0133] In step S102, that is, in the step of generating an introduction video of the target application according to the introduction copywriting, in some embodiments, a voice introduction corresponding to each interface is generated according to the text part; an introduction video is generated according to the video content and the voice introduction corresponding to each interface. ]>

[0134] It can be understood that various video processing technologies can be used in this method, such as video analysis, video splicing and synthesis, video stream processing, virtualization and augmented reality, video voice synchronization, and digital human technology, etc. There is no specific limitation here.

[0135] In some embodiments, a voice introduction corresponding to each interface can be generated according to the text part through digital human technology.

[0136] To understand the above method steps more clearly, please refer to Figure 6 , Figure 6 which is a schematic structural diagram of another introduction video generation method provided by an embodiment of the present application. The flowchart includes at least steps S501 - S503:

[0137] Step S501: Parse each interface of the target application using a multi-modal large model to generate an introduction text for the target application.

[0138] Among them, the introduction text includes a text part and a picture part. The text part is used to describe each interface and each functional component of each interface. The picture part includes the interface pictures of each interface. For the specific description of this step, reference can be made to the relevant descriptions of the above embodiments, and details will not be repeated.

[0139] Step S502: Generate video content based on the introduction text.

[0140] In some embodiments, the picture part is used to generate a picture video according to the interface order of the multiple interfaces and the operation logic of the target application; the text part is integrated with the picture video to generate video content.

[0141] It can be understood that in this process, video generation technology is used to dynamically concatenate the picture part according to the interface order of the target application and the interaction logic of each functional component in each interface, generate a dynamic transition effect, and combine it with the text part of the introduction text to form a complete video.

[0142] Optionally, the video processing technology may include video analysis, video splicing and synthesis, video stream processing, virtualization and augmented reality, video voice synchronization, and digital human technology, etc., which are not specifically limited here.

[0143] Optionally, the dynamic transition effect refers to the conversion of each interface picture in the picture part into a dynamic transition effect. Among them, each interface picture corresponds to an interface in the target application, and whenever a new interface is switched to, it can be presented through a dynamic transition. For example: smooth transition, zoom effect, overlay animation, and addition of dynamic elements, etc., which are not specifically limited here.

[0144] Among them, the smooth transition refers to the interface sliding from left to right or from top to bottom, simulating the transition effect of actual operation. The zoom effect refers to shrinking or enlarging from one interface to another, giving a sense of gradual change. The overlay animation refers to the appearance of stacked interface elements on the screen, enhancing the sense of hierarchy through effects such as fading in and out or flipping. The addition of dynamic elements refers to adding dynamic elements (such as the effect of clicking a button, the cursor blinking in an input box, the dragging of a slider, etc.) between operation steps to simulate actual user operations.

[0145] In the embodiments of the present application, the text part is integrated with the picture video, that is, the content of the text needs to be synchronously displayed with the interface and operation steps.

[0146] In some embodiments, to improve the visual effect, dynamic presentation of text can be adopted to match the text part with the video. For example, fade-in and fade-out, typing effect, and dynamic following, etc., are not specifically limited herein.

[0147] Among them, fade-in and fade-out means that the copywriting can slowly emerge from outside the screen or gradually disappear after being displayed. The typing effect means simulating the appearance of text character by character, simulating the feeling of user input. Dynamic following means that the copywriting follows the operation steps and gradually unfolds as the interface changes.

[0148] Step S503: Through digital human technology, generate a voice introduction corresponding to each interface according to the text part.

[0149] Among them, the voice introduction is a voice introduction with a virtual digital human.

[0150] Digital human technology refers to the use of technologies such as artificial intelligence, computer graphics, and virtual reality to create virtual characters that can simulate human behavior, appearance, and interaction. Digital humans can not only be virtual characters that are extremely similar to real people in appearance and behavior, but also perform functions such as intelligent conversation, task processing, and emotion expression, with a certain degree of autonomy and interactivity.

[0151] Exemplarily, generating a voice introduction through digital human technology according to the introduction copywriting may include inputting the text part into a digital human engine and converting the text part of the introduction copywriting into speech. This process can be achieved through Text-to-Speech (TTS) technology. TTS technology will generate a sound similar to natural speech according to the input text. Then, create a virtual character image to match the speech. This usually involves 3D modeling, facial expression, and motion capture technologies. Finally, synchronize the already generated speech with the animation of the digital human so that when the digital human speaks, actions such as lip shape, eye expression, and facial expression are consistent with the speech content. Thus, a voice introduction containing the digital human image is generated.

[0152] Generating a voice introduction through this digital human technology can not only make the user experience more vivid and immersive through natural speech and actions. Moreover, it can reduce labor costs by replacing manual work, and also reduce human errors because digital humans can accurately execute tasks according to predetermined scripts and algorithms.

[0153] Step S504: Generate an introduction video according to the video content and the voice introduction corresponding to each interface.

[0154] In this method for generating an introduction video, by connecting to the target application and adopting a combination of multi-modal large model technology and digital human technology, an introduction video including a digital human image, voice narration, pictures of the target application interface, and corresponding introduction copy is generated. Since there is no manual intervention in the generation of this introduction video, the efficiency of the introduction video production method is improved, realizing a more intelligent introduction video generation method. Moreover, using a multi-modal large model to generate the introduction video, because this multi-modal large model can process and integrate information from various types, it provides a more comprehensive and accurate understanding and decision-making, generating a more exquisite and high-quality introduction video, avoiding the uneven quality of the introduction video caused by the limited level of people during manual operation. In addition, the digital human technology is combined, and through natural voice and actions, the user experience is made more vivid and immersive. Also, because the digital human can accurately execute tasks according to the predetermined script and algorithm, reducing human errors and improving work efficiency.

[0155] Next, a specific example is used to specifically describe the generation process of the introduction video of the target application.

[0156] Step 1, use the visual understanding ability, parsing ability, and dynamic interaction ability of the multi-modal large model to explore different interfaces and functional characteristics in the target application in real time. Automatically generate and record the interface components of the target application and their corresponding interaction logics. Thus, obtain the interface pictures of each interface, the functional characteristics of each interface, and the interaction logics of each functional component of each interface; among them, the functional characteristics of each interface are used to represent the functions implemented by each interface, and the interaction logic of each functional component is used to represent the interaction situation of each functional component with other interfaces.

[0157] Step 2, based on the text generation ability of the multi-modal large model, summarize the characteristic descriptions of each interface and function point according to the interface pictures of each interface, the functional characteristics of each interface, and the interaction logics of each functional component of each interface, and generate descriptive text.

[0158] Step 3, use the language model of the multi-modal large model to generate a corresponding prompt word template according to the descriptive text, and guide the multi-modal large model to generate introduction copy based on the descriptive text. At the same time, match the interface pictures of the target application with the introduction copy to provide a visual contrast for each description content.

[0159] Step 4, dynamically concatenate the picture parts in the interface order of multiple interfaces and the operation logic of the target application to generate a picture video; fuse the text part with the picture video to generate complete video content; through digital human technology, input the introduction copy into the digital human engine, and generate the voice introduction corresponding to each interface according to the text part, where the language introduction includes the language introduction of the virtual digital human;

[0160] Step 5: Generate a complete introduction video of the functions of the target application based on the video content and the voice introductions of virtual digital humans corresponding to each interface.

[0161] In the embodiments of the present application, the method for obtaining the description text of each interface may include:

[0162] Use a multi-modal large model to obtain the interface pictures of each interface;

[0163] Analyze the interface pictures to obtain the various functional components of each interface;

[0164] Perform simulated operations on the various functional components of each interface, where the simulated operations may include actions such as sliding, dragging, and clicking.

[0165] In response to the simulated operations, obtain the functional characteristics of each interface and the interaction logic of the various functional components of each interface.

[0166] Optionally, in the above Steps 1 to 5 for generating the introduction video of the target application, the generation of the introduction video does not necessarily have to be carried out strictly in this step sequence.

[0167] In some embodiments, the method for generating the introduction video may also have the technology of continuously recording the functional characteristics of different interfaces of the target application during the interaction with the target application, and updating the knowledge base in the case of discovering an update of the target application interface. Thus, the method for generating the introduction video can adapt to different version iterations of the target product. Therefore, when the target application product is updated, through simulating user operations to form a closed-loop learning, different from traditional static interface analysis, the relevant introduction video can be updated in a timely manner, greatly improving the intelligent level.

[0168] The embodiments of the present application also disclose a device for generating an introduction video. Further refer to Figure 1 , including a copywriting generation module 11 and a video generation module 21, where:

[0169] The copywriting generation module 11 is used to parse each interface of the target application by using a multi-modal large model to generate an introduction copywriting of the target application. The introduction copywriting includes a text part and a picture part. The text part is used to describe each interface and the various functional components of each interface, and the picture part includes the interface pictures of each interface;

[0170] The video generation module 21 is used to generate an introduction video of the target application according to the introduction copywriting. The introduction video includes the interface pictures of each interface and the corresponding voice introductions.

[0171] Optionally, the copywriting generation module 11 is further configured to parse each interface of the target application using a multimodal large model to obtain a description text, where the description text is used to describe each interface and each functional component of each interface;

[0172] According to the prompt word template corresponding to the description text, guide the multimodal large model to generate an introduction copy based on the description text.

[0173] It can be understood that the multimodal large model may include a language model for analyzing the description text to generate a corresponding prompt word template.

[0174] Optionally, the copywriting generation module 11 is further configured to use the multimodal large model to obtain the interface picture of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface. The functional characteristics of each interface are used to represent the functions implemented by each interface, and the interaction logic of each functional component is used to represent the interaction situation of each functional component with other interfaces;

[0175] Generate an introduction copy according to the interface picture of each interface, the functional characteristics of each interface, and the interaction logic of each functional component of each interface.

[0176] Optionally, the copywriting generation module 11 is further configured to use the multimodal large model to obtain the interface picture of each interface; analyze the interface picture to obtain each functional component of each interface; perform simulation operations on each functional component of each interface; in response to the simulation operations, obtain the functional characteristics of each interface and the interaction logic of each functional component of each interface.

[0177] Optionally, the copywriting generation module 11 is further configured to connect the multimodal large model to the application program interface of the target application.

[0178] Optionally, the video generation module 21 is further configured to generate video content according to the introduction copy;

[0179] Generate a voice introduction corresponding to each interface according to the text part;

[0180] Generate an introduction video according to the video content and the voice introduction corresponding to each interface.

[0181] It can be understood that the video generation module 21 includes various video processing technologies, such as video analysis, video splicing and synthesis, video stream processing, virtualization and augmented reality, video voice synchronization, and digital human technology, etc., which are not specifically limited here.

[0182] Optionally, the video generation module 21 is further configured to generate a picture video according to the picture part in the order of the interfaces of multiple interfaces and the operation logic of the target application;

[0183] Integrate the text part with pictures and videos to generate video content.

[0184] Optionally, the video generation module 21 is further configured to generate a voice introduction corresponding to each interface according to the text part through digital human technology.

[0185] Based on the above-described method for generating an introduction video and the apparatus for generating an introduction video, an embodiment of the present application further discloses an electronic device, which includes the apparatus for generating an introduction video described above.

[0186] Based on the above-described method for generating an introduction video, the apparatus for generating an introduction video, and the electronic device, an embodiment of the present application further discloses a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above-described method for generating an introduction video is implemented.

[0187] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a ROM, etc.

[0188] Any reference to a memory, storage, database, or other medium as used herein may include non-volatile and / or volatile memory. Suitable non-volatile memory may include ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which is used as an external cache. By way of illustration and not limitation, RAM can be in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus direct RAM (Rambus DRAM, RDRAM), and direct memory bus dynamic RAM (DirectRambus DRAM, DRDRAM).

[0189] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. Those skilled in the art should also be aware that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0190] In various embodiments of the present application, it should be understood that the size of the sequence numbers of the above processes does not necessarily mean the inevitable sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0191] The units described as separate components above may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0192] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0193] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.

[0194] It should be noted that in this article, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0195] The methods disclosed in several method embodiments provided by this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0196] The features disclosed in several product embodiments provided by this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0197] The features disclosed in several method or device embodiments provided by this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0198] The above has introduced in detail the method for generating an introduction video disclosed in the embodiments of this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. At the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for generating an introduction video, characterized in that, Including: Using a multimodal large model to parse each interface of multiple interfaces of a target application to generate an introduction text for the target application, the introduction text including a text part and a picture part, the text part being used to describe each interface and each functional component of each interface, and the picture part including interface pictures of each interface; Generating an introduction video for the target application according to the introduction text, the introduction video including interface pictures of each interface and voice introductions corresponding to each interface.

2. The method according to claim 1, wherein The step of using a multimodal large model to parse each interface of multiple interfaces of a target application to generate an introduction text for the target application includes: Using a multimodal large model to parse each interface of multiple interfaces of a target application to obtain a description text, the description text being used to describe each interface and each functional component of each interface; Guiding the multimodal large model to generate the introduction text based on the description text according to a prompt word template corresponding to the description text.

3. The method according to claim 2, characterized in that, The step of using a multimodal large model to parse each interface of multiple interfaces of a target application to obtain a description text, the description text being used to describe each interface and each functional component of each interface, includes: Using the multimodal large model to obtain interface pictures of each interface, functional characteristics of each interface, and interaction logics of each functional component of each interface, the functional characteristics of each interface being used to represent the functions to be implemented by each interface, and the interaction logic of each functional component being used to represent the interaction situation of each functional component with other interfaces; Obtaining the description text according to the interface pictures of each interface, functional characteristics of each interface, and interaction logics of each functional component of each interface.

4. The method according to claim 3, characterized in that, The step of using the multimodal large model to obtain interface pictures of each interface, functional characteristics of each interface, and interaction logics of each functional component of each interface includes: Using the multimodal large model to obtain interface pictures of each interface; Analyzing the interface pictures to obtain each functional component of each interface; Performing simulated operations on each functional component of each interface; In response to the simulated operations, obtaining the functional characteristics of each interface and the interaction logics of each functional component of each interface.

5. The method according to any one of claims 1 to 4, characterized in that The text part includes description contents of each interface, and the description contents of each interface include function introductions of each interface and function introductions of each functional component of each interface, and the description contents of each interface correspond to the interface pictures of each interface.

6. The method according to any one of claims 1 to 4, characterized in that, The text part further includes: highlights of the usage experience of the target application, and / or, an introduction to the operation logic of the target application.

7. The method according to any one of claims 1 to 4, characterized in that, The step of generating an introduction video for the target application according to the introduction text includes: Generating video content according to the introduction text; Generating voice introductions corresponding to each interface according to the text part; Generating the introduction video according to the video content and the voice introductions corresponding to each interface.

8. The method according to claim 7, wherein Generating video content according to the above-introduced copywriting includes: Generating a picture video from the picture part in the interface sequence of the multiple interfaces and the operation logic of the target application; Fusing the text part with the picture video to generate the video content.

9. The method according to claim 7, characterized in that Generating a voice introduction corresponding to each interface according to the text part; including: Generating a voice introduction corresponding to each interface according to the text part through digital human technology, and the voice introduction is a voice introduction with a virtual digital human.

10. A video introduction generating device, characterized in that, Including a copywriting generation module and a video generation module, where: The copywriting generation module is used to parse each interface of multiple interfaces of the target application by using a multimodal large model to generate the introduction copywriting of the target application. The introduction copywriting includes a text part and a picture part. The text part is used to describe each interface and each functional component of each interface. The picture part includes the interface pictures of each interface; The video generation module is used to generate an introduction video of the target application according to the introduction copywriting, and the introduction video includes the interface pictures of each interface and the corresponding voice introductions.