Video script generation method and device, electronic equipment and storage medium
By selecting highly relevant reference video scripts and using tag sequences to generate target video scripts, the problems of low generation quality, style mismatch, and monotonous content are solved, achieving the generation of high-quality, diverse video scripts that conform to market trends.
Patent Information
- Application Number
- CN202411615427.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2044-11-12
AI Technical Summary
Existing technologies for generating video scripts suffer from problems such as low quality, mismatch between style and business scenario, monotonous content, low diversity, and failure to meet market trends.
By using a large model-based generation method, reference video scripts are selected based on demand information and the relevance of video scripts in the database. Target video scripts are generated using tag sequences to ensure that the style is consistent with the business scenario, highly diverse, and in line with market trends.
It improves the quality and click-through rate of generated video scripts, ensures that the style matches the business scenario, avoids monotonous content, and automatically updates reference video scripts to adapt to market trends.
Smart Images

Figure CN119421014B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the fields of large models, generative models, and the like, and more specifically, the present disclosure provides a video script generation method and device, an electronic device, a storage medium, and a computer program product. BACKGROUND
[0002] With the popularity of short videos, the types of content creation are greatly enriched, and video scripts, as the cornerstone of content creation, can provide inspiration for videos themselves and improve the efficiency of creation. SUMMARY
[0003] The present disclosure provides a video script generation method and device, an electronic device, a storage medium, and a computer program product.
[0004] According to an aspect of the present disclosure, a video script generation method is provided, which comprises: determining a reference video script from a database according to demand information and the correlation between each video script in the database; and generating a target video script based on a large model according to the demand information, the reference video script, and a label sequence corresponding to the reference video script; wherein the reference video script comprises a plurality of subtexts, each subtext corresponding to a label; the label sequence comprises a plurality of labels corresponding to the plurality of subtexts in the reference video script, and the order of the plurality of labels in the label sequence is consistent with the order of the plurality of subtexts in the reference video script.
[0005] According to another aspect of the present disclosure, a video script generation device is provided, which comprises a reference video script determination module and a generation module. The reference video script determination module is configured to determine a reference video script from a database according to demand information and the correlation between each video script in the database. The generation module is configured to generate a target video script based on a large model according to the demand information, the reference video script, and a label sequence corresponding to the reference video script. The reference video script comprises a plurality of subtexts, each subtext corresponding to a label; the label sequence comprises a plurality of labels corresponding to the plurality of subtexts in the reference video script, and the order of the plurality of labels in the label sequence is consistent with the order of the plurality of subtexts in the reference video script.
[0006] According to another aspect of the present disclosure, an electronic device is provided, which comprises at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to perform the method provided by the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method provided by the present disclosure.
[0009] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:
[0011] Figure 1 is an application scenario diagram of the video script generation method and device according to the embodiments of the present disclosure;
[0012] Figure 2 is a schematic flow chart of the video script generation method according to the embodiments of the present disclosure;
[0013] Figure 3 is a schematic principle diagram of the video script generation method according to the embodiments of the present disclosure;
[0014] Figure 4 is a schematic structure block diagram of the video script generation device according to the embodiments of the present disclosure; and
[0015] Figure 5 is a structure block diagram of an electronic device for implementing the video script generation method according to the embodiments of the present disclosure. DETAILED DESCRIPTION
[0016] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, the description below omits the description of well-known functions and structures.
[0017] In the technical scheme of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical scheme of the present disclosure comply with the relevant legal regulations and do not violate public order and good customs.
[0018] In the technical scheme of the present disclosure, the authorization or consent of the user is obtained before the user's personal information is acquired or collected.
[0019] Sometimes a user needs to create a video script, and the user can be an advertiser or another user. In some embodiments, a large model can be fine-tuned based on original capabilities or training data, and then the large model can be used to generate a target video script based on input content provided by the user.
[0020] However, the above technical solutions have the following disadvantages: First, if the process of generating the target video script does not refer to high-quality reference video scripts, the quality of the generated target video script is generally low, resulting in a low click rate of the target video script. Second, if the large model lacks reference video scripts during the process of generating the target video script, or excessively refers to the content of the reference video scripts, the generated target video script may not match the user's demand for the business scenario. For example, the business scenario is medical rehabilitation, and the style of the reference video script is relatively cheerful. In this case, the style of the generated target video script is also likely to be cheerful, which is less suitable for the medical rehabilitation business scenario. Third, when generating a batch of target video scripts, if the reference high-quality scripts are relatively single, the generated target video scripts will have single content and low diversity. Fourth, the original capabilities of the base model and the quality of the fine-tuning data determine the effect of generating the target video script, and the process of iterating the model and constructing the fine-tuning data is relatively complex. If iteration is not performed, the generated target video script may not adapt to market trends, and the attractiveness may decrease.
[0021] The embodiments of the present disclosure provide a video script generation method. The method first determines a reference video script from a database according to demand information and the correlation between each video script in the database. The reference video script includes multiple subtexts, and each subtext corresponds to a label. A label sequence includes multiple labels corresponding to the multiple subtexts in the reference video script, and the order of the multiple labels in the label sequence is consistent with the order of the multiple subtexts in the reference video script. Then, the method generates a target video script based on a large model according to the demand information, the reference video script, and the label sequence corresponding to the reference video script.
[0022] First, the video script generation method provided by the embodiments of the present disclosure can automatically select a reference video script from a database based on correlation, and then generate a target video script based on the reference video script. The reference video script in the database can be a high-quality example, so that the high-quality reference video script can be used as an example for imitation, thereby improving the quality of the generated target video script and ensuring that the target video script has a high click rate.
[0023] Secondly, the video script generation method provided by the embodiment of the disclosure imitates according to the tags and the tag order of each subtext in the reference video script, which is equivalent to imitating the structure of the reference video script, and avoids the large model paying too much attention to the content, style, tone, etc. of the reference video script, so that the content, style, tone, etc. of the reference video script has less influence on the target video script, thereby improving the consistency of the style and the business scenario of the target video script.
[0024] Furthermore, the process of generating the target video script mainly depends on the tags of each subtext in the reference video script, which belongs to structural level imitation, and the range in which the large model can freely play is wider, so that it can avoid generating a large number of target video scripts with similar content, and the diversity is better.
[0025] In addition, the embodiment of the disclosure generates the target video script depending on the reference video script, and in the case that the large model meets the requirements, updating the reference video script as needed (for example, filtering out resources liked by users according to the performance of the resources online, and then updating the database by using the video scripts associated with the resources) can make the target video script meet the market trend, without manually screening the corpus to fine-tune the large model.
[0026] The video script generation method provided by the embodiment of the disclosure is suitable for application scenarios such as product promotion. The technical solutions provided by the disclosure will be described in detail below in combination with the drawings and specific embodiments.
[0027] Figure 1 is a schematic diagram of an application scenario of the video script generation method and device according to the embodiment of the disclosure.
[0028] It should be noted that Figure 1 The system architecture shown is only an example of a system architecture to which the embodiments of the disclosure can be applied, to help those skilled in the art understand the technical content of the disclosure, but does not mean that the embodiments of the disclosure cannot be used in other devices, systems, environments or scenarios.
[0029] As Figure 1 shown, the system architecture 100 according to the embodiment can include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used as a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired and / or wireless communication links, etc.
[0030] A user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. The terminal devices 101, 102, 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers and desktop computers, etc.
[0031] The server 105 can be a server providing various services, for example, a background management server providing support for a website browsed by a user using the terminal device 101, 102, 103 (only as an example). The background management server can perform analysis and the like on received user request and the like data, and feed back the processing result (for example, a target video script generated according to demand information input by the user and the like) to the terminal device.
[0032] It should be noted that the video script generation method provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the video script generation apparatus provided by the embodiments of the present disclosure can generally be arranged in the server 105. The video script generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102, 103 and / or the server 105. Correspondingly, the video script generation apparatus provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102, 103 and / or the server 105.
[0033] It should be understood that Figure 1 The number of terminal devices, networks and servers in the system 100 is only illustrative. According to the needs of implementation, there can be any number of terminal devices, networks and servers.
[0034] Figure 2 FIG. 2 is a schematic flowchart of a video script generation method according to an embodiment of the present disclosure.
[0035] As shown in FIG. 2, the video script generation method 200 can include operations S210-S220. Figure 2
[0036] At operation S210, a reference video script is determined from a database according to demand information and correlation between each video script in the database.
[0037] For example, the user can input the demand information through a front-end operation page, or the user can import a to-be-processed page, and then parse the to-be-processed page to obtain the demand information. The to-be-processed page can be a landing page of some entity, and the entity can be a product, a brand, an enterprise, etc. For example, the to-be-processed page is a detail page of a product, a promotion page of an enterprise, etc. For example, the demand information can include a business scenario, a target audience, etc. The business scenario can include medical rehabilitation, dental implantation, etc. The target audience can include the elderly, children, etc.
[0038] For example, the database can be pre-constructed. In an example, a video, a picture-text, or other types of resources that are popular among people can be screened from a large number of resources, and then a script associated with the resource is added to the database.
[0039] For example, the relevance can be determined by at least one of similarity, category, and description dimension, and for example, a video script that meets a predetermined condition can be taken as a reference video script, and the predetermined condition can include at least one of the following: the similarity between the demand information and the video script is greater than or equal to a threshold value, the demand category in the demand information is the same as the actual category of the video script, and the demand description dimension in the demand information is the same as the description dimension of a certain description text in the video script.
[0040] For example, a plurality of video scripts are stored in the database, each video script can include at least one description text, and each description text can include at least one subtext. For example, a certain video script includes 10 sentences, of which 3 sentences describe the product function, and the three sentences are a description text, and each sentence can be a subtext. Each subtext corresponds to at least one label, and the label represents the category of the subtext, for example, the label can include product information, application scenario, target audience, call to action, user pain point, product highlight, price discount, etc. For example, the subtext can be classified based on a classification model to determine the label corresponding to the subtext. Each video script in the database corresponds to a label sequence. For example, the label sequence corresponding to the reference video script includes a plurality of labels corresponding to a plurality of subtexts in the reference video script, and the order of the plurality of labels in the label sequence is consistent with the order of the plurality of subtexts in the reference video script.
[0041] It should be noted that the label can represent a general category rather than a subdivided category, thereby avoiding the case that the granularity of the category represented by the label is too small, causing the target video script to be too similar to the reference video script, and avoiding the case that the style and business scenario of the target video script are inconsistent.
[0042] In operation S220, the target video script is generated based on the large model according to the demand information, the reference video script, and the label sequence corresponding to the reference video script.
[0043] For example, the demand information, the reference video script, and the prompt information (Prompt) are input into the large model, and the target video script is generated by the large model, wherein the prompt information can require the large model to imitate according to the order and expression content of each label in the label sequence of the reference video script, and can prohibit the target video script from containing relevant entity content in the reference video script, for example, the reference video script mentions a certain brand of earphones, and the generated target video script does not contain the brand.
[0044] In actual applications, the scripts in the database can be texts for describing video, image-text, and other types of resources. In this embodiment, a video script is taken as an example. In some embodiments, the generated target video script can have the same label sequence as the reference video script.
[0045] By using the video script generation method provided in the embodiments of the present disclosure, firstly, the reference video script can be automatically selected from the database based on the correlation, and then the target video script is generated based on the reference video script, which can improve the quality of the generated target video script and ensure that the target video script has a high click rate. Secondly, based on the label sequence, the structure of the reference video script can be imitated, rather than the content of the reference video script, so that the consistency of the style and business scenario of the target video script can be improved. Thirdly, the large model does not need to be fine-tuned by human screening of the corpus. In addition, it can also avoid generating a large number of target video scripts with similar content, and the diversity is better.
[0046] Next, the process of constructing the database is described.
[0047] In this embodiment, the resources can be obtained first, which can be video resources for introducing entities. Then, based on the historical interaction data of each of the plurality of resources, the target resource can be determined from the plurality of resources, for example, the historical interaction data includes at least one of the following: click rate, completion rate, number of collections, number of likes, number of comments, and number of shares. Through the historical interaction data, the target resources that are liked by users can be selected from a large number of resources, which can be referred to as hit resources. Then, the video script associated with the target resource can be used to update the video script in the database.
[0048] In this embodiment, the video script associated with the resource with high heat is added to the database, which ensures that the video script in the database meets the market trend and audience preference, and further ensures that the target video script generated based on the reference video script meets the market trend and audience preference. In addition, in actual applications, the video script in the database can be updated at intervals, for example, daily or weekly, so as to further ensure that the video script in the database and the target video script meet the current popular trend.
[0049] Next, the process of determining the reference video script from the database according to the demand information and the correlation between each video script in the database is described.
[0050] According to the demand information and the similarity between each video script in the database, the video script in the database can be screened, and the screened video script is referred to as an initial video script. It should be noted that the video script in the database includes a plurality of description texts, and the plurality of description texts are used to describe the entity associated with the video script from different dimensions. For example, one description text is used to introduce the function of the product, another description text is used to introduce the hardware information of the product, and another description text is used to introduce the customer feedback.
[0051] Next, each description text in the initial video script can be displayed. If the user is generally satisfied with the initial video script, the initial video script can be confirmed, and the initial video script can be used as a reference video script.
[0052] If the user is satisfied with part of the description texts in the initial video script and is not satisfied with another part of the description texts, the user can select the part of the description texts that is not satisfied. At this time, the device executing the video script generation method detects the selection operation for at least part of the description texts in the initial video script, determines other video scripts of the same category as the initial video script from the database, and then determines a first description text from the other video scripts, wherein the description dimension of the first description text is the same as that of the at least part of the description texts. For example, if the user is not satisfied with the description text for introducing the function of the earphone in a certain video script, the user can select the description text. At this time, the device first determines other video scripts of the same category from the database, for example, screens other video scripts for introducing the earphone from the database, and then takes the description text for describing the function of the earphone in the video script as the first description text.
[0053] Next, the remaining description texts in the initial video script except for the at least part of the description texts selected by the user can be taken as second description texts. It can be seen that the at least part of the description texts selected by the user are the description texts that the user is not satisfied with, and the second description texts are the description texts that the user is satisfied with. Then, the reference video script can be determined according to the first description text and the second description text, for example, the first description text and the second description text are combined, and the combined data is taken as the reference video script.
[0054] In this embodiment, after the initial video script is screened from the database, if the user is not satisfied with part of the description texts in the initial video script, the initial video script can be adjusted based on other video scripts in the database. In this way, the initial video script can be modified according to the actual demand, and the matching degree of the reference video script and the user demand can be improved. In addition, during the modification process, the user does not need to think about the specific modification method, but can refer to other video scripts of the same category in the database for adjustment, and the operation is more convenient.
[0055] According to another embodiment of the present disclosure, the process of determining the initial video script from the database according to the similarity between the requirement information and each video script in the database can include: recalling a plurality of video scripts from the database according to the similarity between the requirement information and each video script in the database. Then, the initial video script is determined from the recalled plurality of video scripts according to at least one of the selection operation of the object and the requirement information.
[0056] For example, if the object performs the selection operation on at least one video script in the recalled plurality of video scripts within the first predetermined period, the selected at least one video script can be taken as the initial video script.
[0057] For another example, if the object does not perform the selection operation on the recalled plurality of video scripts within the first predetermined period, and the requirement information includes the requirement category, at least one video script can be selected from the recalled plurality of video scripts as the initial video script according to the requirement category and the actual category of each of the plurality of video scripts. For example, the video script includes a plurality of description texts, but the video script can focus on some description texts, for example, the character number proportion of a certain description text is the highest in the video script, and the video script focuses on embodying the description text. In this way, among the recalled plurality of video scripts, some video scripts focus on introducing product functions, some video scripts focus on introducing product hardware, and some video scripts focus on introducing user feedback. For example, the video scripts with the same actual category as the requirement category can be filtered from the recalled plurality of video scripts first, and if the number of filtered video scripts is less than or equal to a number threshold, which can be 1 or 2, the filtered video scripts can be taken as the initial video script. If the number of filtered video scripts is greater than the number threshold, the filtered video scripts can be sorted based on a heat index, and the initial video script can be selected from the filtered video scripts in order, and the heat index can include the click rate, the complete playback rate, the collection amount, the like amount, the comment amount, and the sharing amount of resources associated with the video script.
[0058] For another example, if the object does not perform the selection operation on the recalled plurality of video scripts within the first predetermined period, and the requirement information does not include the requirement category, at least one video script can be randomly selected from the recalled plurality of video scripts as the initial video script.
[0059] The embodiment first recalls a plurality of video scripts based on the similarity, then determines the initial video script from the recalled plurality of video scripts based on the selection of the object first, and then determines the initial video script from the recalled plurality of video scripts based on the requirement information first if the user does not perform the selection, and then randomly selects the initial video script from the recalled plurality of video scripts if the object does not propose an explicit requirement. In this way, through multi-level filtering, it can be ensured that the initial video script has high relevance with the requirement and meets the user requirement.
[0060] According to another embodiment of the present disclosure, the number of the first description texts is multiple, and the process of determining the reference video script according to the first description text and the second description text comprises: outputting the multiple first description texts, and then in response to receiving a selection operation on at least part of the multiple first description texts, determining the reference video script according to the at least part of the first description texts and the second description text.
[0061] For example, if the user is not satisfied with the description text for introducing the function of the earphone in the initial video script, the user can select the description text. At this time, the device for executing the video script generation method first determines other video scripts of the same category from the database, for example, filters multiple video scripts for introducing the earphone from the database, and then takes the description text for describing the function of the earphone in the video script as the first description text, so that there are multiple first description texts. Next, the multiple first description texts can be displayed, and the user can select according to personal preference. For example, if a certain first description text is selected, the user-selected first description text and the second description text in the initial video script are combined to obtain the reference video script.
[0062] In this embodiment, the multiple first description texts can be selected from the database, the video scripts to which the multiple first description texts belong are of the same category as the initial video script, and the multiple first description texts are of the same description dimension as the part of the description text that the user is not satisfied with in the initial video script. Then the user can flexibly select according to personal preference, and then obtain the reference video script that meets the user's demand, so that the reference video script can be more in line with the user's personal demand, and the user experience can be improved.
[0063] According to another embodiment of the present disclosure, the video script generation method in this embodiment can include the following operations: determining whether the object selects a reference video script from the database within a second predetermined period of time. If yes, the user-selected video script is taken as the initial video script or directly as the reference video script. If no, the reference video script is determined from the database according to the relevance between the demand information and each video script in the database, and then the target video script is generated based on the large model according to the demand information, the reference video script, and the label sequence corresponding to the reference video script.
[0064] In this embodiment, if the user actively selects, the user-selected video script is preferentially taken as the initial video script, and then the reference video script is determined based on the initial video script, or the user-selected video script can be directly taken as the reference video script. If the user does not actively select, the reference video script can also be automatically filtered from the database, so as to provide a reference for generating the target video script.
[0065] Next, the process of determining the demand information is described.
[0066] If the user can input information through the front-end page, the user input information can be taken as the demand information. If the above input information is not input, but a to-be-processed page is uploaded through the front-end page, in this case, the to-be-processed page can be parsed, for example, characters in the to-be-processed page are extracted through text extraction, optical character recognition and the like, and an abstract can also be extracted through an abstract generation model, so as to obtain a plurality of fields, which can include business, highlights, target audience, abstract and the like. Next, these fields can be taken as the demand information.
[0067] In this embodiment, when the user does not input an explicit demand through the front-end page, but uploads a to-be-processed page, the demand information can be determined based on the to-be-processed page, which is more convenient for user operation and improves user experience. In addition, the fields are extracted from the to-be-processed page first, and then the demand information is determined based on the fields, which can ensure the accuracy of the demand information.
[0068] In the process of determining the demand information, for each field in the plurality of fields, it can be determined whether the field is a key field according to the business scenario of the object associated with the to-be-processed page, and then the demand information is determined according to the key field in the plurality of fields. For example, the to-be-processed page is uploaded by a first object, and according to the pre-established information base, it is known that the first object mainly develops earphone products, so the business scenario is earphone. It can be determined whether the field in the to-be-processed page is associated with the earphone, if not, the field can be deleted. If it is associated, the field can be determined as a key field. In this way, in the process of determining the demand information, the demand information can be generated according to the key field which is more closely associated with the business scenario, rather than all fields, so that the demand information can highlight the content related to the business scenario and ignore the irrelevant content.
[0069] It can be understood that in other embodiments, the fields unrelated to the business scenario can also not be filtered, but the demand information is determined according to all fields.
[0070] Figure 3 is a schematic principle diagram of a video script generation method according to an embodiment of the present disclosure.
[0071] In this embodiment, based on the heat indicators such as click rate, complete play rate, collection amount, like amount, comment amount and sharing amount, a target resource 302 which is widely targeted by the audience can be screened from a large number of online resources 301, and the target resource 302 can be a video resource. Then, a video script 303 associated with the target resource 302 is added to a database 304. The video scripts in the database 304 can be updated regularly, and the update period can be 1 day, 3 days, etc., which is not limited in this embodiment.
[0072] Next, the requirement information 305 of the user can be acquired, and two input forms are provided for the user in this embodiment. In the first input form, the user can directly input the requirement information 305 through the front-end operation page. In the second input form, the user can upload a to-be-processed page through the front-end operation page, and the to-be-processed page can be parsed to obtain the requirement information 305.
[0073] Next, the reference video script 307 can be determined from the database 304. For example, initial video scripts can be recalled from the database 304 according to the similarity between the requirement information 305 and each video script in the database 304, and then the initial video scripts are displayed. If the user is not satisfied with part of the description text in the initial video script, the user can select the part of the description text. A plurality of first description texts can be determined from other video scripts in the database 304 that have the same category as the initial video script, and the plurality of first description texts have the same description dimension as the part of the description text selected by the user. Next, the plurality of first description texts are displayed for the user to select, and the reference video script 307 is determined based on the first description text selected by the user and a second description text in the initial description text that is satisfactory to the user.
[0074] It should be noted that the reference video script 307 includes a plurality of subtexts, and a label set 306 can be preconfigured. The set can be a closed set, and the subtexts in the reference video script 307 can be classified to determine the label of each subtext, so as to obtain a label sequence 308 of the reference video script 307. The order of the plurality of labels in the label sequence 308 is consistent with the order of the plurality of subtexts in the reference video script 307.
[0075] After the reference video script 307 is obtained, the requirement information 305, the reference video script 307, and the label sequence 308 of the reference video script 307 can be input into a large model 309. Thus, the target video script 310 with attraction and influence can be generated based on the market trend and the user requirement according to the large model 309.
[0076] Figure 4 is a schematic structural block diagram of a video script generation apparatus according to an embodiment of the present disclosure.
[0077] As shown in Figure 4 , the video script generation apparatus 400 can include a reference video script determination module 410 and a generation module 420.
[0078] The reference video script determination module 410 is configured to determine a reference video script from the database according to the requirement information and the relevance between each video script in the database. The reference video script includes a plurality of subtexts, each of which corresponds to a label; a label sequence includes a plurality of labels corresponding to the plurality of subtexts in the reference video script, and the order of the plurality of labels in the label sequence is consistent with the order of the plurality of subtexts in the reference video script.
[0079] The generation module 420 is configured to generate a target video script based on the large model according to the requirement information, the reference video script, and the label sequence corresponding to the reference video script.
[0080] According to another embodiment of the present disclosure, the reference video script determination module includes an initial video script determination submodule, a first description text determination submodule, and a reference video script determination submodule. The initial video script determination submodule is configured to determine an initial video script from the database according to the requirement information and the similarity between each video script in the database; wherein the video script in the database includes a plurality of description texts, and the plurality of description texts are used to describe an entity associated with the video script from different dimensions. The first description text determination submodule is configured to determine a first description text from other video scripts in the database that have the same category as the initial video script in response to detecting a selection operation for at least part of the description texts in the initial video script, wherein the description dimension of the first description text is the same as that of the at least part of the description texts. The reference video script determination submodule is configured to determine a reference video script according to the first description text and a second description text; wherein the second description text is the remaining description text in the initial video script except for the at least part of the description texts.
[0081] According to another embodiment of the present disclosure, the initial video script determination submodule includes a recall unit and an initial video script determination unit. The recall unit is configured to recall a plurality of video scripts from the database according to the requirement information and the similarity between each video script in the database. The initial video script determination unit is configured to determine an initial video script from the recalled plurality of video scripts according to at least one of the selection operation of the object and the requirement information.
[0082] According to another embodiment of the present disclosure, the initial video script determining unit comprises a first sub-unit, a second sub-unit and a third sub-unit. The first sub-unit is configured to, in response to detecting that the object performs a selection operation on at least one video script in the plurality of video scripts for recall within the first predetermined period of time, determine the selected at least one video script as the initial video script. The second sub-unit is configured to, in response to detecting that the object does not perform a selection operation on the plurality of video scripts for recall within the first predetermined period of time and the demand information comprises a demand category, select at least one video script from the plurality of video scripts for recall as the initial video script according to the demand category and actual categories of the plurality of video scripts. The third sub-unit is configured to, in response to detecting that the object does not perform a selection operation on the plurality of video scripts for recall within the first predetermined period of time and the demand information does not comprise a demand category, randomly select at least one video script from the plurality of video scripts for recall as the initial video script.
[0083] According to another embodiment of the present disclosure, the number of the first description texts is a plurality; the reference video script determining sub-module comprises an output unit and a reference video script determining unit. The output unit is configured to output the plurality of first description texts. The reference video script determining unit is configured to, in response to receiving a selection operation on at least part of the plurality of first description texts, determine a reference video script according to the at least part of the first description texts and the second description text.
[0084] According to another embodiment of the present disclosure, further comprising a parsing module and a demand information determining module. The parsing module is configured to, in response to receiving the to-be-processed page, parse the to-be-processed page to obtain a plurality of fields. The demand information determining module is configured to determine the demand information according to the plurality of fields.
[0085] According to another embodiment of the present disclosure, the demand information determining module comprises a judgment sub-module and a demand information determining sub-module. The judgment sub-module is configured to, for each field in the plurality of fields, determine whether the field is a key field according to a business scenario of the object associated with the to-be-processed page. The demand information determining sub-module is configured to determine the demand information according to the key fields in the plurality of fields.
[0086] According to another embodiment of the present disclosure, further comprising a triggering module configured to, in response to detecting that the object does not select a reference video script from the database within the second predetermined period of time, trigger an operation of determining the reference video script from the database according to the demand information and the relevance between each video script in the database.
[0087] According to another embodiment of the present disclosure, the method further includes: determining a target resource from the plurality of resources based on historical interaction data of each of the plurality of resources; and determining the video script in the database according to a video script associated with the target resource. The historical interaction data includes at least one of: a click rate, a complete play rate, a number of collections, a number of likes, a number of comments, and a number of shares.
[0088] According to another embodiment of the present disclosure, the resource is a video resource for introducing the entity.
[0089] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, including at least one processor; and a memory connected with the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the video script generation method.
[0090] According to an embodiment of the present disclosure, the present disclosure further provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to enable a computer to perform the video script generation method.
[0091] According to an embodiment of the present disclosure, the present disclosure further provides a computer program product including a computer program, where the computer program, when executed by a processor, implements the video script generation method.
[0092] Figure 5 is a structural block diagram of an electronic device for implementing the video script generation method of the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections, and their functions, as well as their relationships with one another, are merely examples and are not intended to limit the implementations described and / or claimed in this document to the aspects depicted.
[0093] As Figure 5As shown, the device 500 includes a computing unit 501 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0094] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disk, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0095] The computing unit 501 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs various methods and processes described above, such as the video script generation method. For example, in some embodiments, the video script generation method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the video script generation method described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the video script generation method by any other appropriate means, such as by means of firmware.
[0096] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0097] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.
[0098] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0099] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0100] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0101] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0102] It should be understood that various forms of flow shown above can be used, re-ordered, added to, or deleted from without departing from the spirit of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, without departing from the desired results of the technical solutions disclosed in the present disclosure, and this is not limited herein.
[0103] The specific embodiments described above are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the present disclosure. Any further modifications, equivalent substitutions, improvements, combinations and the like not described above are also intended to be encompassed within the scope of the present disclosure.
Claims
1. A method for generating a video script, comprising: Determining an initial video script from the database based on the demand information and similarities between various video scripts in the database; wherein the video script in the database includes a plurality of description texts, and the plurality of description texts are used to describe entities associated with the video script from different dimensions; In response to detecting a selection operation on at least a portion of the description text in the initial video script, determining a first description text from other video scripts in the database that are of the same category as the initial video script, wherein a description dimension of the first description text is the same as a description dimension of the at least portion of the description text; Determine a reference video script according to the first description text and the second description text; wherein the second description text is the remaining description text in the initial video script except for the at least part of the description text; Inputting the requirement information, the reference video script and prompt information into a large model, and generating a target video script by the large model, wherein the prompt information is used to prompt the large model to imitate the order and expression content of each tag in the tag sequence corresponding to the reference video script; The reference video script includes multiple sub-texts, each sub-text corresponds to a label; the label sequence includes multiple labels corresponding to the multiple sub-texts in the reference video script, and the order of the multiple labels in the label sequence is consistent with the order of the multiple sub-texts in the reference video script.
2. The method according to claim 1, wherein Determining the initial video script from the database according to the demand information and the similarity between each video script in the database includes: Recalling a plurality of video scripts from the database according to similarities between the demand information and each video script in the database; and The initial video script is determined from a plurality of recalled video scripts according to at least one of the selection operation of the object and the requirement information.
3. The method according to claim 2, wherein: The determining the initial video script from the recalled multiple video scripts according to at least one of the object selection operation and the demand information comprises: In response to detecting that the subject performs a selection operation on at least one video script among the recalled plurality of video scripts within a first predetermined time period, using the selected at least one video script as the initial video script; In response to detecting that the subject does not perform a selection operation on the recalled multiple video scripts within the first predetermined period, and the demand information includes a demand category, selecting at least one video script from the recalled multiple video scripts as the initial video script based on the demand category and the actual categories of the multiple video scripts; and In response to detecting that the object does not perform a selection operation on the recalled multiple video scripts within the first predetermined time period, and the demand information does not include the demand category, at least one video script is randomly selected from the recalled multiple video scripts as the initial video script.
4. The method according to claim 1, wherein The number of the first description texts is multiple; and determining the reference video script according to the first description texts and the second description texts includes: outputting a plurality of first description texts; and In response to receiving a selection operation on at least a portion of the first description texts among the plurality of first description texts, the reference video script is determined based on the at least a portion of the first description texts and the second description text.
5. The method according to claim 1, further comprising: In response to receiving a page to be processed, parsing the page to be processed to obtain a plurality of fields; as well as The requirement information is determined according to the multiple fields.
6. The method according to claim 5, wherein: The determining the requirement information according to the multiple fields includes: For each of the plurality of fields, determining whether the field is a key field according to a business scenario of an object associated with the page to be processed; and The requirement information is determined according to the key field in the multiple fields.
7. The method according to claim 1, further comprising: In response to detecting that the subject does not select a reference video script from the database within a second predetermined period, an operation of determining the reference video script from the database according to correlations between the demand information and each video script in the database is triggered.
8. The method according to claim 1, further comprising: Determine a target resource from the multiple resources based on their respective historical interaction data; as well as Determining the video script in the database according to the video script associated with the target resource; The historical interaction data includes at least one of the following: click-through rate, completion rate, number of collections, number of likes, number of comments, and number of shares.
9. The method according to claim 8, wherein The resource is a video resource used to introduce the entity.
10. A video script generating device, comprising: An initial video script determination submodule is configured to determine an initial video script from the database based on the demand information and the similarity between each video script in the database; wherein the video script in the database includes a plurality of description texts, and the plurality of description texts are used to describe entities associated with the video script from different dimensions; A first description text determination submodule is configured to, in response to detecting a selection operation on at least a portion of the description text in the initial video script, determine a first description text from other video scripts in the database that are of the same category as the initial video script, wherein a description dimension of the first description text is the same as a description dimension of the at least portion of the description text; A reference video script determination submodule, configured to determine a reference video script based on the first description text and the second description text; wherein the second description text is the remaining description text in the initial video script except for the at least part of the description text; A generation module, configured to input the requirement information, the reference video script, and prompt information into a large model, and generate a target video script by the large model, wherein the prompt information is used to prompt the large model to imitate the order and expression content of each tag in the tag sequence corresponding to the reference video script; The reference video script includes multiple sub-texts, each sub-text corresponds to a label; the label sequence includes multiple labels corresponding to the multiple sub-texts in the reference video script, and the order of the multiple labels in the label sequence is consistent with the order of the multiple sub-texts in the reference video script.
11. The device according to claim 10, wherein The initial video script determination submodule includes: a recall unit, configured to recall a plurality of video scripts from the database according to similarities between the demand information and the respective video scripts in the database; and An initial video script determining unit is configured to determine the initial video script from a plurality of recalled video scripts based on an object selection operation and at least one of the demand information.
12. The device according to claim 11, wherein The initial video script determining unit includes: A first subunit is configured to, in response to detecting that the subject performs a selection operation on at least one video script from the recalled plurality of video scripts within a first predetermined time period, use the selected at least one video script as the initial video script; a second subunit, configured to, in response to detecting that the subject has not performed a selection operation on the multiple recalled video scripts within the first predetermined period, and the demand information includes a demand category, select at least one video script from the multiple recalled video scripts as the initial video script according to the demand category and the actual categories of the multiple video scripts; and The third subunit is used to randomly select at least one video script from the recalled multiple video scripts as the initial video script in response to detecting that the object has not performed a selection operation on the recalled multiple video scripts within the first predetermined time period and the demand information does not include the demand category.
13. The device according to claim 10, wherein The number of the first description texts is multiple; The reference video script determination submodule includes: an output unit, configured to output a plurality of first description texts; as well as The reference video script determining unit is configured to determine the reference video script according to the at least part of the first description text and the second description text in response to receiving a selection operation for at least part of the first description text among the multiple first description texts.
14. The apparatus according to claim 10, further comprising: a parsing module, configured to parse the page to be processed in response to receiving the page to be processed, and obtain a plurality of fields; as well as The requirement information determining module is configured to determine the requirement information according to the multiple fields.
15. The device according to claim 14, wherein The demand information determination module includes: a determination submodule, configured to determine, for each of the plurality of fields, whether the field is a key field according to a business scenario of an object associated with the page to be processed; and The requirement information determination submodule is configured to determine the requirement information according to the key field in the multiple fields.
16. The apparatus according to claim 10, further comprising: The trigger module is used to trigger the operation of determining the reference video script from the database according to the correlation between the demand information and each video script in the database in response to detecting that the object has not selected a reference video script from the database within a second predetermined time period.
17. The apparatus according to claim 10, further comprising: a target resource determination module, configured to determine a target resource from a plurality of resources based on respective historical interaction data of the plurality of resources; as well as A video script determination module, configured to determine a video script in the database based on the video script associated with the target resource; The historical interaction data includes at least one of the following: click-through rate, completion rate, number of collections, number of likes, number of comments, and number of shares.
18. The device according to claim 17, wherein The resource is a video resource used to introduce the entity.
19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.
21. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method and apparatus for managing video content
CN102959542A
Script generation method and system, computer storage medium and computer program product
CN113641859A