Video generation method and related equipment

By obtaining and matching the text and visual resources of the target object, the semantic consistent video material is generated, and the problems of inefficient, high-cost and low-correlation in the prior art video generation are solved, and high-quality and low-cost video generation is achieved.

CN119967260APending Publication Date: 2025-05-09BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510138260.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

Existing automated video generation tools rely on preset templates, making it difficult to effectively highlight the personalized characteristics of the object, the dubbing content and video materials are low correlation, and it is impossible to fully demonstrate the characteristics of the object description, and it is difficult to generate videos that meet the characteristics of the target object at low cost and batch.

Method used

By obtaining the text resources and visual resources of the target object, generate semantically matching target text and target visual materials, and combining these resources to generate target videos, improve the quality and accuracy of video generation and enhance the correlation between video content and describing object information.

Benefits of technology

It realizes efficient and low-cost generation of videos that meet the characteristics of the target object, significantly improves the quality and accuracy of video generation, reduces manpower and time investment, and enhances the correlation between video content and object description information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967260A_ABST
    Figure CN119967260A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and related equipment. The method comprises the following steps: acquiring a text resource and a visual resource about a target object; generating a target text and a target visual material which are semantically matched based on the text resource and the visual resource; and generating a target video based on the target text and the target visual material.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video processing, and in particular to a video generation method and related equipment. Background Art

[0002] At present, automated video generation tools often rely on preset templates to generate videos. They lack a deep understanding of the objects described in the videos and find it difficult to effectively highlight the personalized characteristics of the objects being described. The dubbing content is also often less relevant to the video material and cannot fully demonstrate the characteristics of the objects being described. Summary of the invention

[0003] The present disclosure proposes a video generation method and related equipment, which at least to a certain extent solve the technical problems of high cost, low efficiency, low matching degree between video content and description object information in video generation scenarios in related technologies.

[0004] In a first aspect, the present disclosure provides a video generation method, comprising:

[0005] Obtain textual and visual resources about the target object;

[0006] Generate semantically matching target text and target visual material based on the text resource and the visual resource;

[0007] A target video is generated based on the target text and the target visual material.

[0008] In a second aspect of the present disclosure, a video generation device is provided, comprising:

[0009] An acquisition module, used to acquire text resources and visual resources about the target object;

[0010] A matching module, used for generating a semantically matching target text and target visual material based on the text resource and the visual resource;

[0011] A generation module is used to generate a target video based on the target text and the target visual material.

[0012] According to a third aspect of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in the first aspect is implemented.

[0013] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the first aspect.

[0014] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method according to the first aspect.

[0015] As can be seen from the above, the video generation method and related equipment provided by the present disclosure combine the multimodal resources of the target object to obtain semantically matching target text and visual materials, and then synthesize the target text and visual materials to form a target video, thereby improving the quality and accuracy of video generation and enhancing the relevance between video content and related descriptions. It is possible to efficiently and low-cost generate a target video that meets the characteristics of the target object, greatly reducing the investment in manpower and time. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0017] Figure 1 A schematic diagram of a video generation architecture according to an embodiment of the present disclosure.

[0018] Figure 2 The figure is a schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.

[0019] Figure 3 The figure is a flowchart of a video generation method according to an embodiment of the present disclosure.

[0020] Figure 4 A schematic diagram of the principles of video generation according to an embodiment of the present disclosure.

[0021] Figure 5 A schematic diagram of preprocessing of visual resources according to an embodiment of the present disclosure.

[0022] Figure 6 The figure is a schematic diagram of the principle of the multimodal model of the embodiment of the present disclosure.

[0023] Figure 7 A schematic diagram of video rendering according to an embodiment of the present disclosure.

[0024] Figure 8 Schematic diagram of a video generating device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0025] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in combination with specific embodiments and with reference to the accompanying drawings.

[0026] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Including" or "comprising" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0027] It is understandable that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, and usage scenarios of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations. For example, in response to receiving an active request from the user, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application, server, or storage medium that performs the operation of the technical solution of the present disclosure according to the prompt message.

[0028] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that meet the relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0029] Figure 1 FIG. 1 is a schematic diagram showing a video generation architecture of an embodiment of the present disclosure. Figure 1 The video generation architecture 100 may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 may be connected via a wired or wireless network 130. The server 110 may be an independent physical server, or a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and CDN.

[0030] The terminal 120 may be implemented in hardware or software. For example, when the terminal 120 is implemented in hardware, it may be various electronic devices having a display screen and supporting page display, including but not limited to smart phones, tablet computers, e-book readers, laptop portable computers, and desktop computers, etc. When the terminal 120 device is implemented in software, it may be installed in the electronic devices listed above; it may be implemented as multiple software or software modules (such as software or software modules used to provide distributed services), or it may be implemented as a single software or software module, which is not specifically limited here.

[0031] It should be noted that the video generation method provided in the embodiment of the present application can be executed by the terminal 120 or by the server 110. It should be understood that Figure 1 The number of terminals, networks and servers in the embodiment is only for illustration and is not intended to limit the number of terminals, networks and servers. Any number of terminals, networks and servers may be provided as required.

[0032] Figure 2 FIG. 2 shows a schematic diagram of the hardware structure of an exemplary electronic device 200 provided in an embodiment of the present disclosure. Figure 2 As shown, the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208 and a bus 210. The processor 202, the memory 204, the network module 206 and the peripheral interface 208 are connected to each other in communication within the electronic device 200 through the bus 210.

[0033] Processor 202 may be a central processing unit (CPU), a neural network processor (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits. Processor 202 may be used to perform functions related to the technology described in this disclosure. In some embodiments, processor 202 may also include multiple processors integrated into a single logical component. For example, Figure 2 As shown, processor 202 may include a plurality of processors 202a, 202b, and 202c.

[0034] The memory 204 may be configured to store data (eg, instructions, computer code, etc.). Figure 2As shown, the data stored in the memory 204 may include program instructions (e.g., program instructions for implementing the video generation method of the embodiment of the present disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). The processor 202 may also access the program instructions and data stored in the memory 204, and execute the program instructions to operate on the data to be processed. The memory 204 may include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 204 may include a random access memory (RAM), a read-only memory (ROM), an optical disk, a magnetic disk, a hard disk, a solid-state drive (SSD), a flash memory, a memory stick, etc.

[0035] The network module 206 can be configured to provide communication with other external devices to the electronic device 200 via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC) etc.), a cellular network, the Internet or a combination thereof. It is understood that the type of network is not limited to the above specific examples. In some embodiments, the network module 206 can include any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc., in any combination.

[0036] The peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to achieve information input and output. For example, the peripheral devices can include input devices such as a keyboard, a mouse, a touch pad, a touch screen, a microphone, and various sensors, and output devices such as a display, a speaker, a vibrator, and an indicator light.

[0037] The bus 210 can be configured to transmit information between various components of the electronic device 200 (e.g., the processor 202, the memory 204, the network module 206, and the peripheral interface 208), such as an internal bus (e.g., a processor-memory bus), an external bus (USB port, PCI-E bus), etc.

[0038] It should be noted that, although the architecture of the electronic device 200 only shows the processor 202, the memory 204, the network module 206, the peripheral interface 208 and the bus 210, in the specific implementation process, the architecture of the electronic device 200 may also include other components necessary for normal execution. In addition, it can be understood by those skilled in the art that the architecture of the electronic device 200 may also only include the components necessary for implementing the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.

[0039] The production of high-quality videos usually requires a lot of manpower and time. Although some automated video generation tools can generate videos based on templates, these tools often rely on preset templates and lack a deep understanding of the objects described in the video. The generated videos are difficult to effectively highlight the personalized characteristics of the described objects. In addition, the dubbing content is often less relevant to the video material and cannot fully demonstrate the core characteristics of the product. For large-scale description objects, it is difficult to generate corresponding videos in batches at low cost, especially to generate customized video content in a targeted manner according to the different characteristics of different description objects. Therefore, how to improve the correlation between video content and related descriptions, improve the quality and efficiency of video generation, and reduce costs, realize personalized and batch generation of videos have become technical problems that need to be solved urgently.

[0040] In view of this, the embodiments of the present disclosure provide a video generation method and related equipment. By combining the multimodal resources of the target object, semantically matching target text and visual materials are obtained, and then the target text and visual materials are synthesized to form a target video, thereby improving the quality and accuracy of video generation and enhancing the relevance between video content and related descriptions. The target video that meets the characteristics of the target object can be generated efficiently and at low cost, greatly reducing the investment of manpower and time.

[0041] See also Figure 3 , Figure 3 The schematic flow chart of the video generation method according to the embodiment of the present disclosure is shown. The video generation method according to the embodiment of the present disclosure can be deployed on a server or a terminal. Figure 3 In the video generation method 300, the video generation method 300 may further include the following steps.

[0042] In step S310, text resources and visual resources about the target object are obtained.

[0043] Among them, the target object can refer to the entity or thing to be displayed. The text resource can be a text description of the target object, including a title, attribute description, specification parameters, etc. The text resource can be text information. The text information of the target object can be obtained directly from the text description of the target object, or it can be obtained by converting the voice information about the target object into text information (for example, converted from the audio or displayed characters in the visual resource). The visual resources of the target object can refer to image resources and video resources, which can intuitively display the appearance, function and usage scenarios of the target object and enhance the understanding and interest of the target object. The text resources and visual resources of the target object can be automatically captured from a database, website or other data source through an interface or API. The text resources and visual resources of the target object can also be uploaded by the user to a specified platform or system. It can be seen that the text resources and visual resources are descriptions and displays of different modes of the target object, and the two have a high correlation and certain complementarity in information expression. The acquisition of text resources and visual resources is conducive to improving the accuracy of subsequent target video generation and provides important basic data.

[0044] Specifically, the target object can be a product, such as a new smartphone. The text resources of the product can include various attributes and features of the product, such as brand, model, color, size, weight, processor model, memory size, battery capacity, camera configuration, operating system, etc. The text resources of the product can also include the selling points, functional features, usage scenarios, etc. of the product. For example, "This smartphone is equipped with the latest A processor, with powerful performance, and can easily handle various large-scale games and complex applications. At the same time, it is also equipped with an aaa megapixel ultra-clear main camera, which can take clearer and more delicate photos and videos. Whether it is daily photography or professional photography, it can meet your needs."

[0045] The visual resources of the product may include product image resources, which may include the product's physical pictures, detail pictures, and usage scene pictures. Physical pictures can show the overall appearance and color of the product; detail pictures can highlight a certain feature or function of the product, such as the layout of the camera, the display effect of the screen, etc.; usage scene pictures can show the effect of the product in actual use. The visual resources of the product may also include product video resources, such as product demonstration videos and function introduction videos. Demonstration videos can show the operation process and use effect of the product, such as boot speed, application opening speed, game running effect, etc.; function introduction videos can introduce the various functions and features of the product in detail, such as photo effects, video functions, intelligent recognition, etc.

[0046] In step S320, semantically matching target text and target visual material are generated based on the text resource and the visual resource.

[0047] Among them, semantic matching can refer to the consistency and coordination between the target text and the target visual material in terms of content, style, etc. The target text can refer to the text description or narration generated based on text resources and used in the target video, which needs to match the target visual material semantically. The target visual material can refer to the images and video clips used in the target video, which needs to match the target text semantically. Specifically, the target text can accurately describe the content displayed by the target visual material, while the target visual material should intuitively reflect the information and style conveyed by the target text, thereby enhancing the user's understanding of the target object through multimodal forms. For example, the target text "This mobile phone is equipped with a powerful A processor, which can easily handle various large-scale games and complex applications", and the target visual material can be "the smooth picture of the mobile phone playing games", then the target visual material and the "can easily handle various large-scale games" mentioned in the above target text are semantically matched.

[0048] Specifically, see Figure 4 , Figure 4 A schematic diagram showing the principles of video generation according to an embodiment of the present disclosure is shown. Figure 4 In the method, visual resources (such as image resources and / or video resources) can be preprocessed to obtain visual materials, and the visual materials and text resources can be input into the multimodal model to obtain the target text and the target visual materials that match the text semantically. The target text and the target visual materials with relevance and consistency are video rendered to generate the target video. It can be seen that compared with the prior art, the video generation method of the embodiment of the present disclosure improves the matching degree and expressiveness of the video content by preprocessing the visual resources, generating the target text and the target visual materials using the multimodal model, and performing video rendering to generate the target video. The multimodal model can further understand the semantic relationship between the text and the visual materials, thereby generating the target visual materials that are highly semantically consistent with the target text. The relevance and consistency between the target text and the target visual materials make the video content more fluent and natural in expression, and can bring better visual and auditory experience to the audience. Through video rendering technology, text and visual materials can be organically combined to form video content with storyline, emotional expression or information transmission. The application of preprocessing and multimodal models can automatically generate target text and target visual materials, greatly reducing the workload of manual editing and screening and improving the efficiency of video production.

[0049] In some embodiments, generating a target text and a target visual material with semantic matching based on the text resource and the visual resource includes:

[0050] generating an initial text about the target object based on the text resource, and obtaining visual material based on the visual resource;

[0051] generating a description text of the visual material based on the text resource and the visual content of the visual material;

[0052] Based on the text resource, the initial text and the description text, the initial text is updated to obtain the target text, and the target visual material matching the target text is determined from the visual material.

[0053] Among them, the initial text may refer to a preliminary text description of the target object directly generated based on the text resource. The initial text may include basic information, functional features, usage scenarios, etc. of the target object. Visual materials may refer to images or video clips selected or edited from visual resources, which are used to intuitively display the appearance, functions, usage scenarios, etc. of the target object. Descriptive text may refer to descriptive text generated based on the visual content of text resources and visual materials. The descriptive text can closely fit the content of the visual material and is used to explain, supplement or emphasize the information displayed by the visual material, so that it is more consistent and coordinated with the description in the text resource. By combining text resources and visual materials to generate descriptive text, and updating the initial text based on the descriptive text and the text resources, a high degree of matching between the target text and the target visual material in terms of content, meaning and emotion is ensured. It helps to improve the coherence and credibility of the target video, making it easier for the audience to understand and accept the information conveyed by the video. The close integration of the target text and the target visual material enables the text description and visual elements in the video to complement and reinforce each other, which helps to improve the overall quality and attractiveness of the video. At the same time, it also reduces the workload of manual editing and review, and improves the efficiency and speed of video generation. At the same time, it also makes video production more flexible and customizable, and can be adjusted and optimized according to different needs and goals.

[0054] In some embodiments, obtaining visual material based on the visual resource includes:

[0055] Preprocessing the video resources in the visual resources to obtain the video materials in the visual materials;

[0056] The preprocessing includes at least one of the following: extracting video slices with different semantics, removing duplicate video slices, identifying key frames in the visual material, or identifying key segments in the visual material.

[0057] Preprocessing refers to data processing of visual resources to obtain high-quality, representative video materials. Through the preprocessing step, redundant and repeated content in the video can be removed, and the most representative and unique video slices, key frames and clips can be retained, thereby improving the overall quality of the video material.

[0058] See also Figure 5 , Figure 5A schematic diagram showing preprocessing of visual resources according to an embodiment of the present disclosure is shown. Figure 5 In the process of extracting video slices with different semantics, the video resource can be divided into multiple video slices with different semantics. Specifically, by identifying different scenes, actions or dialogues in the video, it can be divided into multiple independent video slices with clear semantics, which helps to more accurately locate and use the video content in subsequent analysis, thereby improving the overall quality of the video material. Removing duplicate video slices can refer to the fact that there may be duplicate or similar video content in the video resource. By removing these duplicate video slices, it can be ensured that the material for subsequent analysis is unique and representative, and the interference of redundant information is avoided. Keyframes are representative frames in the video content, such as turning points, climaxes or important information in the video. Similar to keyframes, key segments are continuous video sequences with complete meaning and appeal in the video content, such as highlight segments. Identifying keyframes or key segments in visual materials can more effectively extract key information in the video, providing strong support for subsequent video generation or analysis. Optical character recognition (OCR) and automatic speech recognition (ASR) can also be performed on video resources. Optical character recognition is used to extract text information from videos, while automatic speech recognition is used to extract speech content from videos. They can work together on video resources to convert text and voice information in the video into analyzable text data, further enriching the understanding and application of video content.

[0059] In some embodiments, obtaining visual material based on the visual resource includes:

[0060] Screening the image resources in the visual resources to obtain the image materials in the visual materials;

[0061] Among them, the screening includes at least one of the following: evaluating the image resources based on preset evaluation dimensions to obtain evaluation scores; and determining the image resources whose evaluation scores are greater than or equal to the preset scores as the image materials; or, removing image resources whose similarity is lower than the preset similarity.

[0062] The preset evaluation dimensions may include image clarity, color saturation, composition rationality, theme clarity, and other aspects. Figure 5As shown, each image resource can be scored on these dimensions through an image evaluation model or other algorithms to obtain a comprehensive evaluation score. A preset score can be set as a screening threshold, and image resources with an evaluation score greater than or equal to the preset score are determined as image materials, thereby screening out the most visually attractive pictures. The image evaluation model can be based on deep learning or other machine learning algorithms, which can automatically identify and evaluate aesthetic elements in pictures, such as composition, color, light and shadow, etc. In order to ensure the uniqueness and representativeness of image materials, similar or repeated image resources can be removed. For example, the similarity between different image resources can be calculated through an image similarity algorithm. A preset similarity is set, and image resources with a similarity lower than the preset similarity are removed to avoid duplication and redundancy. After the above screening and preprocessing steps, high-quality video clips and image materials can be generated. These visual materials will be used for subsequent target text matching and video generation, which is conducive to improving the quality of the target video.

[0063] In some embodiments, based on the text resource, the initial text and the description text, updating the initial text to obtain the target text includes:

[0064] Determining, from the visual material, an intermediate visual material matching the initial text based on the text resource, the initial text and the description text;

[0065] The initial text is updated based on the description text corresponding to the intermediate visual material to obtain the target text.

[0066] Among them, the visual material is matched with the initial text according to the text resources (such as a series of keywords, phrases or sentences related to the topic), the initial text (such as a preliminary text description or script) and the description text (such as a brief description or label of the content of the visual material). The text and the description text can be semantically analyzed to understand the relevance and similarity between them, and then find the intermediate visual material that is most relevant or most matching to the content of the initial text from the visual material library, such as the visual material with the highest similarity to the initial text. For example, by comparing the keywords, themes, emotional colors and other elements of the initial text and the description text, the visual material that matches the content of the initial text can be screened out as a candidate. Once the intermediate visual material is determined, the initial text can be updated based on the description text corresponding to the intermediate visual material. For example, the update can include adjusting the expression of the initial text to make it more suitable for the content of the intermediate visual material. For example, some key information or details in the description text can be integrated into the initial text to enhance the richness and accuracy of the text. The updated text can be used as the target text for subsequent voice narration production, video generation or other application scenarios.

[0067] It should be noted that in actual applications, the process of updating the initial text can also go through multiple iterations and optimizations. This is because the matching between text and visual materials is not a simple one-to-one correspondence, but requires comprehensive consideration of multiple factors (such as semantic relevance, emotional color, visual style, etc.) for refined adjustments.

[0068] In some embodiments, determining the target visual material matching the target text from the visual material includes:

[0069] The visual material is voted on multiple times based on a voting mechanism to determine the visual material with the highest relevance to the target text as the target visual material.

[0070] Among them, the visual material combination most relevant to the target text can be screened out through a voting mechanism to obtain the final target visual material. For example, the material combination can be scored or voted according to preset standards (such as relevance, attractiveness, etc.). According to the voting results, the material combination with a low score (for example, lower than a preset value, and the preset value of each round of voting can be different) can be eliminated, or the combination with a higher score (for example, greater than a preset value) can be retained to enter the next round. After multiple rounds of cyclic voting, the material combination with the highest final score or the most votes will be determined as the target visual material that best matches the target text. This can ensure a high degree of consistency between the selected material and the target text in terms of content, theme or emotion. By performing multiple rounds of cyclic voting on the material combination of visual materials based on the voting mechanism, the target visual material that best matches the target text can be effectively determined, thereby improving the accuracy of the match.

[0071] Specifically, see Figure 6 , Figure 6 A schematic diagram of the principle of a multimodal model according to an embodiment of the present disclosure is shown. Figure 6 In the example of the target object being a product, the multimodal model can analyze the text resources of the product to generate an initial text (such as an initial script), which can include the main selling points, functional features and usage scenarios of the product, and can use concise and clear voice to ensure that the product information is transmitted to the maximum extent in a short time. For example, the text resource "product name: XXX; product selling points: XXX..." and other information can be input into the multimodal model, and the initial text first sentence, initial text second sentence, and so on can be output.

[0072] The multimodal model generates a description text related to the product features for each visual material based on the preprocessed visual material, which helps to understand the relevance of each visual material to the product. For example, the text resource "Product name: XXX; Product selling point: XXX...", the visual material and the corresponding visual material identifier (such as the material ID) can be input into the multimodal model, and the description text of the visual material can be output, such as the visual material identifier and the corresponding description text (for example, a description script).

[0073] The multimodal model further optimizes the initial text and matches the relevant visual material based on the descriptive text of the visual material. Specifically, the text resource, the initial text, and the descriptive text can be input into the multimodal model to match the visual material and update the initial text. For example, the text resource "Product name: XXX; Product selling point: XXX..." and other information, the first sentence of the initial text, the second sentence of the initial text, ..., as well as the visual material identifier and the corresponding descriptive text can be input into the multimodal model to output [Visual material identifier 1] the first sentence of the target text, [Visual material identifier 2] the second sentence of the target text, ....

[0074] The multimodal model also continuously adjusts the association between the target text and the visual material through a multi-round voting mechanism to select the optimal material combination and the corresponding target text. The target visual material finally determined will be used for video generation together with the target text. Specifically, the text resource, the target text, the visual material identifier and the corresponding description text are input into the multimodal model to determine whether the currently matched intermediate visual material perfectly matches the corresponding target text (e.g., the target script). If so, no modification is made, otherwise, the most relevant visual material is changed. For example, the text resource "Product name: XXX; Product selling point: XXX..." and other information, [Visual material identifier 1] target text sentence 1, [Visual material identifier 2] target text sentence 2, ..., as well as the visual material identifier and the corresponding description text can be input into the multimodal model, and the output is [Target visual material identifier 1] target text sentence 1, [Target visual material identifier 2] target text sentence 2, ....

[0075] In step S330, a target video is generated based on the target text and the target visual material.

[0076] Among them, a target video with high content relevance and coherence can be generated based on the target text and target visual material.

[0077] In some embodiments, generating a target video based on the target text and the target visual material includes:

[0078] generating a target audio based on the target text;

[0079] Aligning the target audio and the target visual material based on a timeline;

[0080] The target video is generated based on the aligned target audio and target visual material.

[0081] Specifically, see Figure 7 , Figure 7 A schematic diagram of video rendering according to an embodiment of the present disclosure is shown. Figure 7 In the process, the target text can be preprocessed, such as including word segmentation, part-of-speech tagging, and removal of stop words, to ensure that the text content is suitable for subsequent audio generation. Text-to-speech (TTS) technology can be used to convert the target text into a target audio. For example, a pre-trained TTS model can be used, which can generate a corresponding speech waveform based on the text content. During the conversion process, parameters such as speech rate, intonation, and volume can be adjusted to meet the emotional and style requirements of the target visual material. After the target audio is generated, the timeline of the target audio and the target visual material in the target video can be aligned. Specifically, the start time and end time of each target visual material, as well as the corresponding relationship between the target visual material and the target audio, can be determined. Video editing software or programming tools can be used to accurately align the target audio and the target visual material on the timeline. The playback speed of the target visual material can also be adjusted, and the clips can be cropped or extended to ensure that the target visual material matches the rhythm of the target audio. After the target audio and the target visual material are aligned, the two are synthesized into a video file, such as superimposing the target audio track on the target video track, while ensuring that the two are synchronized in time.

[0082] In some embodiments, method 300 further includes: eliminating existing audio data in the target visual material.

[0083] If the matched target visual material contains original audio data, it can be muted to avoid conflicts with the generated target audio. After video rendering, the final generated target video can be exported to multiple formats for display on different multimedia platforms. You can generate target videos for multiple products in batches, or fine-tune a single video to meet different needs.

[0084] It can be seen that according to the method of the embodiment of the present disclosure, a target video that meets the characteristics of the target object can be generated efficiently and at low cost, reducing the investment in manpower and time, and significantly improving the quality and accuracy of automated content generation.

[0085] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will generate videos with each other to complete the described method.

[0086] It should be noted that some embodiments of the present disclosure are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0087] Based on the same technical concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a video generation device, see Figure 8 , the video generating device, the device comprising:

[0088] An acquisition module, used to acquire text resources and visual resources about the target object;

[0089] A matching module, used for generating a semantically matching target text and target visual material based on the text resource and the visual resource;

[0090] A generation module is used to generate a target video based on the target text and the target visual material.

[0091] For the convenience of description, the above device is described by dividing it into various modules according to its functions. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0092] The device of the above embodiment is used to implement the corresponding video generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0093] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the video generation method described in any of the above embodiments.

[0094] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0095] The computer instructions stored in the storage medium of the above embodiments are used to enable the computer to execute the video generation method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0096] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0097] In addition, to simplify the description and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the known power / ground connections to the integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, the device can be shown in the form of a block diagram to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure will be implemented (that is, these details should be fully within the scope of understanding of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it is apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with changes in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0098] Although the present disclosure has been described in conjunction with specific embodiments of the present disclosure, many replacements, modifications and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0099] The embodiments of the present disclosure are intended to cover all such substitutions, modifications and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A video generation method, comprising: Obtain textual and visual resources about the target object; Obtaining semantically matching target text and target visual material based on the text resource and the visual resource; A target video is generated based on the target text and the target visual material.

2. The method according to claim 1, wherein: Obtaining a target text and a target visual material with semantic matching based on the text resource and the visual resource includes: generating an initial text about the target object based on the text resource, and obtaining visual material based on the visual resource; generating a description text of the visual material based on the text resource and the visual content of the visual material; Based on the text resource, the initial text and the description text, the initial text is updated to obtain the target text, and the target visual material matching the target text is determined from the visual material.

3. The method according to claim 2, wherein: Based on the text resource, the initial text and the description text, updating the initial text to obtain the target text includes: Determining, from the visual material, an intermediate visual material matching the initial text based on the text resource, the initial text and the description text; The initial text is updated based on the description text corresponding to the intermediate visual material to obtain the target text.

4. The method according to claim 3, wherein: Determining the target visual material matching the target text from the visual material includes: The visual material is voted on multiple times based on a voting mechanism to determine the visual material with the highest relevance to the target text as the target visual material.

5. The method according to claim 1, wherein: Generating a target video based on the target text and the target visual material includes: generating a target audio based on the target text; Aligning the target audio and the target visual material based on a timeline; The target video is generated based on the aligned target audio and target visual material.

6. The method according to claim 5, further comprising: Eliminate the existing audio data in the target visual material.

7. The method according to claim 2, wherein: Obtaining visual materials based on the visual resources includes: Preprocessing the video resources in the visual resources to obtain the video materials in the visual material; wherein the preprocessing includes at least one of the following: extracting video slices with different semantics, removing duplicate video slices, identifying key frames in the visual material, or identifying key segments in the visual material; and / or, The image resources in the visual resources are screened to obtain the image materials in the visual material; wherein the screening includes at least one of the following: evaluating the image resources based on a preset evaluation dimension to obtain an evaluation score; and determining the image resources having the evaluation score greater than or equal to a preset score as the image material; or, removing image resources having a similarity lower than a preset similarity.

8. A video generating device, comprising: An acquisition module, used to acquire text resources and visual resources about the target object; A matching module, used for generating a semantically matching target text and target visual material based on the text resource and the visual resource; A generation module is used to generate a target video based on the target text and the target visual material.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the video generation method according to any one of claims 1 to 7 when executing the program. 10 . A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the video generation method according to claim 1 .