Image-text combination-based video generation method and system
Through AI, AI automatically generates videos that combine pictures and text, analyzes user input content and assists in decision-making, solving the problem of high creative threshold for ordinary users and achieving fast and convenient video generation.
Patent Information
- Application Number
- CN202510487438.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-01
AI Technical Summary
In the process of traditional video production, especially the creation of pictures and texts, professional skills and tools are required, which leads to high creative threshold for ordinary users and it is difficult to quickly generate videos that meet their needs.
Through AI, AI automatically generates videos that combine pictures and texts, analyzes user input content, determines video generation needs, and assists users in making quick decisions in the terminal operation space, lowering the threshold for creation.
Without the need for users to process text content, picture design, soundtrack, editing and other operations by themselves, users can easily and efficiently complete video creation, lowering the creation threshold, and improving the convenience and user experience of video generation.
Smart Images

Figure CN120416619A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method and system for generating videos based on a combination of text and images. Background Art
[0002] Currently, with the continuous growth of the demand for digital content, videos have become an important carrier for information dissemination and entertainment.
[0003] However, traditional video production often requires professional skills and tools, and many ordinary users face a relatively high technical threshold when making videos, especially in the process of creating videos that combine text and images. Users not only need to process text content, but also perform cumbersome operations such as graphic design, background music selection, and video editing, which makes it difficult for non-professionals to quickly generate videos that meet their requirements.
[0004] Therefore, there is an urgent need for an innovative solution to simplify the video production process, lower the creation threshold, and enable more users to easily and efficiently complete the creation of videos that combine text and images. Summary of the Invention
[0005] One of the objectives of the present invention is to provide a method for generating videos based on a combination of text and images, which analyzes the obtained user input content to determine the video generation requirements, and based on the video generation requirements, AI automatically generates videos that combine text and images. In the process of creating videos that combine text and images, users do not need to process text content, graphic design, background music selection, video editing, etc. by themselves, which helps them easily and efficiently complete the creation of videos that combine text and images and lowers the creation threshold.
[0006] A method for generating videos based on a combination of text and images provided by an embodiment of the present invention includes: Obtain user input content; Analyze the user input content to determine the video generation requirements; Based on the video generation requirements, AI automatically generates videos that combine text and images.
[0007] Optionally, when the determination of the video generation requirements fails, generate a terminal operation space based on multiple fuzzy expressions of the user input content; Assist the user in quickly making a decision on the video generation requirements based on the terminal operation space.
[0008] Optionally, the generating a terminal operation space based on multiple fuzzy expressions of the user input content includes: Sort each fuzzy expression in ascending order of its respective fuzzy degree to obtain an expression sequence; Determine multiple target local sequences from the expression sequences; wherein, there is content overlap or content generation time overlap between the first support contents of each pair of adjacent fuzzy expressions in each target local sequence in the user input content; Traverse each target local sequence in descending order according to the number of fuzzy expressions in it; Each time when traversing, generate a spatial visual label based on the feature distribution of the traversed target local sequence, and set it at the label position corresponding to the current traversal order in the initial space; wherein, the smaller the current traversal order, the more preferentially the user can visually view the corresponding label position when entering the initial space; Endow the label position corresponding to the current traversal order with an operation support mechanism; After the sequential traversal is completed, use the initial space in which all label positions are set with spatial visual labels and are endowed with the operation support mechanism as the terminal operation space; Wherein, the operation support mechanism includes: When the user stays and views the label position corresponding to the current traversal order for more than the threshold duration, provide the user with the first support content corresponding to each fuzzy expression in the traversed target local sequence for viewing, and record the viewing history; Based on the supplementary intention reflected by the viewing history, crawl the intention-related content, and provide it for the user to view and confirm.
[0009] Optionally, during the process of AI automatically generating a video with pictures and texts, if the user inputs a request to view the preliminary video, provide the user with a linear view of the preliminary video automatically generated by the AI from beginning to end, and record the viewing progress; When the viewing progress is about to reach the trigger progress, based on the viewing history of the user linearly viewing the preliminary video in the past, determine the video playback active control strategy for the local video between the trigger progress where the viewing progress in the preliminary video is about to reach and the next trigger progress; When the viewing progress reaches the trigger progress, start the corresponding playback control for the local video based on the video playback active control strategy.
[0010] Optionally, the steps for determining the trigger progress are as follows: Determine the playback progress that meets the progress constraints from the preliminary video, and use it as the trigger progress; Wherein, the progress constraints include: The content type set of the playback content within the preset progress range before and after the playback progress in the preliminary video matches the standard content type set.
[0011] Optionally, the video playback active control strategy for the local video between the trigger progress where the viewing progress in the preliminary video is about to reach and the next trigger progress, determined based on the viewing history of the user linearly viewing the preliminary video in the past, includes: When the viewing history indicates that there is unacceptable content in the user's history, the decision-making active control strategy includes: When the weighted sum of the user's non-acceptance degree of the unacceptable content and the relevance between the video content of the local video and the unacceptable content exceeds the threshold sum, control the local video to skip playback; Otherwise, based on the second supporting content of the local video in the video generation requirements, match the first target control rule; Based on the first target control rule, control the local video playback; And / or, When the viewing history indicates that there is acceptable content in the user's history, the decision-making active control strategy includes: based on the relevant content distribution of the acceptable content in the local video, match the second target control rule; Based on the second target control rule, control the local video playback.
[0012] A video generation system based on the combination of pictures and texts provided by an embodiment of the present invention includes: A user input content acquisition module for acquiring user input content; A video generation requirement determination module for analyzing the user input content to determine the video generation requirements; An AI video generation module for automatically generating a video combining pictures and texts based on the video generation requirements by AI.
[0013] Optionally, the video generation requirement determination module is further configured to: When the determination of the video generation requirements fails, generate a terminal operation space based on multiple fuzzy expressions of the user input content; Assist the user in quickly making a decision on the video generation requirements based on the terminal operation space.
[0014] Optionally, generating the terminal operation space based on multiple fuzzy expressions of the user input content includes: Sort each fuzzy expression in ascending order of its respective fuzzy degree to obtain an expression sequence; Determine multiple target local sequences from the expression sequence; wherein, for each target local sequence, there is content overlap or content generation time overlap between the first supporting contents of two adjacent fuzzy expressions in the user input content; Traverse each target local sequence in descending order of the number of fuzzy expressions in it; Each time when traversing, generate a spatial visual label based on the feature distribution of the traversed target local sequence, and set it at the label position corresponding to the current traversal order in the initial space; wherein, the smaller the current traversal order, the more priority the user has to view the corresponding label position when entering the initial space; Endow the label position corresponding to the current traversal order with an operation support mechanism; After the sequential traversal is completed, all tag positions are set as the terminal operation space with the completed spatial visual tags and the initial space endowed with the completion operation support mechanism; Among them, the operation support mechanism includes: When the user stays to view the tag position corresponding to the current traversal order for more than the threshold duration, the user is provided with the first support content corresponding to each fuzzy expression in the target local sequence traversed, and the viewing history is recorded; Based on the supplementary intention reflected by the viewing history, content related to the intention is crawled, and provided for the user to view and confirm.
[0015] Optionally, the AI video generation module is also used for: During the process of AI automatically generating a video combining text and images, if the user inputs a request to view the preliminary video, the user is provided with the preliminary video that has been automatically generated by AI to view linearly from beginning to end, and the viewing progress is recorded; When the viewing progress is about to reach the trigger progress, based on the viewing history of the user linearly viewing the preliminary video in the past, the video playback active control strategy for the local video between the trigger progress where the viewing progress in the preliminary video is about to reach and the next trigger progress is determined; When the viewing progress reaches the trigger progress, based on the video playback active control strategy, corresponding playback control of the local video is started.
[0016] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the drawings.
[0017] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings
[0018] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings: Figure 1 It is a schematic diagram of a method for generating a video combining text and images in an embodiment of the present invention; Figure 2 It is a schematic diagram of a system for generating a video combining text and images in an embodiment of the present invention. Detailed Embodiments
[0019] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to explain and illustrate the present invention, and are not used to limit the present invention. Embodiment 1
[0020] An embodiment of the present invention provides a method for generating a video based on a combination of text and images, as Figure 1 shown, including: S1. Obtain user input content; S2. Analyze the user input content to determine the video generation requirements; S3. Based on the video generation requirements, use AI to automatically generate a video combining text and images.
[0021] The user input content at least includes: the video style, video length, displayed text content, displayed pictures, etc. that the user expects to generate in the video; the video generation requirements obtained by analyzing the user input content are the user's requirements for the generated video. The neural network is pre-trained using a large number of processes for generating videos combining text and images on demand, artificial experience, etc. until convergence to obtain an AI that can automatically generate videos combining text and images on demand. Based on the video generation requirements, use AI to automatically generate a video combining text and images.
[0022] Specifically, for example: if the video generation requirements are a relaxed and concise video style, a 1-minute video length, the text content involves information related to healthy diet, the picture materials include fruit pictures, and the background music used is relaxing background music, then the AI automatically generates a 1-minute video. At the beginning of the video, display the text "Healthy diet makes life better" and match it with fresh fruit pictures; then, display some suggestions for healthy diet, such as the text "Eat more fruits and vegetables every day and eat less greasy food", accompanied by pictures of vegetables and relaxing background music, and finally end with "Choose a healthy diet and embrace a beautiful life".
[0023] This application analyzes the obtained user input content to determine the video generation requirements. Based on the video generation requirements, the AI automatically generates a video combining text and images. In the process of creating a video combining text and images, the user does not need to process the text content, screen design, background music, editing, etc. by himself, which helps him complete the creation of a video combining text and images easily and efficiently, and reduces the creation threshold. Embodiment 2
[0024] When the user has too many requirements for the video generated by the AI, it may be impossible to express the requirements accurately, systematically and comprehensively when inputting the content. At this time, the system will have difficulties in understanding when analyzing the user input content, that is, the video generation requirements determination fails. At this time, if the system asks the user to sort out the details in the input content one by one to clarify the specific requirements, it will often increase the burden on the user. Especially when the user lacks sufficient understanding of the video generation process or the required technical details, it is not only easy to make the user feel confused, but may even cause their irritable mood, thinking that the video generation process is both cumbersome and inefficient.
[0025] Therefore, solving the above problems is the key to improving the quality of AI-generated video services. In one embodiment, when the determination of video generation requirements fails, a terminal operation space is generated based on multiple fuzzy expressions of the user input content; Assist the user in quickly making decisions about video generation requirements based on the terminal operation space.
[0026] Fuzzy expressions refer to the fuzzy video generation requirements expressed in the user input content. Through semantic understanding technology, by analyzing and reasoning the user input content, potential intentions and requirements can be extracted to form fuzzy expressions. Specifically, first, the user input content is tokenized and syntactically analyzed to identify key information such as style, theme, or plot; then, through context reasoning and semantic speculation, possible requirements are inferred. For example, if the user says "hoping the video is a bit mysterious", it is speculated that the user expects the video to contain suspense or fantasy elements, which is used as a fuzzy expression.
[0027] Based on each fuzzy expression, a terminal operation space is generated. In the terminal operation space, the user can quickly make decisions about video generation requirements based on each fuzzy expression, and the system will also assist the user in quickly making decisions about video generation requirements based on the terminal operation space.
[0028] In the embodiment of the present invention, when the determination of video generation requirements fails, a terminal operation space is generated based on multiple fuzzy expressions of the user input content, and the user is assisted in quickly making decisions about video generation requirements based on the terminal operation space, helping the user quickly make decisions about video generation requirements, avoiding directly asking the user to sort out the details in their input content one by one, which would increase their burden, and more importantly, avoiding confusing the user, etc., improving the convenience of video generation and enhancing the user experience. Embodiment 3
[0029] In one embodiment, the generating a terminal operation space based on multiple fuzzy expressions of the user input content includes: Sort each fuzzy expression in ascending order of its respective fuzzy degree to obtain an expression sequence; Determine multiple target local sequences from the expression sequence; where, for each pair of adjacent fuzzy expressions in each target local sequence, there is content overlap or content generation time overlap between their first supporting contents in the user input content; Traverse each target local sequence in descending order of the number of fuzzy expressions in it; Each time when traversing, based on the feature distribution of the traversed target local sequence, generate a spatial visual label and set it at the label position corresponding to the current traversal order in the initial space; where, the smaller the current traversal order, the more preferentially the user can view the corresponding label position when entering the initial space; Endow the label position corresponding to the current traversal order with an operation support mechanism; After the sequential traversal is completed, all tag positions are set as the terminal operation space with the completed spatial visual tags and the initial space given the completion operation support mechanism; Among them, the operation support mechanism includes: When the user stays at the tag position corresponding to the current traversal order for more than the threshold duration, the user is provided with the first supporting content corresponding to each fuzzy expression in the target local sequence traversed, and the viewing history is recorded; Based on the supplementary intention reflected by the viewing history, the content related to the intention is crawled, and provided for the user to view and confirm.
[0030] The ambiguity represents the degree of ambiguity of the fuzzy expression. For example, when processing the user input content based on semantic understanding technology, the smaller the possibility of inferring that the user has the need for fuzzy expression (such as the fewer keywords related to the fuzzy expression in the user input content), the greater the degree of ambiguity.
[0031] The first supporting content refers to the relevant content in the user input content that supports the inference that the user has a fuzzy expression when processing the user input content based on semantic understanding technology. For example, if the fuzzy expression is that the video expected by the user contains suspense or fantasy elements, the first supporting content is the expected statement "hoping the video is a bit mysterious".
[0032] There is content overlap means that there is overlapping content between two first supporting contents, and there is generation time overlap means that the time periods of the two first supporting contents input by the user overlap.
[0033] Sort each fuzzy expression from smallest to largest to form an expression sequence, so that the ambiguity of adjacent fuzzy expressions is as close as possible. Constrained by the condition that there is content overlap or generation time overlap between the first supporting contents corresponding to two adjacent fuzzy expressions, multiple target local sequences are determined from the expression sequence, so that multiple fuzzy intentions in the same target local sequence are the intentions that the user is entangled or unable to clarify the expression idea at a certain stage when inputting the content (at a certain stage when inputting the content, a clear intention should be expressed. If entangled or unable to clarify the expression idea, the input content may be logically disordered, resulting in the expression of multiple fuzzy intentions, and the content expressing these fuzzy intentions will overlap in content or generation time).
[0034] When determining the intentions that the user is entangled or unable to clarify the expression idea at a certain stage when inputting the content, it is necessary to start helping the user solve them one by one. The larger the number of fuzzy expressions in the target local sequence, the more intentions the user is entangled or unable to clarify the expression idea in the corresponding stage, and the more priority is given to solving them. That is, traverse each target local sequence in descending order according to the number of fuzzy expressions.
[0035] The feature distribution of the target local sequence at least includes: the types of fuzzy expressions in the target local sequence, the types of the respective corresponding first support contents, etc. Based on this feature distribution, a spatial visual label is generated, and the spatial visual label is a virtual label for displaying this feature distribution. The initial space is a three-dimensional virtual space, in which there are multiple label positions, each corresponding to a current traversal order. The current traversal order refers to the order of traversing to the target local sequence currently, that is, which target local sequence it is. The smaller the traversal order, the more priority the user needs to view and solve correspondingly after entering the initial space. Then, the more priority the corresponding label position is visible to the user when entering the initial space. The spatial visual label is set at the label position corresponding to the current traversal order in the initial space.
[0036] An operation support mechanism is given to the label position corresponding to the current traversal order. In the operation support mechanism, the threshold duration is the duration representing a relatively long stay and view time. When the user stays and views the label position corresponding to the current traversal order for more than the threshold duration, it indicates that the user is interested in the spatial visual label beside the label position and starts to solve it, so as to provide the first support content corresponding to each fuzzy expression in the target local sequence traversed for the user to view. The user can recall ideas during the viewing process. The system records its viewing history, at least including: the types of viewed contents, the order of viewing different contents, etc. The viewing history will reflect the current supplementary intention of the user. For example, if the viewing history reflects that the user continuously views multiple photos of the same fruit, the reflected supplementary intention is to hope to supplement more fruit photos. Based on the supplementary intention, content related to the intention is crawled, such as: more fruit photos. The content related to the intention is provided for the user to view and confirm, such as: the user views more fruit photos and confirms some of the photos directly as video generation requirements.
[0037] When the user enters the created terminal operation space through a smart phone, etc., the user will view the spatial visual label that needs to be processed with higher priority. When the user stays and views its label position for a long time, the operation support mechanism given to the label position will continue to execute, helping the user quickly view and confirm the content related to the intention whose supplementary intention is reflected in its viewing history, and finally realizing the accurate and clear determination of video generation requirements.
[0038] In the embodiments of the present invention, a terminal operation space is generated based on multiple fuzzy expressions in the user input content, which effectively helps the user clarify their thoughts and quickly locate their needs, thereby improving the accuracy of the subsequent AI-generated video combining text and images. By sorting the fuzzy expressions according to their degrees of fuzziness and determining the target local sequence, the system can identify the intentions that the user has not fully clarified during the input process, and preferentially display the positions of the tags to be processed after the user enters the virtual space, improving the efficiency of the user in processing these unclarified intentions. When the user stays at a tag position for a long time, the system provides the user with access to the first supporting content corresponding to each fuzzy expression in the target local sequence, and further crawls relevant content for the user to select according to the supplementary intentions reflected by the user's access history, accurately meeting the needs and reducing the troubles caused by the ambiguity of the expression, significantly improving the user experience as a whole. Embodiment 4
[0039] During the process of AI automatically generating a video, the user may want to view the preliminary video that the AI has already generated. However, when viewing the preliminary video, the user often adopts a skipping viewing method to quickly judge whether the video meets their actual expectations. Such a viewing method may lead to misjudgment of the overall effect of the video by the user. Especially when the user only focuses on some segments and ignores the overall content, it is easy for them to prematurely think that the AI cannot generate a video that meets their needs. For example, the user may only evaluate the quality of a certain part of the content in the preliminary video, without comprehensively considering the overall structure of the video, the generation of subsequent content, and the detailed adjustment, thus making an inaccurate judgment on the generation ability of the AI.
[0040] Therefore, to solve the above problems, in one embodiment, during the process of AI automatically generating a video combining text and images, if the user inputs a request to view the preliminary video, the user is provided with a linear view of the preliminary video that the AI has automatically generated from beginning to end, and the viewing progress is recorded; When the viewing progress is about to reach the trigger progress, based on the viewing history of the user linearly viewing the preliminary video in the past, determine the active control strategy for video playback of the local video between the trigger progress where the viewing progress in the preliminary video is about to reach and the next trigger progress; When the viewing progress reaches the trigger progress, start the corresponding playback control of the local video based on the active control strategy for video playback.
[0041] Linear video viewing means watching video content in a continuous and sequential manner, where the user watches the video from start to end without being able to skip certain segments. The viewing progress refers to the progress of viewing the preliminary video. When the viewing progress is about to reach the trigger progress, it indicates that the user may misunderstand the subsequent video playback content and misjudge the overall effect of the video. Therefore, it is necessary to actively control the playback of the local video between the trigger progress that the viewing progress is about to reach and the next trigger progress to avoid this situation. When the viewing progress reaches the trigger progress, based on the active video playback control strategy, corresponding playback control of the local video starts.
[0042] The embodiments of the present invention effectively solve the problem of misjudgment that may occur when users view the preliminary samples of AI-generated videos. By forcing users to watch the video linearly and evaluate the video content completely from beginning to end, one-sided judgments caused by skipping viewing are avoided. In addition, when the user's viewing progress is about to reach the trigger progress, the system will actively control the video playback based on the user's historical viewing behavior to ensure that the user can fully understand each segment of content and avoid misjudging the overall effect of the video due to neglecting certain details or subsequent generated parts. Embodiment 5
[0043] In one embodiment, the steps for determining the trigger progress are as follows: Determine the playback progress that meets the progress constraints from the preliminary video and use it as the trigger progress; Among them, the progress constraints include: The set of content types of the playback content within the preset progress range before and after the playback progress in the preliminary video matches the standard content type set.
[0044] The progress range can be a 5-second playback range. The standard content type set contains multiple content types that jointly represent sudden changes in picture style (such as switching from a natural landscape or a clear picture to a fast-flashing, complex picture or abstract art), sudden changes in sound effects (switching from a quiet or soft sound effect to a loud sound effect), or sudden changes in language style (suddenly switching from formal and professional language to slang). When the set of content types of the playback content within the preset progress range before and after the playback progress in the preliminary video matches the standard content type set, when the viewing progress is about to reach the corresponding playback progress, the user may misunderstand the subsequent video playback content and misjudge the overall effect of the video.
[0045] The embodiments of the present invention determine the playback progress that meets the progress constraints from the preliminary video by setting progress constraints and use it as the trigger progress, improving the accuracy and efficiency of trigger progress determination. Embodiment 6
[0046] In one embodiment, the video playback active control strategy for the partial video between the trigger progress that is about to be reached in the viewing progress of the preliminary sample video and the next trigger progress, which is determined based on the viewing history of the user linearly viewing the preliminary sample video in the past, includes: When the viewing history indicates that there is unacceptable content in the user's history, the decision-making active control strategy includes: When the weighted sum of the user's non-acceptance degree of the unacceptable content and the relevance between the video content of the partial video and the unacceptable content exceeds the threshold sum, control the partial video to skip playback; Otherwise, based on the second supporting content of the partial video in the video generation requirements, match the first target control rule; Based on the first target control rule, control the playback of the partial video; And / or, When the viewing history indicates that there is acceptable content in the user's history, the decision-making active control strategy includes: based on the relevant content distribution of the acceptable content in the partial video, match the second target control rule; Based on the second target control rule, control the playback of the partial video.
[0047] The unacceptable content refers to the video content that the user found unacceptable when viewing the preliminary sample video in the past. It has a non-acceptance degree, which represents the degree to which the user cannot accept it. The user can set the unacceptable content and the non-acceptance degree according to their actual viewing situation. The relevance between the video content of the partial video and the unacceptable content refers to the degree of relevance between the two contents. When calculating the weighted sum of the two, the non-acceptance degree and the relevance are respectively multiplied by their preset weights to obtain two products, and the sum of them is the weighted sum. The specific formula is: weighted sum = (non-acceptance degree × non-acceptance degree weight) + (relevance × relevance weight). The threshold sum is a preset threshold representing a large weighted sum. When the weighted sum exceeds the threshold sum, it means that if the partial video is actively controlled to skip playback (skip this video content), it can prevent the user from seeing the content that may make them unacceptable again. Otherwise, it means that it can be viewed by the user, but it is necessary to match the first target control rule based on the second supporting content (the second supporting content is the basis for the AI to generate the partial video in the video generation requirements. If the partial video is a fruit cutting demonstration screen, the second supporting content is the requirement to display the fruit demonstration screen) and execute it. For example, when controlling the playback of the partial video, display the second supporting content so that the user can understand why the AI does this.
[0048] The content to be received is the video content set by the user. The distribution of the relevant content of the received content in the partial video refers to the distribution position of the content related to the received content in the partial video. Based on its matching with the second target control rule, for example: controlling the partial video to play the content related to the received content in the partial video in a jumping manner in sequence, so that the user can be satisfied with the partial video as soon as possible.
[0049] The embodiments of the present invention can effectively avoid the user's misjudgment of the overall video effect due to some content not meeting expectations.
[0050] The embodiments of the present invention provide a video generation system based on the combination of pictures and texts, as Figure 2 shown, including: A user input content acquisition module 1, which is used to acquire user input content; A video generation requirement determination module 2, which is used to analyze the user input content and determine the video generation requirement; An AI video generation module 3, which is used to automatically generate a video combining pictures and texts based on the video generation requirement by AI.
[0051] The video generation requirement determination module is further used for: When the determination of the video generation requirement fails, generating a terminal operation space based on multiple fuzzy expressions of the user input content; Assisting the user to quickly make a decision on the video generation requirement based on the terminal operation space.
[0052] The generating of the terminal operation space based on multiple fuzzy expressions of the user input content includes: Sorting each fuzzy expression in ascending order of its respective fuzzy degree to obtain an expression sequence; Determining multiple target partial sequences from the expression sequence; wherein, for each target partial sequence, there is content overlap or content generation time overlap between the first support contents of two adjacent fuzzy expressions in the user input content; Traversing each target partial sequence in descending order of the number of fuzzy expressions in it in turn; Each time when traversing, generating a spatial visual label based on the feature distribution of the traversed target partial sequence, and setting it at the label position corresponding to the current traversal order in the initial space; wherein, the smaller the current traversal order, the more preferentially the user can visually see the corresponding label position when entering the initial space; Endowing the label position corresponding to the current traversal order with an operation support mechanism; After the sequential traversal is completed, the initial space with all label positions set with spatial visual labels and endowed with the completed operation support mechanism is used as the terminal operation space; Wherein, the operation support mechanism includes: When the user stays at the label position corresponding to the current traversal order for a duration exceeding the threshold, provide the user with access to the first supporting content corresponding to each fuzzy expression in the target local sequence traversed, and record the access history. Based on the supplementary intention reflected in the access history, crawl the content related to the intention, and provide it for the user to access and confirm.
[0053] The AI video generation module is also used for: During the process of AI automatically generating a video combining text and images, if the user inputs a request to view the preliminary video, provide the user with a linear view of the preliminary video that has been automatically generated by the AI from start to finish, and record the viewing progress. When the viewing progress is about to reach the trigger progress, based on the viewing history of the user linearly viewing the preliminary video in the past, determine the active video playback control strategy for the local video between the trigger progress where the viewing progress in the preliminary video is about to reach and the next trigger progress. When the viewing progress reaches the trigger progress, based on the active video playback control strategy, start the corresponding playback control for the local video.
[0054] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A video generation method based on the combination of text and images, characterized in that, Including: Obtain the user input content; Analyze the user input content to determine the video generation requirements; Based on the video generation requirements, the AI automatically generates a video combining pictures and texts.
2. The video generation method based on the combination of graphics and text according to claim 1, wherein When the determination of the video generation requirements fails, generate a terminal operation space based on multiple fuzzy expressions of the user input content; Assist the user to quickly make a decision on the video generation requirements based on the terminal operation space.
3. The video generation method based on the combination of graphics and text according to claim 2, wherein, The generating of the terminal operation space based on multiple fuzzy expressions of the user input content includes: Sort each fuzzy expression in ascending order of its respective fuzzy degree to obtain an expression sequence; Determine multiple target local sequences from the expression sequence; wherein, there is content overlap or content generation time overlap between the first support contents of each pair of adjacent fuzzy expressions in each target local sequence in the user input content; Traverse each target local sequence in descending order of the number of fuzzy expressions in it in turn; Each time when traversing, generate a spatial visual label based on the feature distribution of the traversed target local sequence, and set it at the label position corresponding to the current traversal order in the initial space; wherein, the smaller the current traversal order, the more preferentially the user can visually see the corresponding label position when entering the initial space; Endow the label position corresponding to the current traversal order with an operation support mechanism; After the sequential traversal is completed, use the initial space with all label positions set with spatial visual labels and endowed with the completed operation support mechanism as the terminal operation space; Wherein, the operation support mechanism includes: When the user stays and views the label position corresponding to the current traversal order for more than the threshold duration, provide the user with the first support content corresponding to each fuzzy expression in the traversed target local sequence for viewing, and record the viewing history; Crawl intention-related content based on the supplementary intention reflected by the viewing history, and provide it for the user to view and confirm.
4. The video generation method based on the combination of graphics and text according to claim 1, wherein, During the process of the AI automatically generating a video combining pictures and texts, if the user inputs a request to view the preliminary video, provide the user with a linear view of the preliminary video automatically generated by the AI from beginning to end, and record the viewing progress; When the viewing progress is about to reach the trigger progress, based on the viewing history of the user's linear viewing of the preliminary video in the past, determine the video playback active control strategy for the local video between the trigger progress where the viewing progress in the preliminary video is about to reach and the next trigger progress; When the viewing progress reaches the trigger progress, start corresponding playback control of the local video based on the video playback active control strategy.
5. The video generation method based on combination of graphics and text according to claim 4, wherein The steps for determining the trigger progress are as follows: Determine the playback progress that meets the progress constraint from the preliminary video and use it as the trigger progress; Wherein, the progress constraint includes: The content type set of the playback content within the preset progress range before and after the playback progress in the preliminary video matches the standard content type set.
6. The video generation method based on the combination of graphics and text according to claim 4, wherein The determining of the video playback active control strategy for the local video between the trigger progress where the viewing progress in the preliminary video is about to reach and the next trigger progress based on the viewing history of the user's linear viewing of the preliminary video in the past includes: When the viewing history indicates that there is unacceptable content in the user's history, determine the active control strategy, including: When the weighted sum of the user's non - acceptance degree of non - accepted content and the relevance between the video content of the local video and the non - accepted content exceeds the threshold sum, control the local video to skip playback; Otherwise, based on the second supporting content of the local video in the video generation requirements, match the first target control rule; Based on the first target control rule, control the local video playback; And / or, When the viewing history indicates that there is accepted content in the user's history, determine the active control strategy, including: based on the relevant content distribution of the accepted content in the local video, match the second target control rule; Based on the second target control rule, control the local video playback.
7. A video generation system based on the combination of images and texts, characterized in that, Including: A user input content acquisition module for acquiring user input content; A video generation requirement determination module for analyzing the user input content to determine the video generation requirements; An AI video generation module for AI - automatically generating a video combining text and images based on the video generation requirements.
8. The video generation system based on the combination of graphics and text according to claim 7, wherein The video generation requirement determination module is also used for: When the determination of video generation requirements fails, generate a terminal operation space based on multiple fuzzy expressions of the user input content; Assist the user in quickly making a decision on video generation requirements based on the terminal operation space.
9. The video generation system based on the combination of graphics and text according to claim 8, characterized in that, Generating the terminal operation space based on multiple fuzzy expressions of the user input content includes: Sort each fuzzy expression in ascending order of its respective fuzzy degree to obtain an expression sequence; Determine multiple target local sequences from the expression sequence; wherein, for each target local sequence, there is content overlap or content generation time overlap between the first supporting content of two adjacent fuzzy expressions in the user input content; Traverse each target local sequence in descending order of the number of fuzzy expressions in it; Each time when traversing, generate a spatial visual label based on the feature distribution of the traversed target local sequence, and set it at the label position corresponding to the current traversal order in the initial space; wherein, the smaller the current traversal order, the more preferentially the user can view the corresponding label position when entering the initial space; Endow the label position corresponding to the current traversal order with an operation support mechanism; After the sequential traversal is completed, use the initial space with all label positions set with spatial visual labels and endowed with the completed operation support mechanism as the terminal operation space; Wherein, the operation support mechanism includes: When the user stays and views the label position corresponding to the current traversal order for more than the threshold duration, allow the user to view the first supporting content corresponding to each fuzzy expression in the traversed target local sequence and record the viewing history; Based on the supplementary intention reflected by the viewing history, crawl the intention - related content and allow the user to view and confirm the selection.
10. The video generation system based on the combination of graphics and text according to claim 7, wherein The AI video generation module is also used for: During the process of AI - automatically generating a video combining text and images, if the user inputs a request to view the preliminary video sample, allow the user to linearly view the AI - automatically generated preliminary video sample from beginning to end and record the viewing progress; When the viewing progress is about to reach the trigger progress, based on the viewing history of the user linearly viewing the preliminary video sample in the past, determine the active control strategy for video playback of the local video between the trigger progress where the viewing progress in the preliminary video sample is about to reach and the next trigger progress; When the viewing progress reaches the trigger progress, based on the active video playback control strategy, corresponding playback control of the local video starts.