Demand-oriented AIGC video generation method and generation system

By collecting and analyzing user multimodal demand information, generating feature label sets, establishing a selection pool and predicting duration, the problem of insufficient understanding of user needs in AIGC video generation is solved, and video generation that better meets user expectations is achieved.

CN120111313BActive Publication Date: 2025-10-17TIANJIN TIANKAI RUIAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510350241.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-10-17
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

Existing AIGC video generation methods find it difficult to accurately capture and understand users' multi-dimensional needs, resulting in a large deviation between the generated video content and user expectations.

Method used

By collecting multimodal demand information of target users, extracting video features, performing feature analysis and label set merging, establishing a selection pool, recording user choices, and using the three-point estimation principle to predict duration, AIGC videos that meet user expectations are generated.

Benefits of technology

Accurately capture and understand users' multi-dimensional needs and generate AIGC videos that better meet user expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120111313B_ABST
    Figure CN120111313B_ABST
Patent Text Reader

Abstract

The application discloses a demand-oriented AIGC video generation method and a generation system, and relates to the field of video content generation. The method comprises the following steps: obtaining a first target demand of a target user; extracting a first feature in a predetermined video feature, performing feature analysis on first demand modal information, and obtaining a first feature label set; performing set operation analysis to obtain a first target label list; establishing a label selection pool and obtaining demand selection records; extracting time length demand information in the first target demand, and obtaining a first most probable estimated time length by using a three-point estimation principle; and generating a target AIGC video by taking the second target demand and the first most probable estimated time length as constraints. The technical problems that the existing video generation cannot accurately capture the multi-dimensional demand of the user and leads to a large deviation between the generated video content and the user's expectation are solved, and the technical effect of accurately capturing the multi-dimensional demand of the user and generating an AIGC video that is more in line with the user's expectation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of video content generation, in particular to a demand-oriented AIGC video generation method and system. BACKGROUND

[0002] AIGC (Artificial Intelligence Generated Content) refers to the process of using artificial intelligence technology to automatically or semi-automatically generate text, images, audio, video and other content, which has high importance in the current digital content creation field. Current video generation mainly relies on preset video templates or general video generation algorithms, which generate video content by receiving user input basic information such as theme, style, etc. However, the current method lacks in-depth understanding and fine processing of user demand, resulting in a large deviation between the generated video content and user expectations, especially when dealing with user requirements with multiple complex demand modalities (such as emotion, scene, style, etc.), the problem is particularly prominent.

[0003] In the related art, AIGC video generation has the technical problem of being unable to accurately capture and understand the multi-dimensional requirements of users, resulting in a large deviation between the generated video content and user expectations. SUMMARY

[0004] The present application provides a demand-oriented AIGC video generation method and system, which collects multi-modal demand information of target users, extracts video features, analyzes demand information, generates a feature label set, merges the feature label set, forms a target label list, establishes a selection pool based on the target label list, records user selection, determines specific requirements, extracts time length information from the requirements, uses the three-point estimation principle to predict the most likely time length, and generates a target AIGC video based on the user's specific requirements and the predicted time length. Technical means, achieve the technical effect of accurately capturing and understanding the multi-dimensional requirements of users, and generating AIGC videos that better meet user expectations.

[0005] The application provides a demand-oriented AIGC video generation method, including: obtaining a first target demand of a target user, wherein the first target demand includes multiple demand modality information; extracting a first feature in a predetermined video feature, and performing feature analysis on first demand modality information in the multiple demand modality information based on the first feature to obtain a first feature label set; performing set operation analysis on the first feature label set to obtain a first target label list; establishing a label selection pool based on the first target label list, and obtaining a demand selection record of the target user on the label selection pool, denoted as a second target demand; extracting time length demand information in the first target demand, and obtaining a first most probable estimated time length by using a three-point estimation principle; generating a target AIGC video by taking the second target demand and the first most probable estimated time length as constraints.

[0006] In a possible implementation, the following processing is performed: the predetermined video feature includes a video scene feature, a video style feature, and a video element feature.

[0007] In a possible implementation, the following processing is performed when generating a target AIGC video by taking the second target demand and the first most probable estimated time length as constraints: a first label in the second target demand is extracted, the first label has a first selection order and a first selection time identifier; a second label in the second target demand is extracted, the second label has a second selection order and a second selection time identifier, and the second selection order and the first selection order are adjacent orders; a second consideration time length of the second label is obtained by comparing the first selection time and the second selection time; a second weight of the second label is determined based on the second selection order and the second consideration time length, and denoted as a second proportion; a label material database is read, and a second target material set of the second label is obtained in combination with the second proportion; a target material set is established based on the second target material set, and the target AIGC video is generated in combination with the first most probable estimated time length.

[0008] In a possible implementation, the following processing is performed when reading a label material database and obtaining a second target material set of the second label in combination with the second proportion: a second material set corresponding to the second label is screened in the label material database; materials of the second proportion in the second material set are randomly extracted to form the second target material set.

[0009] In a possible implementation, the following processing is performed: an average value of a first demand time length and a second demand time length in the time length demand information is taken, denoted as the first most probable estimated time length.

[0010] In a possible implementation, after generating the target AIGC video under the constraints of the second target demand and the first most probable estimated duration, the following processing is further performed: obtaining a target duration modification demand of the target user on the target AIGC video; when the target duration modification demand is to expand the duration, analyzing the duration demand information to obtain a first pessimistic estimated duration; and performing expansion processing on the target AIGC video under the constraint of the first pessimistic estimated duration.

[0011] In a possible implementation, the expansion processing on the target AIGC video under the constraint of the first pessimistic estimated duration is performed by performing the following processing: comparing the first most probable estimated duration and the first pessimistic estimated duration to obtain a target duration modification ratio; adjusting the second ratio based on the target duration modification ratio to obtain a second target material adjustment set; and replacing the second target material set with the second target material adjustment set to assemble the target material set.

[0012] In a possible implementation, after obtaining the target duration modification demand of the target user on the target AIGC video, the following processing is further performed: when the target duration modification demand is to shorten the duration, analyzing the duration demand information to obtain a first optimistic estimated duration; and performing compression processing on the target AIGC video under the constraint of the first optimistic estimated duration.

[0013] The application further provides a demand-oriented AIGC video generation system, comprising: a first target demand obtaining module, configured to obtain a first target demand of a target user, wherein the first target demand comprises a plurality of demand modal information; a feature analysis module, configured to extract a first feature in a predetermined video feature, and perform feature analysis on first demand modal information in the plurality of demand modal information based on the first feature to obtain a first feature label set; a set operation analysis module, configured to perform set operation analysis on the first feature label set to obtain a first target label list; a second target demand obtaining module, configured to assemble a label selection pool based on the first target label list, and obtain a demand selection record of the target user on the label selection pool, denoted as a second target demand; a duration estimation module, configured to extract duration demand information in the first target demand, and obtain a first most probable estimated duration by using a three-point estimation principle; and a target AIGC video generation module, configured to generate a target AIGC video under the constraints of the second target demand and the first most probable estimated duration.

[0014] The demand-oriented AIGC video generation method and system provided in the application can first obtain a first target demand of a target user, wherein the first target demand includes multiple demand modal information, then extract a first feature in a predetermined video feature, and perform feature analysis on first demand modal information in the multiple demand modal information based on the first feature to obtain a first feature label set, further perform set operation analysis on the first feature label set to obtain a first target label list, then establish a label selection pool based on the first target label list, and obtain a demand selection record of the target user on the label selection pool, denoted as a second target demand, further extract time length demand information in the first target demand, and obtain a first most probable estimated time length by using a three-point estimation principle, and finally generate a target AIGC video by taking the second target demand and the first most probable estimated time length as constraints. The technical effect of accurately capturing and understanding the multi-dimensional demand of the user is achieved, so that the AIGC video more in line with the user's expectation is generated. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings of the embodiments of the application will be briefly introduced below. The flowchart is used to illustrate the operations performed by the system according to the embodiments of the application in the present application. It should be understood that the foregoing or the following operations are not necessarily performed in sequence. On the contrary, various steps can be processed in reverse order or simultaneously according to needs. At the same time, other operations can be added to these processes, or one or more steps of operation can be removed from these processes.

[0016] Figure 1 The flowchart of the demand-oriented AIGC video generation method provided for the embodiments of the application.

[0017] Figure 2 The structure diagram of the demand-oriented AIGC video generation system provided for the embodiments of the application.

[0018] Explanation of reference numerals: first target demand acquisition module 10, feature analysis module 20, set operation analysis module 30, second target demand acquisition module 40, time length estimation module 50, target AIGC video generation module 60. DETAILED DESCRIPTION

[0019] The above description is only a summary of the technical solutions of the application. In order to more clearly understand the technical means of the application, the application can be implemented according to the content of the specification, and in order to make the above and other purposes, features and advantages of the application more obvious and easy to understand, the following specific embodiments of the application are described.

[0020] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings, and the described embodiments should not be regarded as limitations of the present application. All other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0021] In the following description, "some embodiments" are involved, which describe a subset of all possible embodiments, but it can be understood that "some embodiments" can be the same or different subsets of all possible embodiments, and can be combined with each other without conflict. The term "first\second" involved only distinguishes similar objects and does not represent a specific order of the objects. The terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or server including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or modules that are not clearly listed or inherent to these processes, methods, products or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application.

[0022] The embodiments of the present application provide a demand-oriented AIGC video generation method, as shown in Figure 1 The method comprises the following steps:

[0023] In step S100, a first target demand of a target user is obtained, wherein the first target demand comprises a plurality of demand modal information.

[0024] Specifically, the specific demands of the target user for video generation are collected through user research, questionnaire survey, online form, interactive interface and the like. These demands include the theme, style, content, audience, duration and the like of the video. The collected demand information is sorted and analyzed to form the first target demand comprising a plurality of demand modal information. The target user refers to a specific individual or organization that hopes to use the AIGC video generation method to generate a video. The demand modal information refers to the demands of the user for different aspects of video generation, such as theme, style, content, etc., which exist in different forms or modalities.

[0025] In step S200, a first feature in a predetermined video feature is extracted, and a first demand modal information in the plurality of demand modal information is analyzed based on the first feature to obtain a first feature label set.

[0026] In particular, the predetermined video features are a series of features predefined in video processing and analysis, which are used to describe the content or attributes of a video. The predetermined video features can include color, texture, shape, motion, scene, etc. Any one of these predetermined features is extracted as a first feature, such as color, scene, etc. The first demand modality information (any one of the plurality of demand modality information, such as the theme of the video) is analyzed by feature analysis, and the key information related to the first feature is extracted. According to these information, a first feature label set is generated, which corresponds to the selected feature and demand modality information.

[0027] In one possible implementation, step S200 further includes step S210, and the predetermined video features include video scene features, video style features, and video element features. Specifically, the predetermined video features include three main aspects: video scene features, video style features, and video element features. The video scene features are used to describe the scene setting in the video, the video style features are used to describe the overall visual and feeling style of the video, and the video element features are used to describe the specific elements appearing in the video.

[0028] Step S300, performing a set operation analysis on the first feature label set to obtain a first target label list.

[0029] In particular, a set operation is performed on all labels in the first feature label set, that is, all labels are merged and duplicates are removed. The result obtained is the first target label list, which contains all unique labels related to the predetermined video features.

[0030] The following is a specific example. A target user is an online education platform, and they want to generate a video introducing a programming course. Through user research and online forms, the following first target demand is collected, which contains multiple demand modality information:

[0031] Demand modality information 1 (theme): introduction of programming course

[0032] Demand modality information 2 (style): modern and concise

[0033] Demand modality information 3 (content): course highlights, faculty strength, and student evaluation display

[0034] Demand modality information 4 (audience): college students and new job seekers

[0035] The predetermined video features include predetermined video feature 1 (video scene feature), predetermined video feature 2 (video style feature), and predetermined video feature 3 (video element feature). These predetermined video features are analyzed with each of the demand modality information in the first target demand to generate a first feature label set, as shown in Table 1.

[0036] Table 1: First feature label set

[0037]

[0038] The first feature label set is sorted, duplicates are removed, and the first target label list is combined, as shown in Table 2.

[0039] Table 2: First target label list

[0040] Core theme Introduction to programming courses Visual style Modern, simple, retro cartoon, technological education atmosphere Content focus Course highlights, faculty strength, student evaluation Target audience College students, new employees, users interested in programming courses, people who value efficient learning and practical skills Scene design Classroom scene (introduction of the lecturer), online learning platform scene (course demonstration), outdoor teaching scene (practical cases), simple and bright layout Visual elements Lecturer image, student interaction, programming code animation, computer screen interface, data chart visualization Animation and graphics Simple lines / icons, dynamic chart display, gradient transition effect Production logic Switch scenes according to content, unify style

[0041] Step S400, based on the first target label list, a label selection pool is established, and the demand selection record of the target user for the label selection pool is obtained, denoted as the second target demand.

[0042] Specifically, according to the first target label list, a label selection pool containing multiple label categories is established, such as a target scene label list, a target style label list, and a target element label list. Through an interactive interface or a questionnaire, etc., let the target user select the labels they need from the label selection pool. Record the user's selection results to form the second target demand. The second target demand reflects the user's specific demand and preference for video generation.

[0043] Step S500, extracting the duration demand information in the first target demand, and using the three-point estimation principle to obtain the first most possible estimated duration.

[0044] Specifically, the specific demand information about the video duration is extracted from the first target demand. The three-point estimation principle is used to estimate the most possible duration of video generation. For example, the user provides a duration range of 2-5 minutes, and the three-point estimation formula is applied: wherein, is the first most possible estimated duration, is the shortest possible duration (the minimum value provided by the user), is the most possible duration (the middle value provided by the user), is the longest possible duration (the maximum value provided by the user), calculated as follows: = = 3.5 minutes. If the user provides duration demand information in relative description (such as "short video" or "long video"), it is converted into a specific numerical value according to the preset rule. The preset duration rule is shown in Table 3, for example.

[0045] Table 3: Preset duration rule table

[0046] User input type Optimistic duration (minutes) Possible duration (minutes) Pessimistic duration (minutes) Short video 1 1.5 2 Medium video 2 3 5 Long video 5 7 10

[0047] In a possible implementation, step S500 further includes step S510, taking the mean of the first demand duration and the second demand duration in the duration demand information as the first most probable estimated duration.

[0048] Specifically, specific numerical value or range information related to video duration is identified and extracted from the first target demand. These information includes user explicitly specified duration (such as “3-minute video”), duration range (such as “between 2-5 minutes”) or relative description of duration (such as “short video” “long video” and the like, which needs to be converted into specific numerical value in combination with system preset duration standard). In the extracted duration demand information, at least two different duration values or duration ranges are identified as the first demand duration and the second demand duration. If the user only provides one duration value or range, a second duration value or range similar or related to the first one is generated according to preset rules for calculation. The first demand duration and the second demand duration are mathematically averaged. If the duration is in numerical form, arithmetic mean is directly calculated; if it is in range form, the upper and lower limits of the range need to be averaged respectively, or the range is converted into a single numerical value (such as taking the midpoint) according to system rules and then averaged. The calculated mean is taken as the first most probable estimated duration and recorded for subsequent steps. This implementation can generate a video with duration more consistent with user's expectation by considering multiple duration information provided by the user, thereby improving user's satisfaction.

[0049] Step S600, generating a target AIGC video with the second target demand and the first most probable estimated duration as constraints.

[0050] Specifically, according to the second target demand (user's selection result in the tag selection pool) and the first most probable estimated duration (the most probable duration of video generation), a target video is generated using AIGC technology, including multiple steps of generation, editing, synthesis and the like of video content, to ensure that the generated video meets the user's demand and preference and is completed within the specified duration. The embodiments of the present application adopt the technical means of collecting multi-modal demand information of target users, extracting video features, analyzing features of demand information, generating feature label set, merging feature label set, forming target label list, establishing selection pool based on target label list, recording user selection, determining specific demand, extracting duration information from demand, predicting the most probable duration using three-point estimation principle, combining user's specific demand and predicted duration, and generating target AIGC video, to achieve the technical effects of accurately capturing and understanding user's multi-dimensional demand, thereby generating AIGC video more consistent with user's expectation.

[0051] In one possible implementation, the target AIGC video is generated with the second target demand and the first most likely estimated duration as constraints, and step S600 further includes step S610, extracting the first tag in the second target demand, wherein the first tag has an identifier of the first selection order and the first selection time. Specifically, by parsing the user's selection record, which contains the order and timestamp of the user's selection of each tag, the first tag in the second target demand is extracted, and its first selection order and first selection time are identified. Among them, the first tag / second tag is extracted from the target user's demand selection record for the tag selection pool, representing the user's specific preferences for video content, style, elements, etc. Each tag carries an identifier of the selection order and selection time, which is used to analyze the user's preference strength and order. The selection order indicates the order in which the user selects the tags, reflecting the user's preliminary judgment on the importance of different tags.

[0052] Step S620 extracts the second tag from the second target requirement. The second tag is identified by a second selection order and a second selection time, and the second selection order is adjacent to the first selection order. Specifically, the second tag from the second target requirement is extracted and its second selection order and second selection time are identified. This step is also performed based on the user's selection history, ensuring that selection information for all relevant tags is captured.

[0053] Step S630: Compare the first selection time with the second selection time to obtain a second consideration duration for the second tag. Specifically, the first selection time and the second selection time are compared, and the difference between the selection times of adjacent tags is calculated to obtain the second consideration duration for the second tag, which reflects the user's hesitation between different tags.

[0054] Step S640: Determine the second weight of the second tag based on the second selection order and the second consideration time, and record it as the second ratio. Specifically, determine the second weight of the second tag based on the second selection order and the second consideration time. This step is achieved through a weight calculation model that takes into account two factors: the selection order (the earlier the tag is selected, the greater the weight) and the consideration time (the shorter the selection time, the greater the weight). Specifically, a weighting function is designed that takes the selection order and the consideration time as input and outputs a weight value. Weight formula: ,in, is the weight of the i-th label, is the order in which the user selects the i-th tag (such as 1, 2, 3, etc.), is the consideration time for the user to select the i-th tag, and is the weight coefficient.

[0055] The example calculation is as follows: set = 0.6, = 0.4 (which can be adjusted according to actual needs), the order of the user selecting label A is 1, and the consideration duration is 5 seconds; the order of label B is 2, and the consideration duration is 10 seconds. = 0.68, = 0.34, normalize the weights of all labels, and the normalized weight .

[0056] Step S650, read the label material database, and combine the second proportion to obtain a second target material set of the second label. Specifically, the label material database stores a large number of materials associated with labels, including pictures, video clips, audio clips, etc., which are classified and marked as being associated with specific labels for generating AIGC videos that meet user needs. Read the label material database, and combine the second proportion to obtain a second target material set of the second label, that is, search for materials associated with the second label in the database, and filter according to the weight, the greater the weight, the higher the proportion of the corresponding material in the final material set.

[0057] Step S660, based on the second target material set, assemble a target material set, and combine the first most probable estimated duration to generate the target AIGC video. Specifically, by combining the materials corresponding to different labels, a complete target material set is formed, and then using video editing software or algorithm, according to the first most probable estimated duration, the target AIGC video is finally generated by editing, synthesizing, etc. that meets the user's needs. This implementation way analyzes the user's selection order and time of the label, to determine the importance of different labels in the user's preference, and accordingly filters the most suitable material from the label material database to generate the target AIGC video, which can more accurately capture the user's video generation needs, and improve the quality and satisfaction of video generation.

[0058] In a possible implementation, reading the label material database and combining the second proportion to obtain a second target material set of the second label, step S650 further includes step S651, filtering a second material set corresponding to the second label in the label material database. Specifically, access the label material database through a database query interface, and the system sends a request containing the second label as a query condition. Then, the database returns a list or set of all materials associated with the second label according to this query condition, that is, the second material set.

[0059] Step S652, randomly extract the second proportion of materials in the second material set to form the second target material set. After obtaining the second material set, the materials in the second material set are randomly selected according to the second proportion. For example, a random sampling method can be used, and a random number is generated by the system to select a material from the second material set. This process is repeated multiple times until the number of materials specified by the second proportion is reached. This implementation can more accurately capture the user's video generation needs by calculating the weight of each label (second proportion), and accordingly filter the most suitable materials from the database, improve the quality and satisfaction of video generation, and reduce unnecessary calculation and resource consumption.

[0060] In a possible implementation, after generating the target AIGC video with the second target demand and the first most likely estimated duration as constraints, the method further includes step S700 of obtaining a target duration modification demand of the target user for the target AIGC video. Specifically, the input of the target user is received through a user interface (UI), and the user can express the modification demand for the video duration by dragging the time axis, inputting specific numbers, or using preset options. That is, the target duration modification demand is the modification requirement of the target user for the generated target AIGC video in terms of duration, which may be to make the video longer or shorter.

[0061] Step S800, when the target duration modification demand is to expand the duration, a first pessimistic estimated duration is obtained by analyzing the duration demand information. Specifically, when the target duration modification demand is to expand the duration, based on the duration demand information in the first target demand and the target duration modification demand, a three-point estimation principle is used to calculate the first pessimistic estimated duration for video expansion processing. At this time, in the three-point estimation formula, is the first pessimistic estimated duration, remains unchanged, and is adjusted according to the target duration modification demand.

[0062] Step S900, expand the target AIGC video with the first pessimistic estimated duration as a constraint. Specifically, according to the first pessimistic estimated duration, an expansion plan is made, and corresponding editing tools and resources are called to execute the plan, including adding new video segments, adjusting the playback speed of existing segments, adding transition effects, background music or narration, etc., to ensure that the final video duration meets the user's requirements while maintaining the quality and coherence of the video. This implementation can accurately expand the video duration without damaging the overall coherence and quality of the video, and meet the user's modification demand for the video duration.

[0063] In a possible implementation, the target AIGC video is expanded with the first pessimistic estimated length as a constraint, and step S900 further includes step S910 of comparing the first most likely estimated length with the first pessimistic estimated length to obtain a target length modification ratio. Specifically, the first most likely estimated length (i.e., the length estimated according to the original user demand) and the first pessimistic estimated length (i.e., the longest length estimated by the system to meet the user's demand for an expanded length) are obtained. Then, the target length modification ratio is obtained by calculating the difference or proportional relationship between the two. For example, the first most likely estimated length is T1, and the first pessimistic estimated length is T2 (T2 is greater than T1), and the target length modification ratio R = (T2-T1) / T1. This ratio reflects the relative degree of increase in the user's demand for the video length.

[0064] Step S920, adjusting the second ratio based on the target length modification ratio to obtain a second target material adjustment set. Specifically, according to the target length modification ratio R, the second ratio (i.e., the material weight or display length ratio) of each label determined previously is adjusted. The principle of adjustment is to keep the relative importance between labels unchanged, but the overall length is expanded by a ratio of R. The specific method is as follows: for each label, the original second ratio is P(i), and the adjusted ratio P'(i) = P(i)*(1+R). Then, the system reselects the materials in the label material database according to the adjusted ratio P'(i) to form a second target material adjustment set. The number of materials in this set will be adjusted according to the new ratio to meet the user's demand for an expanded length.

[0065] Step S930, the target material set is assembled by replacing the second target material set with the second target material adjustment set. Specifically, the second target material set previously filtered based on the original second proportion is replaced with the second target material adjustment set. Then, based on the new target material set (i.e., the set containing the adjusted materials) and the first pessimistic estimated length (as the new length constraint), the target material set is reassembled and the expanded AIGC video is generated. The specific implementation is as follows: the system first clears the previous target material set, and then adds the materials in the second target material adjustment set to the target material set according to the adjusted proportion and order. Finally, the system clips and edits the target material set according to the first pessimistic estimated length to ensure that the final generated AIGC video length meets the user's requirements. This implementation determines the relative degree of increase in video length that the user wants by calculating the target length modification proportion. Then, based on this proportion, the second proportion of the label is adjusted to ensure that the relative importance between the labels remains unchanged while expanding the video length. Finally, the original material set is replaced with the adjusted material set, and the expanded AIGC video is generated based on the new length constraint, which meets the user's length requirement while maintaining the coherence of the video content and the rationality of the label weight, thereby improving the user experience and satisfaction.

[0066] In a possible implementation, after obtaining the target user's target length modification demand for the target AIGC video, the method further includes step S1000 of analyzing the length demand information to obtain a first optimistic estimated length when the target length modification demand is to shorten the length. Specifically, similar to step S800, when the target length modification demand is to shorten the length, based on the length demand information in the first target demand and the target length modification demand, a three-point estimation principle is used to calculate the first optimistic estimated length for video compression processing.

[0067] Step S1100, the target AIGC video is compressed with the first optimistic estimated length as a constraint. Specifically, the system formulates a video compression strategy according to the first optimistic estimated length. Video editing techniques such as editing, acceleration, and merging are used to compress the target AIGC video. During the compression process, the system constantly monitors the length and content coherence of the video to ensure that the final output video meets the user's desired length and maintains the integrity and coherence of the content. The system presents the compressed video to the user for viewing and feedback. This implementation can accurately shorten the video length while maintaining the core content and quality of the video through intelligent length compression, meeting the user's modification demand for the video length and improving the compactness and watchability of the video.

[0068] In the foregoing, reference is made to Figure 1The demand-oriented AIGC video generation method according to the embodiments of the present application is described in detail. Next, the demand-oriented AIGC video generation method according to the embodiments of the present application will be described in detail with reference to the accompanying drawings. Figure 2 The demand-oriented AIGC video generation system according to the embodiments of the present application is described.

[0069] The demand-oriented AIGC video generation system according to the embodiments of the present application is used to solve the technical problem that it is difficult to accurately capture and understand the multi-dimensional requirements of users in the prior art, resulting in a large deviation between the generated video content and the user's expectations, to achieve the technical effect of accurately capturing and understanding the multi-dimensional requirements of users, thereby generating AIGC videos that are more in line with the user's expectations. The demand-oriented AIGC video generation system includes a first target requirement acquisition module 10, a feature analysis module 20, a set operation analysis module 30, a second target requirement acquisition module 40, a duration estimation module 50, and a target AIGC video generation module 60.

[0070] The first target requirement acquisition module 10 is configured to acquire a first target requirement of a target user, wherein the first target requirement includes a plurality of requirement modal information; the feature analysis module 20 is configured to extract a first feature from a predetermined video feature, and perform feature analysis on a first requirement modal information from the plurality of requirement modal information based on the first feature, to obtain a first feature label set; the set operation analysis module 30 is configured to perform set operation analysis on the first feature label set, to obtain a first target label list; the second target requirement acquisition module 40 is configured to form a label selection pool based on the first target label list, and acquire a requirement selection record of the target user on the label selection pool, denoted as a second target requirement; the duration estimation module 50 is configured to extract duration requirement information from the first target requirement, and obtain a first most likely estimated duration using a three-point estimation principle; and the target AIGC video generation module 60 is configured to generate a target AIGC video with the second target requirement and the first most likely estimated duration as constraints.

[0071] Next, the specific configuration of the feature analysis module 20 will be described in detail. As described above, the feature analysis module 20 can further include a predetermined video feature construction unit configured to construct predetermined video features, wherein the predetermined video features include video scene features, video style features, and video element features.

[0072] In the following, the specific configuration of the target AIGC video generation module 60 will be described in detail. As described above, the target AIGC video generation module 60 can further include a first label extraction unit configured to extract a first label in the second target demand, the first label having an identification of a first selection order and a first selection time; a second label extraction unit configured to extract a second label in the second target demand, the second label having an identification of a second selection order and a second selection time, and the second selection order being adjacent to the first selection order; a second consideration time length calculation unit configured to obtain a second consideration time length of the second label by comparing the first selection time and the second selection time; a second weight determination unit configured to determine a second weight of the second label based on the second selection order and the second consideration time length, and record the second weight as a second proportion; a second target material set obtaining unit configured to read a label material database and obtain a second target material set of the second label based on the second proportion; and a target AIGC video generation unit configured to generate the target AIGC video based on the second target material set and the first most probable estimated time length.

[0073] In the above, the second target material set obtaining unit can further include a second material set screening sub-unit configured to screen a second material set corresponding to the second label in the label material database; and a material extraction sub-unit configured to randomly extract materials in the second proportion in the second material set to form the second target material set.

[0074] In the following, the specific configuration of the time length estimation module 50 will be described in detail. As described above, the time length estimation module 50 can further include a first most probable estimated time length determination unit configured to take the mean value of the first demand time length and the second demand time length in the time length demand information as the first most probable estimated time length.

[0075] In the above, after generating the target AIGC video based on the second target demand and the first most probable estimated time length, the system can further include a target time length modification demand obtaining module configured to obtain a target time length modification demand of the target user for the target AIGC video; a first pessimistic estimated time length obtaining module configured to, when the target time length modification demand is an extended time length, analyze the time length demand information to obtain a first pessimistic estimated time length; and a video expansion module configured to expand the target AIGC video based on the first pessimistic estimated time length.

[0076] In the following, the specific configuration of the video expansion module will be described in detail. As described above, the target AIGC video is expanded with the first pessimistic estimated length as a constraint, the video expansion module can further include: a target length modification ratio obtaining unit configured to obtain a target length modification ratio by comparing the first most likely estimated length with the first pessimistic estimated length; a second ratio adjusting unit configured to adjust the second ratio based on the target length modification ratio to obtain a second target material adjustment set; and a target material set assembling unit configured to replace the second target material set with the second target material adjustment set to assemble the target material set.

[0077] In the following, the specific configuration of the video expansion module will be described in detail. As described above, the target AIGC video is expanded with the first pessimistic estimated length as a constraint, the video expansion module can further include: a target length modification ratio obtaining unit configured to obtain a target length modification ratio by comparing the first most likely estimated length with the first pessimistic estimated length; a second ratio adjusting unit configured to adjust the second ratio based on the target length modification ratio to obtain a second target material adjustment set; and a target material set assembling unit configured to replace the second target material set with the second target material adjustment set to assemble the target material set.

[0078] The demand-oriented AIGC video generation system provided by the embodiments of the present application can execute the demand-oriented AIGC video generation method provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0079] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, however, any number of different modules can be used and run on the user terminal and / or server, and each unit and module included is only divided according to the functional logic, but is not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy mutual differentiation, and do not limit the protection scope of the present application.

[0080] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application. In some cases, the actions or steps described in the present application can be executed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.

Claims

1. A demand-oriented AIGC video generation method, characterized in that: include: Acquire a first target demand of a target user, wherein the first target demand includes multiple demand modal information; Performing feature analysis on the predetermined video features and each demand modality information in the first target demand to generate a first feature label set; Performing a union operation on the first feature label set to obtain a first target label list; A tag selection pool is formed based on the first target tag list, and a demand selection record of the target user for the tag selection pool is obtained and recorded as a second target demand; Extracting the duration requirement information from the first target requirement and obtaining a first most likely estimated duration using a three-point estimation principle; generating a target AIGC video based on the second target requirement and the first most likely estimated duration as constraints; The step of generating a target AIGC video based on the second target requirement and the first most likely estimated duration as constraints includes: Extracting a first tag from the second target requirement, where the first tag has identifiers of a first selection order and a first selection time; Extracting a second tag from the second target requirement, where the second tag has an identifier of a second selection order and a second selection time, and the second selection order is adjacent to the first selection order; Comparing the first selection time with the second selection time to obtain a second consideration time for the second tag; Determine a second weight of the second tag based on the second selection order and the second consideration time, and record it as a second ratio; Reading a label material database, and obtaining a second target material set of the second label in combination with the second ratio; A target material set is constructed based on the second target material set, and the target AIGC video is generated in combination with the first most likely estimated duration.

2. The demand-oriented AIGC video generation method according to claim 1, characterized in that: The predetermined video features include video scene features, video style features and video element features.

3. The demand-oriented AIGC video generation method according to claim 1, characterized in that: Reading the label material database and obtaining the second target material set of the second label in combination with the second ratio includes: Filtering the second material set corresponding to the second label in the label material database; The second proportion of materials in the second material set is randomly extracted to form the second target material set.

4. The demand-oriented AIGC video generation method according to claim 1, characterized in that: An average of the first required duration and the second required duration in the duration requirement information is taken as the first most likely estimated duration.

5. The demand-oriented AIGC video generation method according to claim 1, characterized in that: After generating a target AIGC video based on the second target requirement and the first most likely estimated duration as constraints, the method further includes: Obtaining the target user's requirement for modifying the target duration of the target AIGC video; When the target duration modification requirement is to extend the duration, analyzing the duration requirement information to obtain a first pessimistic estimated duration; The target AIGC video is expanded using the first pessimistic estimated duration as a constraint.

6. The demand-oriented AIGC video generation method according to claim 5, characterized in that: The target AIGC video is extended based on the first pessimistic estimated duration as a constraint, including: Comparing the first most likely estimated duration with the first pessimistic estimated duration to obtain a target duration modification ratio; Adjusting the second ratio based on the target duration modification ratio to obtain a second target material adjustment set; The second target material set is replaced with the second target material adjustment set to form the target material set.

7. The demand-oriented AIGC video generation method according to claim 5, characterized in that: After obtaining the target user's requirement for modifying the target duration of the target AIGC video, the method further includes: When the target duration modification requirement is to reduce the duration, analyzing the duration requirement information to obtain a first optimistic estimated duration; The target AIGC video is compressed using the first optimistic estimated duration as a constraint.

8. The demand-oriented AIGC video generation system is characterized by: The system is used to implement the demand-oriented AIGC video generation method according to any one of claims 1 to 7, and the system includes: A first target demand acquisition module is used to acquire a first target demand of a target user, wherein the first target demand includes a plurality of demand modal information; A feature analysis module, configured to perform feature analysis on predetermined video features and each demand modality information in the first target demand to generate a first feature tag set; a union operation analysis module, configured to perform a union operation analysis on the first feature label set to obtain a first target label list; A second target demand acquisition module is configured to construct a tag selection pool based on the first target tag list, and obtain the target user's demand selection record for the tag selection pool as a second target demand; a duration estimation module, configured to extract duration requirement information from the first target requirement and obtain a first most likely estimated duration using a three-point estimation principle; The target AIGC video generation module is configured to generate a target AIGC video based on the second target requirement and the first most likely estimated duration as constraints.

Citation Information

Patent Citations

  • Video generation method, device, system, equipment and medium

    CN118869906A

  • Video generation method and device, electronic equipment and storage medium

    CN119474454A