Digital creativity cultivation method and system based on image video analysis

By constructing a structured contradiction map and an intent clarification inquiry mechanism, the problem of creators being unable to clearly express their creative ideas when they are experiencing a lack of inspiration or are in the early stages of learning is solved, generating creative materials that better meet the needs of creators and improving the efficiency and quality of digital creative cultivation.

CN121880583APending Publication Date: 2026-04-17SHANDONG NEW NONGSHANG SOFTWARE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG NEW NONGSHANG SOFTWARE TECHNOLOGY CO LTD
Filing Date
2026-01-08
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In the field of digital art and content creation, creators often struggle to clearly express their creative ideas when they are creatively uninspired or in the early stages of learning. This makes it difficult for the system to accurately capture their true intentions, resulting in disappointing creative materials that further disorient creators.

Method used

By using a digital creative incubation method based on image and video analysis, reference media content is received, visual information is extracted, a structured contradiction map is constructed, an intent clarification inquiry mechanism is initiated, the contradiction map is dynamically adjusted, creative materials are generated, and the interaction behavior between creators and creative materials is continuously analyzed to iteratively optimize the generation strategy.

Benefits of technology

It effectively solves the problems of ambiguous creator intentions and contradictory input content, generates creative materials that better meet the needs of creators, improves the efficiency of early concept design, helps creators overcome inspirational obstacles, and avoids logical confusion that increases bewilderment.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention discloses a digital creative cultivation method and system based on image video analysis, and the method comprises the following steps: a visual extraction step: receiving reference media content, and carrying out the content analysis of the reference media content, so as to extract visual information; a contradictory graph construction step: constructing a structured contradictory graph representing conflicts between creative directions in the reference media content based on the extracted visual information; a contradiction map adjusting step: based on the contradiction map, starting an intention clarification inquiry mechanism, outputting a guiding problem, receiving feedback of guiding inquiry, and dynamically adjusting the contradiction map according to the feedback to converge the creation intention; and an iterative optimization step: generating a creative material based on the converged creative intention, and continuously analyzing an interactive behavior between the creator and the creative material so as to iteratively optimize a generation strategy. According to the method, complex image video analysis, contradiction map construction and adjustment and creative material generation and optimization processes can be carried out, and the complexity and workload of manual operation are remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital creative cultivation, and more specifically, to a method and system for digital creative cultivation based on image and video analysis. Background Technology

[0002] In the fields of digital art and content creation, creators often utilize auxiliary tools to inspire their work. By submitting reference images or video clips, the system analyzes the visual information and generates structured creative materials, such as color schemes, composition templates, or animation drafts, to improve the efficiency of early-stage concept design. However, this approach requires creators to have a clear creative intent and a high level of aesthetic judgment, and to be able to provide reference materials with a distinct theme and consistent style.

[0003] When creators are experiencing creative blocks or are in the early stages of learning, they often struggle to clearly express their creative ideas. The reference materials they submit may be disorganized, with significant differences or even contradictions in visual style, theme, and emotional expression, making it difficult for the system to accurately capture their true intentions. For example, a student designing a game character might simultaneously upload sci-fi mecha concept art, screenshots of classical oil paintings, and clips from cartoons. The conflict between these materials results in the system obtaining chaotic and contradictory data when extracting visual information.

[0004] Because the input materials lack inherent consistency, the creative materials generated by the system based on this chaotic information are often disappointing, potentially recommending incongruous color schemes or compositional drafts that stitch together contradictory elements. These outputs not only fail to provide effective inspiration but also increase the creator's confusion due to their logical inconsistencies, making them even more lost in their creative direction. Summary of the Invention

[0005] This invention discloses a digital creative cultivation method based on image and video analysis, which aims to solve the problem in the existing field of digital art and content creation where creators are unable to clearly express their creative ideas when they are experiencing a lack of inspiration or are in the early stages of learning. This makes it difficult for the system to accurately capture their true intentions, resulting in the generation of disappointing creative materials and causing creators to lose their creative direction.

[0006] The technical solution of the present invention is as follows: In a first aspect, the present invention discloses a digital creative cultivation method based on image and video analysis, comprising the following steps: The visual extraction step involves receiving reference media content and parsing the content to extract visual information. The steps for constructing the contradiction map are as follows: Based on the extracted visual information, a structured contradiction map representing the conflict between creative directions in the reference media content is constructed. The steps for adjusting the contradiction map are as follows: Based on the contradiction map, an intent clarification inquiry mechanism is initiated, guiding questions are output and feedback on the guiding questions is received, and the contradiction map is dynamically adjusted according to the feedback in order to converge the creative intent. The iterative optimization process generates creative materials based on the converged creative intent and continuously analyzes the interaction between the creator and the creative materials to iteratively optimize the generation strategy.

[0007] Through this technical solution, the present invention can proactively explore the creator's true intention. Even when faced with ambiguous and contradictory input, it can gradually converge the creative intention by constructing and adjusting the contradiction map, thereby generating creative materials that better meet the creator's needs. This effectively solves the problem of the system failing as a mere visual element "calculator" in the prior art.

[0008] Secondly, the present invention also discloses a digital creative incubation system for performing any of the above methods.

[0009] Through this technical solution, the present invention provides a system that can be practically deployed and operated, transforming the above-mentioned methods into an operable tool, and providing an integrated solution for the cultivation of digital creativity.

[0010] Beneficial effects This invention discloses a digital creative cultivation method based on image and video analysis. By receiving reference media content and extracting visual information, a structured contradiction map representing conflicts between creative directions is constructed. The method further activates an intent clarification inquiry mechanism to dynamically adjust the contradiction map based on creator feedback, thereby converging creative intentions. Finally, creative materials are generated based on the converged creative intentions, and the interaction between creators and creative materials is continuously analyzed to iteratively optimize the generation strategy.

[0011] This method effectively addresses the problem in existing technologies where creators, when experiencing creative blocks or in their early stages, struggle to clearly express their ideas, leading to the system's inability to accurately capture their true intentions and generating disappointing creative materials that further disorient the creator. By introducing a contradiction map and an intent-clarification inquiry mechanism, this invention proactively explores what the creator "wants," rather than passively analyzing "what is provided." Even when faced with disorganized reference materials that differ significantly or even contradict each other in visual style, theme, and emotional expression, this invention can clearly present the conflicts between creative directions through a structured contradiction map. Through interaction with the creator, it gradually guides them to clarify and refine their creative intentions. This proactive guidance and iterative optimization mechanism transforms the system from a mere visual element "calculator" into an intelligent partner that explores and grows alongside the creator, generating creative materials that better meet their actual needs. This effectively improves the efficiency of early-stage concept design, helps creators overcome inspirational obstacles, and avoids confusion caused by logical inconsistencies. Detailed Implementation

[0012] The following will provide a clear and complete description of the technical solutions. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0013] To address this, the present invention proposes a digital creativity cultivation method based on image and video analysis, comprising the following steps: The visual extraction step involves receiving reference media content and parsing the reference media content to extract visual information. The contradiction map construction step involves constructing a structured contradiction map representing the conflict between creative directions in the reference media content based on the extracted visual information. The steps for adjusting the contradiction map are as follows: Based on the contradiction map, an intent clarification inquiry mechanism is initiated, guiding questions are output and feedback on the guiding questions is received, and the contradiction map is dynamically adjusted according to the feedback to converge the creative intent. The iterative optimization step involves generating creative materials based on the converged creative intent and continuously analyzing the interaction behavior between the creator and the creative materials to iteratively optimize the generation strategy.

[0014] This invention proactively identifies and resolves contradictions in creator input, and combines an intent clarification inquiry mechanism and iterative optimization strategy to effectively address the problems of ambiguous creator intent and contradictory input content. This allows for a more accurate understanding of the creator's deeper needs and the generation of more inspiring and targeted creative materials, significantly improving the efficiency and quality of digital creative cultivation.

[0015] In the visual extraction step, the system receives reference media content provided by the creator and parses it to extract visual information. For example, a creator can upload a collage containing elements of various styles as reference media content. The system can extract visual information in several ways. One approach is to utilize traditional image processing techniques, such as edge detection, color histogram analysis, and texture analysis, to extract low-level visual features from the image. For video content, this analysis can be performed frame-by-frame, combined with techniques such as motion estimation to extract dynamic visual information. Another approach is to employ deep learning-based methods, such as using a pre-trained convolutional neural network model, to extract features from the reference media content, thereby obtaining higher-level semantic visual information, such as object recognition, scene classification, and art style recognition. This extracted visual information will serve as the basis for subsequent steps.

[0016] In the contradiction mapping construction step, based on the extracted visual information, the system constructs a structured contradiction map representing the conflict between creative directions in the reference media content. In the iterative optimization step, creative materials are generated based on the converged creative intent, and the interaction between the creator and the creative materials is continuously analyzed to iteratively optimize the generation strategy.

[0017] The digital creative cultivation method based on image and video analysis proposed in this invention significantly improves the intelligence level of digital creative assistance tools by introducing proactive contradiction identification and intent clarification mechanisms. Traditional methods often passively analyze reference media content provided by creators, and the system struggles to provide effective assistance when the content contains contradictions or the creator's intent is ambiguous. The core innovation of this invention lies in its ability not only to extract visual information but also to further construct a structured contradiction map, making the potential conflicts in the creator's input explicit.

[0018] Furthermore, the introduction of iterative optimization steps enables the method of this invention to continuously learn and improve. By analyzing the interaction between creators and generated creative materials, the system can continuously optimize its generation strategy, ensuring that subsequently generated creative materials more accurately meet the needs of creators. For example, if creators show a higher acceptance of materials in a certain style, the system will increase the weight of materials in that style in subsequent generation. This closed-loop optimization mechanism allows the method of this invention to adapt to the ever-changing creative needs of creators and provide personalized creative nurturing services.

[0019] In summary, this invention establishes a complete and efficient digital creative incubation process through four core steps: visual extraction, contradiction map construction, contradiction map adjustment, and iterative optimization. It not only overcomes the limitations of traditional methods in handling ambiguous and contradictory inputs but also provides creators with a more intelligent and personalized creative support tool through proactive interaction and continuous optimization, significantly improving the efficiency and quality of digital creative incubation.

[0020] The present invention further proposes that the above-mentioned visual extraction steps include: Semantic sub-dimensional analysis is performed on the art styles in the extracted visual information to identify semantic sub-dimensional differences within the art styles; Cross-cognitive domain tension analysis is performed on the thematic concepts in the extracted visual information to identify cross-cognitive domain tensions between the thematic concepts; Monitor the sequence of reference media content uploaded iteratively by creators, and analyze the changing trends of visual information between different batches of reference media content in the sequence in order to identify the drift of creators' creative intentions; Based on the semantic sub-dimension differences, the cross-cognitive domain tensions, and the drift of creative intentions, a multi-dimensional contradiction map is constructed. The multi-dimensional contradiction map represents the macro-style and theme conflicts, the semantic sub-dimension differences, the cross-cognitive domain theme tensions, and the drift of creative intentions.

[0021] Specifically, semantic sub-dimensional analysis of artistic styles extracted from visual information refers to further refining a macro-level artistic style (such as "Impressionism" or "Cyberpunk") into multiple quantifiable semantic sub-dimensions. For example, Impressionism can be broken down into sub-dimensions such as color brightness, brushstroke thickness, and light and shadow treatment; Cyberpunk can be broken down into sub-dimensions such as neon color saturation, mechanical texture complexity, and density of futuristic elements. By analyzing these sub-dimensions, subtle differences or potential conflicts within the same artistic style can be identified. For example, in an Impressionist work, the color brightness is very high, but the brushstrokes are exceptionally delicate, which may constitute an internal contradiction within the style.

[0022] Therefore, based on the aforementioned semantic sub-dimension differences, cross-cognitive domain tensions, and the drift of creative intentions, a richer and more refined multi-dimensional contradiction map can be constructed. This multi-dimensional contradiction map not only represents macro-level style and theme conflicts but also further incorporates semantic sub-dimension differences within artistic styles, cross-cognitive domain theme tensions between thematic concepts, and the dynamic drift of the creator's creative intentions. Its purpose is to provide a more comprehensive and in-depth view of creative conflict, offering a more precise basis for subsequent creative guidance and material generation.

[0023] This invention significantly enhances the depth and dynamism of the visual extraction process by introducing semantic sub-dimensional analysis, cross-cognitive domain tension analysis, and creative intent drift monitoring. Traditional visual extraction methods may only identify macroscopic visual elements, while this invention delves into the internal structure of artistic styles and the deep semantic connections of thematic concepts. It is precisely this meticulous analysis of the semantic sub-dimensions of artistic styles that allows the system to capture subtle conflicts and innovative potential within the same style, avoiding simplistic style categorization. Simultaneously, through cross-cognitive domain tension analysis of thematic concepts, the system can identify seemingly unrelated combinations that actually contain deep creative tension, thereby broadening the boundaries of creative exploration. Furthermore, by continuously monitoring the sequence of reference media content uploaded by creators, the system can dynamically perceive the evolution and drift of creators' intentions. This makes the construction of the contradiction map no longer static, but rather reflects the creator's latest direction of thought in real time. This multi-dimensional information is integrated into a multi-dimensional conflict map, which enables the map to more comprehensively and accurately represent the conflicts between creative directions, thus providing a more solid and dynamic foundation for subsequent conflict map adjustments and creative material generation.

[0024] Through the above technical solutions, this invention can significantly improve the accuracy and adaptability of digital creative incubation. Compared to solutions that only extract basic visual information, this invention, by deeply analyzing the semantic sub-dimension differences of artistic styles and the cross-cognitive tension of thematic concepts, can identify deeper and more inspiring points of creative conflict, thus providing creators with more insightful guidance. Furthermore, real-time monitoring of the creator's shifting creative intentions allows the system to dynamically adjust its understanding of the creator's intent, ensuring that the provided creative guidance and material generation always align with the creator's latest thinking, avoiding guidance deviations caused by changes in intent. Therefore, the constructed multi-dimensional contradiction map can more comprehensively and accurately reflect the creator's true needs and the potential tension of creative content, greatly improving the effectiveness of the creative incubation process and the innovativeness of the final creative outcome.

[0025] Specifically, in the above visual extraction steps, the step of performing semantic sub-dimension analysis on the artistic style in the extracted visual information to identify the semantic sub-dimension differences within the artistic style can be further refined.

[0026] The steps for identifying semantic sub-dimension differences within the art style include: For each art style in the extracted visual information, identify its multiple semantic sub-dimensions, and define a set of visual feature descriptors for each semantic sub-dimension. The visual feature descriptors include color saturation, texture roughness, line smoothness, light and shadow contrast, and spatial layout density. For each part of the reference media content, calculate the matching degree of its visual feature descriptor on each semantic sub-dimension to obtain the contribution of that part to each semantic sub-dimension. For semantic sub-dimensions within the same art style, when the matching degree of their visual feature descriptors overlaps, the differences in the semantic sub-dimensions are quantified by comparing the contribution degree and combining it with the pixel proportion or frequency of occurrence of the corresponding part in the reference media content.

[0027] Specifically, when processing the extracted visual information, it is first necessary to conduct an in-depth analysis of each art style to identify the multiple semantic sub-dimensions it contains. For example, an "Impressionist" art style may include semantic sub-dimensions such as "light and shadow capture," "brushstroke expression," and "color usage." To ensure accurate quantification of these sub-dimensions, a specific set of visual feature descriptors is defined for each semantic sub-dimension. These descriptors are quantifiable indicators; for example, color saturation measures the vividness of colors, texture roughness describes the quality of details in the image, line smoothness reflects the fluidity of lines, light and shadow contrast focuses on the difference between light and dark areas, and spatial layout density assesses the compactness of the distribution of elements in the image. These descriptors together constitute a multi-dimensional feature space used to accurately capture and characterize the subtle differences in art styles.

[0028] Furthermore, for each part of the reference media content, such as a region in an image or a frame in a video, the matching degree of its visual feature descriptors in each semantic sub-dimension is calculated. This matching degree reflects the extent to which the part conforms to the visual feature descriptors defined by a specific semantic sub-dimension. Through this calculation, the contribution of the part to each semantic sub-dimension can be obtained, that is, the extent to which the part embodies the features of a certain semantic sub-dimension.

[0029] This invention employs a detailed semantic sub-dimension analysis of artistic styles and defines a set of quantified visual feature descriptors for each sub-dimension, thereby accurately capturing the intrinsic composition of artistic styles from multiple dimensions. By calculating the contribution of each part of the reference media content to these semantic sub-dimensions and combining this with contextual information such as pixel proportion or frequency of occurrence, even when visual feature descriptors overlap, it can effectively decouple and quantify the differences in semantic sub-dimensions within the same artistic style. This method avoids the crude analysis of artistic styles as a single whole, instead delving into the level of their constituent elements, thus providing more refined and accurate foundational data for subsequent identification of creative direction conflicts and construction of contradiction maps.

[0030] The aforementioned technical solutions enable refined identification and quantification of semantic sub-dimension differences within artistic styles. This allows the system not only to identify macro-level artistic styles but also to deeply understand subtle variations and potential conflicts within the same style. For example, within the "minimalist" style, it can distinguish between "geometric minimalism" and "natural minimalism." This meticulous analysis helps to more accurately capture the creator's underlying intentions and identify potential contradictions in creative direction, thus providing richer and more precise input for constructing a multi-dimensional contradiction map, significantly enhancing the depth and effectiveness of digital creative cultivation.

[0031] In some of the above embodiments, cross-cognitive domain tension analysis is performed on the thematic concepts in the extracted visual information to identify the cross-cognitive domain tensions between the thematic concepts. Specifically, the steps for identifying the cross-cognitive domain tensions between thematic concepts include: Identify the set of concrete visual symbols associated with the theme concept, the set of visual symbols including objects, scenes, colors, light and shadow and compositional elements directly associated with the theme concept; The presence of visual symbols from the set of visual symbols in the reference media content is analyzed, and the degree of concretization of the theme concept is evaluated based on the frequency of occurrence, visual salience, and contextual relevance of the visual symbols in the reference media content. For thematic concepts with low figurative degree, identify the metaphorical visual cues that may exist in the reference media content. The metaphorical visual cues include visual elements that are not directly related but can evoke associations, emotional expressions, and narrative hints. Based on the degree of figuration and the metaphorical visual cues, the abstract semantics of the topic concept are inferred, and the semantic distance between the abstract semantics is calculated; Based on the semantic distance and the degree of concretization of the topic concepts, the cross-cognitive domain tension between the topic concepts is quantified.

[0032] Furthermore, for thematic concepts with a low degree of concreteness, they may be expressed through metaphorical visual cues. These metaphorical visual cues refer to visual elements that do not directly depict the thematic concept but convey its deeper meaning through association, symbolism, or suggestion. For example, when depicting the theme of "loneliness," the word or symbol of "loneliness" may not appear directly. Instead, it may evoke feelings of loneliness in the viewer through non-directly related visual elements such as empty scenes, a single figure, and cool-toned lighting. Emotional expression and narrative suggestion are also important metaphorical visual cues. They construct a specific emotional atmosphere or story background through the combination and arrangement of visual elements, thereby indirectly conveying abstract thematic concepts.

[0033] This invention identifies and analyzes a set of concrete visual symbols representing a theme concept, enabling accurate assessment of the explicitness of that concept within reference media content. By further identifying metaphorical visual cues of theme concepts with lower levels of concreteness, this invention captures indirect, deeper creative expressions, thus comprehensively understanding the semantic connotations of the theme concept. By combining the degree of concreteness and metaphorical visual cues to infer abstract semantics and calculate semantic distance, this invention quantifies the inherent connections and conflicts between different theme concepts. Ultimately, by quantifying cross-cognitive domain tension based on semantic distance and the degree of concreteness, the system can more accurately identify potential, multi-dimensional conflicts between creative directions, providing more refined input for subsequent construction and adjustment of the contradiction map.

[0034] Through the above technical solution, this invention can more comprehensively and deeply understand the expression of thematic concepts in reference media content, not only limited to explicit visual symbols, but also capturing implicit, deep-seated semantic cues. This enables the system to identify potential tensions existing between different cognitive domains that are difficult to discover using traditional methods, thereby more accurately quantifying conflicts between creative directions. This refined tension identification helps to construct a more insightful map of contradictions, providing creators with more inspiring creative directions, effectively avoiding the creative limitations caused by focusing only on surface visual elements, and significantly improving the depth and breadth of digital creative cultivation.

[0035] In some embodiments of the present invention described above, the drift of the creator's creative intent is identified by monitoring the sequence of reference media content uploaded iteratively by the creator and analyzing the changing trends of visual information between different batches of reference media content in the sequence. However, in practical applications, relying solely on the changing trends of visual information to judge the drift of creative intent may have limitations. Creators may engage in exploratory attempts, minor adjustments, or accidental operations during the creative process; the visual changes resulting from these actions may not be a fundamental drift of the true intent. If this distinction is not made, the system may misjudge the creator's intent, leading to generated creative materials that do not match the creator's actual needs, thus affecting the efficiency and accuracy of creative development.

[0036] In response, this invention further proposes a method for identifying drift in a creator's creative intent, the steps of which include: Visual features are extracted from the reference media content uploaded in each batch of the sequence to obtain a set of visual features for each batch. Calculate the visual difference between each batch of visual feature sets and the previous batch of visual feature sets; The visual differences are semantically calibrated by combining the text descriptions or tags that the creators may have attached when uploading this batch of materials; For batches where the visual differences exhibit weak or non-linear changes, a time decay weighting mechanism is introduced, assigning lower weights to the visual feature sets of earlier batches and higher weights to the visual feature sets of the latest batches. By considering whether the creator viewed, modified, or deleted the materials shortly after uploading them, we can determine whether the visual changes in this batch were an exploratory attempt or a shift in the true intent. Based on the judgment result, update the direction and intensity of the creative intention drift.

[0037] This invention effectively addresses the limitations of relying solely on visual information trends to identify creative intent drift by introducing a multi-dimensional and refined analysis mechanism. First, by extracting visual features and calculating visual differences for each batch of reference media content, a quantitative foundation is laid for subsequent intent analysis. Second, semantic calibration of visual differences is performed by incorporating text descriptions or tags that may accompany the upload, enabling the system to comprehensively consider both visual and semantic information, avoiding misjudgments caused by purely visual changes, and thus more accurately understanding the creator's intent. Third, a time decay weighting mechanism is introduced for batches with weak or non-linear changes, allowing the system to focus more on the creator's latest creative trends, effectively filtering out early or unimportant exploratory content, ensuring the timeliness and relevance of intent drift judgments. Finally, by analyzing whether the creator viewed, modified, or deleted materials within a short period after uploading, the system can distinguish between exploratory attempts and genuine intent drift. It is precisely because of these multi-layered analyses that the system is able to capture the dynamic changes of the creator's intentions more comprehensively and accurately, avoiding misjudging exploratory behavior as intention drift, thus providing a more reliable basis for subsequent adjustment of the contradiction map and generation of creative materials.

[0038] Through the above technical solutions, this invention can significantly improve the accuracy and robustness of identifying creators' drifting creative intentions. Specifically, by introducing semantic calibration, the system can integrate visual and textual information, avoiding misjudgments that may result from relying solely on visual information, thus enabling a more comprehensive and in-depth understanding of the creator's intentions. Furthermore, the introduction of a time decay weighting mechanism allows the system to dynamically adjust its focus on different batches of content, ensuring more timely and effective capture of the latest creative intentions, thereby improving the real-time performance of intention drift identification. More importantly, by analyzing the creator's subsequent interactive behavior, this invention can effectively distinguish between exploratory attempts and genuine intention drift, avoiding misjudging the creator's tentative operations as a fundamental change in intention. This reduces deviations in creative material generation caused by misjudgments, making the digital creative incubation process more accurately respond to the creator's real needs, ultimately improving the efficiency of creative incubation and user satisfaction.

[0039] In some embodiments of the present invention described above, when identifying the drift of a creator's creative intent, semantic calibration of visual differences is performed by combining the text descriptions or tags that the creator may have attached when uploading the batch of materials. However, in practical applications, the text descriptions or tags provided by the creator may semantically conflict with the visual content, or the information may be ambiguous or incomplete, affecting the accuracy of semantic calibration and potentially leading to misjudgments of the drift direction and intensity of the creator's true intent. If these problems are not addressed, subsequently generated creative materials may not match the creator's actual intent, reducing the efficiency of creative development and user satisfaction.

[0040] In response, this invention further proposes a step for semantic calibration of the aforementioned visual dissimilarity, which specifically includes: Keyword extraction and sentiment analysis are performed on the text descriptions or tags to obtain a set of semantic features; Visual saliency analysis and topic concept recognition are performed on the visual feature set to obtain a visual semantic set; Identify semantic elements that conflict with the set of semantic features and the set of visual semantics, and calculate the conflict weight based on the number and intensity of the conflicting elements; Identify ambiguous or missing semantic elements between the semantic feature set and the visual semantic set, and calculate uncertainty weights based on the number and importance of the ambiguous or missing elements; The calibration intensity of the text description or label for the visual difference is dynamically adjusted based on the conflict weight and the uncertainty weight.

[0041] Specifically, keyword extraction and sentiment analysis are performed on the aforementioned text descriptions or tags to extract core concepts and potential emotional tendencies from unstructured text information, thereby forming a structured set of semantic features. For example, Natural Language Processing (NLP) techniques, such as TF-IDF and TextRank algorithms, can be used for keyword extraction, and sentiment dictionaries or deep learning models can be used for sentiment polarity judgment. Visual saliency analysis and topic concept recognition on the aforementioned visual feature set can be understood as identifying the most eye-catching areas (visual saliency) in an image or video at the pixel level using computer vision techniques, and the specific objects, scenes, or abstract concepts (topic concepts) represented by these areas, thus obtaining a visual semantic set. For example, deep learning models (such as YOLO and Mask R-CNN) can be used for object detection and scene recognition, combined with attention mechanisms to evaluate visual saliency.

[0042] This invention effectively identifies potential semantic conflicts and information gaps by comparing and analyzing the semantic feature set of textual descriptions or tags with the visual semantic set of visual feature sets. Specifically, keyword extraction and sentiment analysis can extract the creator's initial intentions and emotional cues from the text; visual saliency analysis and topic concept recognition can extract objective visual semantics from image or video content. When these two semantic sets are identified as conflicting—for example, if the textual description and the image content express inconsistent themes or emotions—the system calculates a conflict weight, thereby reducing the influence of textual information in calibrating visual differences and avoiding erroneous intent drift judgments due to textual misleading. Simultaneously, when textual information is ambiguous or lacks information, the system calculates an uncertainty weight, indicating that the textual information's calibration of visual differences may be insufficient or inaccurate, prompting the system to rely more on the visual information itself or seek further clarification during the calibration process. This dynamic adjustment mechanism makes the semantic calibration process more intelligent and robust, more accurately reflecting the creator's true intentions.

[0043] Through the above technical solution, this invention effectively solves the problem of inaccurate calibration caused by inconsistencies, ambiguities, or missing information between the text descriptions or tags provided by creators and the visual content in traditional semantic calibration methods. By quantifying semantic conflicts and uncertainties, the system can dynamically adjust the calibration intensity of the text description on the visual differences, thereby avoiding misjudgments of the creator's intent. This makes the identification of creator intent drift more accurate, ensuring that the subsequently generated creative materials can more accurately match the creator's actual needs, significantly improving the efficiency of digital creative incubation and user experience.

[0044] In some embodiments of the present invention described above, when identifying the drift of a creator's creative intent, a time decay weighting mechanism is introduced for batches exhibiting weak or non-linear changes in visual differences. Lower weights are assigned to the visual feature sets of earlier batches, and higher weights are assigned to the visual feature sets of the most recent batches. However, in specific scenarios where creators frequently upload sequences of reference media content within a short period, relying solely on time decay weighting may not accurately capture the subtle evolution of creative intent. This is especially true when visual changes are not significant but the semantic direction has shifted, potentially leading to insufficient sensitivity or accuracy in judging the drift of creative intent.

[0045] To address this, the present invention further proposes a digital creative cultivation method based on image and video analysis. This method involves: extracting visual features from the reference media content uploaded in each batch of the sequence to obtain a visual feature set for each batch; calculating the visual difference between each batch's visual feature set and the previous batch's visual feature set; semantically calibrating the visual difference by incorporating any accompanying text descriptions or tags that the creator may have attached when uploading the batch of materials; identifying and extracting the core semantic elements of each batch from batches uploaded frequently within a short period; calculating the semantic evolution distance between the core semantic elements of adjacent batches; increasing the weight of the latest batch when the semantic evolution distance exceeds a preset threshold; otherwise, decreasing the weight of the latest batch; determining whether the visual changes in the batch are exploratory attempts or a genuine intention shift by considering whether the creator viewed, modified, or deleted the material within a short period after uploading it; and updating the direction and intensity of the creative intention shift based on the determination result.

[0046] This invention identifies and extracts the core semantic elements of each batch from frequent uploads within a short period, and calculates the semantic evolution distance between the core semantic elements of adjacent batches. This allows for a deeper understanding of the creator's true intentions during rapid iteration. Traditional time-decay-based weighting mechanisms may only focus on the timeliness of input, ignoring the semantic evolution of the content itself. By introducing semantic evolution distance as the basis for weight adjustment, this invention can directly assess the continuity or leaps in the creator's intentions at the semantic level. When the semantic evolution distance is large, it indicates that the creator's intentions may have undergone a substantial change. In this case, increasing the weight of the latest batch allows the system to respond to this change more quickly. Conversely, when the semantic evolution distance is small, even with frequent uploads, it may only be a minor adjustment or exploration. In this case, decreasing the weight of the latest batch prevents the system from overreacting to non-substantial changes. Thus, this solution can more accurately capture the drift of the creator's intentions, especially in scenarios where visual differences are not obvious but semantic directions have shifted.

[0047] Through the above technical solution, this invention can more accurately and sensitively identify the drift of a creator's creative intent, especially when the creator is engaged in rapid and frequent iterations. This solution avoids misjudgments that may result from relying solely on visual differences or simple time decay, and can distinguish between exploratory attempts and genuine shifts in intent. By analyzing core semantic elements and quantifying the distance of semantic evolution, the system can gain a deeper understanding of the creator's creative thought process, thereby providing more precise creative nurturing guidance and material generation, significantly improving the efficiency of digital creative nurturing and the user experience.

[0048] Specifically, the steps for determining whether the visual changes in this batch are exploratory attempts or a true intention drift include: After the creator uploads the batch of materials, behavioral sequence analysis is performed on the creator's browsing time, number of modifications, types of modified content, and deletion operations within a preset time window; Identify whether there are contradictory or inconsistent operation patterns in the sequence of behaviors. The contradictory or inconsistent operation patterns include repeated modification of the same visual element, modification followed by deletion, or deletion followed by re-uploading. Based on the contradictory or inconsistent operating patterns, and in conjunction with the visual difference, the degree of hesitation of the creator's intention is calculated; Based on the degree of hesitation, determine whether the visual changes in the batch are due to hesitation or repetition of true intentions, or to random fluctuations in system interaction or user habits.

[0049] The preset time window refers to a specific time range during which the system continuously monitors the creator's subsequent actions after the creator uploads the reference media content. The length of this time window can be flexibly configured according to the actual application scenario and user behavior patterns; for example, it can be set to several minutes, hours, or days. Its purpose is to capture immediate feedback and subsequent actions closely related to the upload. Behavioral sequence analysis specifically refers to the recording and analysis of the time sequence and type of a series of operations (such as browsing, modifying, and deleting) performed by the creator on the uploaded material within the preset time window. Browsing duration can reflect the creator's level of attention to the content; the number of modifications and the type of modified content can reveal the creator's adjustments to details and exploration of direction; deletion operations may indicate the creator's dissatisfaction with the current solution or abandonment. Through the serialization analysis of these discrete behaviors, the creator's deeper intentions and decision-making processes can be revealed.

[0050] Furthermore, contradictory or inconsistent operational patterns refer to specific behavioral patterns identified in behavioral sequence analysis that indicate uncertainty or repetition in the creator's intentions. For example, repeatedly modifying the same visual element may indicate that the creator is wavering between multiple possibilities; modifying and then deleting may mean that the creator tried a certain direction but ultimately abandoned it; deleting and then re-uploading may reflect the creator's dissatisfaction with the initial plan and a desire to try something new. These patterns are key clues for judging the degree of hesitation in the creator's intentions. The degree of hesitation in the creator's intentions is a quantitative indicator used to measure the degree of uncertainty or repetition exhibited by the creator in a specific batch of visual changes. This degree of hesitation is calculated by analyzing the frequency and intensity of contradictory or inconsistent operational patterns and their correlation with visual disparity. For example, when the visual disparity is large and accompanied by frequent repeated modification operations, the degree of hesitation will increase accordingly.

[0051] Therefore, determining whether the batch of visual changes reflects hesitation or repetition of true intentions, or is merely a random fluctuation in system interaction or user habits, is the final decision-making process. By comprehensively considering the calculated degree of hesitation, the system can distinguish between two situations: one is that the creator genuinely hesitates or repeatedly explores the creative direction, which is part of the true intention; the other is superficial behavioral fluctuations caused by randomness in system interaction (such as accidental operation) or the creator's personal usage habits (such as habitual saving and frequent previewing), which are not drifts in true intentions.

[0052] This invention, by introducing in-depth analysis of creator behavior sequences, can more precisely capture the true intentions of creators in the digital creative nurturing process. Traditionally, judging the drift of creative intention solely based on the degree of difference in visual content may fail to distinguish between exploratory attempts and genuine directional shifts. For example, a creator might upload a series of visually disparate materials, but this could simply be a rapid experimentation with different styles or themes within a short period. Without considering subsequent behavior, such visual differences might be misjudged as intention drift. By analyzing the creator's browsing time, number of modifications, types of modified content, and deletion operations within a preset time window, the system can construct a more comprehensive user behavior profile. When contradictory or inconsistent operational patterns are identified, such as repeated modifications to the same visual element or modifying before deleting, these behavioral patterns directly reflect the creator's hesitation and uncertainty in the decision-making process. Combined with visual differences, these behavioral patterns are used to calculate the degree of hesitation in the creator's intention. The higher the degree of hesitation, the more likely the creator is in an exploratory or iterative phase, rather than firmly turning to a new direction. Thus, the system is able to distinguish visual changes caused by exploratory attempts or hesitation from drifts in genuine creative intent, avoiding misjudgment of the creator's intentions.

[0053] Through the above technical solution, this invention can significantly improve the accuracy and precision of judging the drift of creators' creative intentions. Compared with methods that rely solely on visual content analysis, this solution introduces a deep insight into the creator's behavioral sequence, enabling the system to effectively distinguish between the creator's exploratory attempts, hesitations and repetitions, and genuine intention drift. Specifically, by quantifying the degree of hesitation in the creator's intention, the system can avoid misjudging transient, exploratory visual changes as fundamental shifts in creative direction, thereby reducing unnecessary adjustments to the conflicting graph and the generation of creative materials, and improving the efficiency and relevance of the creative incubation process. In addition, this solution can also identify and eliminate superficial behavioral fluctuations caused by the randomness of system interactions or user habits, making the capture of the creator's true intention more accurate, providing a more reliable basis for subsequent adjustments to the conflicting graph and the generation of creative materials, and ultimately improving the overall effect of digital creative incubation and user experience.

[0054] In some embodiments of the present invention described above, when determining whether batch visual changes are exploratory attempts or a drift in true intent, the degree of hesitation of the creator's intent is calculated by identifying contradictory or inconsistent operational patterns in the behavioral sequence. However, in the actual creative nurturing process, the creator's behavioral patterns may be more complex and intertwined. For example, they may simultaneously make exploratory modifications to multiple visual elements or make meticulous and repeated adjustments to a single visual element. If these complex and intertwined repeated modification patterns are not effectively decoupled and identified, the assessment of the degree of hesitation of the creator's true intent may be inaccurate, thus affecting the accurate judgment of creative intent drift. To address this, the present invention further proposes a more refined method, which, within a preset time window, timestamps and records the operation type of the creator's operation sequence for each visual element, and based on this, identifies parallel exploration patterns and single-element hesitation patterns, thereby decoupling and identifying complex and intertwined repeated modification patterns.

[0055] Within a preset time window, the creator's operation sequence for each visual element is timestamped and the operation type is recorded; the degree of temporal overlap of operations for different visual elements in the operation sequence is identified; when the temporal overlap of operations for multiple visual elements exceeds a preset threshold, these overlapping operations are classified as parallel exploration mode; when the operation of a single visual element appears repeatedly in a short period of time, and its operation type has subtle semantic differences, these operations are classified as single-element hesitation mode; based on the parallel exploration mode and the single-element hesitation mode, the complex and intertwined repeated modification mode is decoupled and identified.

[0056] Specifically, within a preset time window, the system is configured to record in detail all operations performed by the creator interacting with visual elements on the creative interface. This record includes the timestamp of each operation and the specific type of operation, such as adding, deleting, moving, scaling, color adjustment, texture modification, etc. Each visual element, such as an image, video clip, text box, or graphic, will have its own independent operation sequence tracked. Identifying the degree of temporal overlap between operations on different visual elements within the operation sequence refers to the system analyzing the operation timestamps of different visual elements to determine whether they occurred within similar time periods. The degree of overlap can be quantified by calculating the ratio of the intersection to the union of two operation time windows.

[0057] Furthermore, when the overlap in the operation time of multiple visual elements exceeds a preset threshold, these overlapping operations are categorized as parallel exploration mode. This means that the creator may be simultaneously trying different creative directions or making preliminary adjustments to multiple elements to explore the overall effect. For example, the creator may simultaneously adjust the background color, the position of the foreground object, and the font of the text within a short period of time, indicating that they are conducting multi-faceted parallel exploration. In addition, when the operation of a single visual element occurs repeatedly within a short period of time, and the operation type has subtle semantic differences, these operations are categorized as single-element hesitation mode. This indicates that the creator may be making fine adjustments or repeatedly deliberating on a specific visual element, such as repeatedly adjusting the saturation, contrast, or cropping area of ​​the same image to seek the best performance. The subtle semantic differences can refer to minor changes in operation parameters or slight fluctuations in operation intent. Thus, based on the parallel exploration mode and the single-element hesitation mode, the system can decouple and identify the complex, intertwined repeated modification patterns. This means that the system no longer simply regards all repeated modifications as "contradictory or inconsistent," but can distinguish between multi-element parallel exploration and single-element deep hesitation, thereby more accurately understanding the creator's true intention.

[0058] This invention captures more detailed user behavior data by introducing timestamps and operation type records for creator operation sequences. By analyzing the temporal overlap of operations on different visual elements, the system can distinguish whether the creator is simultaneously exploring multiple creative directions (parallel exploration mode) or deeply optimizing or hesitating on a specific element (single-element hesitation mode). This refined pattern recognition allows the system to decouple what was previously considered a general "repeated modification" behavior into "parallel exploration" and "single-element hesitation" with different intentions. For example, when a creator modifies multiple unrelated visual elements within a short period, the system identifies it as parallel exploration, indicating the creator may be in a creative brainstorming phase; while when the creator repeatedly adjusts a certain attribute of the same visual element, it is identified as single-element hesitation, indicating the creator may be experiencing decision-making difficulties or pursuing perfection on that element. This decoupled recognition mechanism allows the system to more accurately understand the deeper intentions behind the creator's behavior, thus avoiding misjudging exploratory attempts as a drift of true intentions or misjudging deep optimization as simple indecisiveness.

[0059] Through the above technical solution, this invention can analyze creators' modification behavior more precisely, thereby more accurately determining whether batch visual changes are exploratory attempts or a drift in true intent. Compared to merely identifying contradictory or inconsistent operational patterns, this invention, by distinguishing between parallel exploration patterns and single-element hesitation patterns, can effectively decouple and identify the creator's complex and intertwined iterative modification patterns. This allows the system to gain a deeper understanding of the creator's creative psychology and intent at different stages; for example, identifying whether the creator is engaging in multi-dimensional creative brainstorming or repeatedly refining a particular detail. Consequently, it can significantly improve the accuracy of calculating the creator's degree of intentional hesitation, avoiding misjudgments caused by the complexity of behavioral patterns, and thus providing more precise guidance for subsequent creative material generation and strategy optimization, improving the efficiency of digital creative cultivation and user experience.

[0060] Specifically, this invention also proposes a digital creative incubation system configured to execute the aforementioned digital creative incubation method based on image and video analysis. Specifically, the digital creative incubation system can be one or more computing devices, such as servers, workstations, personal computers, mobile devices, or cloud computing platforms. The system is designed to receive, process, and analyze image and video data, and interact with creators. The system typically includes one or more processors, memory, and a communication interface, wherein the processor executes instructions, the memory stores program code and data, and the communication interface exchanges data with other devices or networks.

[0061] As a preferred implementation, the digital creative incubation system can adopt a modular design, where each module is responsible for executing one or more specific steps in the method. For example, the visual information processing module can use a deep learning model to parse the reference media content to extract visual information. The contradiction graph management module can construct and manage contradiction graphs based on graph database technology and implement an intent clarification inquiry mechanism through natural language processing technology. The creative generation and optimization module can integrate generative models such as generative adversarial networks (GANs) or variational autoencoders (VAEs) to generate creative materials, and monitor the interaction between creators and creative materials through a user behavior analysis module to iteratively optimize the generation strategy.

[0062] The digital creative incubation system of this invention automates and efficiently executes the method by instantiating each step of the aforementioned image and video analysis-based digital creative incubation method into system components or software modules. Specifically, when the system receives reference media content, the visual information processing module is activated to parse the content and extract visual information. Subsequently, the contradiction graph management module uses this visual information to construct and dynamically adjust the contradiction graph, converging creative intentions through interaction with the creator. Finally, the creative generation and optimization module generates creative materials based on the converged creative intentions and iteratively optimizes the generation strategy by continuously monitoring the creator's interactive behavior. Thus, the system provides creators with a fully intelligent auxiliary environment from creative germination to material generation, effectively improving the efficiency and quality of digital creative incubation.

[0063] Through the above technical solution, this invention provides a concrete digital creative incubation system, enabling the above methods to be practically deployed and operated. This system can automatically perform complex image and video analysis, contradiction mapping construction and adjustment, and creative material generation and optimization processes, significantly reducing the complexity and workload of manual operations. Furthermore, through systematic implementation, it ensures the consistency and stability of method execution, improves the efficiency and accuracy of creative incubation, and provides creators with a stable, efficient, and intelligent creative support platform, thereby promoting the development of the digital creative industry.

[0064] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A digital creativity cultivation method based on image and video analysis, characterized in that, Includes the following steps: The visual extraction step involves receiving reference media content and parsing the reference media content to extract visual information. The contradiction map construction step involves constructing a structured contradiction map representing the conflict between creative directions in the reference media content based on the extracted visual information. The steps for adjusting the contradiction map are as follows: Based on the contradiction map, an intent clarification inquiry mechanism is initiated, guiding questions are output and feedback on the guiding questions is received, and the contradiction map is dynamically adjusted according to the feedback to converge the creative intent. The iterative optimization step involves generating creative materials based on the converged creative intent and continuously analyzing the interaction behavior between the creator and the creative materials to iteratively optimize the generation strategy.

2. The digital creativity cultivation method based on image and video analysis according to claim 1, characterized in that, The visual extraction step includes: Semantic sub-dimensional analysis is performed on the art styles in the extracted visual information to identify semantic sub-dimensional differences within the art styles; Cross-cognitive domain tension analysis is performed on the thematic concepts in the extracted visual information to identify cross-cognitive domain tensions between the thematic concepts; Monitor the sequence of reference media content uploaded iteratively by creators, and analyze the changing trends of visual information between different batches of reference media content in the sequence in order to identify the drift of creators' creative intentions; Based on the semantic sub-dimension differences, the cross-cognitive domain tensions, and the drift of creative intentions, a multi-dimensional contradiction map is constructed. The multi-dimensional contradiction map represents the macro-style and theme conflicts, the semantic sub-dimension differences, the cross-cognitive domain theme tensions, and the drift of creative intentions.

3. The digital creativity cultivation method based on image and video analysis according to claim 2, characterized in that, The steps for identifying semantic sub-dimension differences within the art style include: For each art style in the extracted visual information, identify its multiple semantic sub-dimensions, and define a set of visual feature descriptors for each semantic sub-dimension. The visual feature descriptors include color saturation, texture roughness, line smoothness, light and shadow contrast, and spatial layout density. For each part of the reference media content, calculate the matching degree of its visual feature descriptor on each semantic sub-dimension to obtain the contribution of that part to each semantic sub-dimension. For semantic sub-dimensions within the same art style, when the matching degree of their visual feature descriptors overlaps, the differences in the semantic sub-dimensions are quantified by comparing the contribution degree and combining it with the pixel proportion or frequency of occurrence of the corresponding part in the reference media content.

4. The digital creativity cultivation method based on image and video analysis according to claim 2, characterized in that, The steps for identifying cross-cognitive domain tensions between the aforementioned topical concepts include: Identify the set of concrete visual symbols associated with the theme concept, the set of visual symbols including objects, scenes, colors, light and shadow and compositional elements directly associated with the theme concept; The presence of visual symbols from the set of visual symbols in the reference media content is analyzed, and the degree of concretization of the theme concept is evaluated based on the frequency of occurrence, visual salience, and contextual relevance of the visual symbols in the reference media content. For thematic concepts with low figurative degree, identify the metaphorical visual cues that may exist in the reference media content. The metaphorical visual cues include visual elements that are not directly related but can evoke associations, emotional expressions, and narrative hints. Based on the degree of figuration and the metaphorical visual cues, the abstract semantics of the topic concept are inferred, and the semantic distance between the abstract semantics is calculated; Based on the semantic distance and the degree of concretization of the topic concepts, the cross-cognitive domain tension between the topic concepts is quantified.

5. The digital creativity cultivation method based on image and video analysis according to claim 2, characterized in that, The steps to identify drift in a creator's creative intent include: Visual features are extracted from the reference media content uploaded in each batch of the sequence to obtain a set of visual features for each batch. Calculate the visual difference between each batch of visual feature sets and the previous batch of visual feature sets; The visual differences are semantically calibrated by combining the text descriptions or tags that the creators may have attached when uploading this batch of materials; For batches where the visual differences exhibit weak or non-linear changes, a time decay weighting mechanism is introduced, assigning lower weights to the visual feature sets of earlier batches and higher weights to the visual feature sets of the latest batches. By considering whether the creator viewed, modified, or deleted the materials shortly after uploading them, we can determine whether the visual changes in this batch were an exploratory attempt or a shift in the true intent. Based on the judgment result, update the direction and intensity of the creative intention drift.

6. The digital creativity cultivation method based on image and video analysis according to claim 5, characterized in that, The steps for semantically calibrating the visual dissimilarity include: Keyword extraction and sentiment analysis are performed on the text descriptions or tags to obtain a set of semantic features; Visual saliency analysis and topic concept recognition are performed on the visual feature set to obtain a visual semantic set; Identify semantic elements that conflict with the set of semantic features and the set of visual semantics, and calculate the conflict weight based on the number and intensity of the conflicting elements; Identify ambiguous or missing semantic elements between the semantic feature set and the visual semantic set, and calculate uncertainty weights based on the number and importance of the ambiguous or missing elements; The calibration intensity of the text description or label for the visual difference is dynamically adjusted based on the conflict weight and the uncertainty weight.

7. The digital creativity cultivation method based on image and video analysis according to claim 5, characterized in that, Visual features are extracted from the reference media content uploaded in each batch of the sequence to obtain a set of visual features for each batch. Calculate the visual difference between each batch of visual feature sets and the previous batch of visual feature sets; The visual differences are semantically calibrated by combining the text descriptions or tags that the creators may have attached when uploading this batch of materials; In batches uploaded frequently within a short period of time, identify and extract the core semantic elements of each batch; Calculate the semantic evolution distance between core semantic elements in adjacent batches; When the semantic evolution distance exceeds a preset threshold, the weight of the latest batch is increased; Otherwise, reduce the weight of the latest batch; By considering whether the creator viewed, modified, or deleted the materials shortly after uploading them, we can determine whether the visual changes in this batch were an exploratory attempt or a shift in the true intent. Based on the judgment result, update the direction and intensity of the creative intention drift.

8. The digital creativity cultivation method based on image and video analysis according to claim 5, characterized in that, The steps to determine whether the visual changes in this batch are exploratory attempts or a true intention drift include: After the creator uploads the batch of materials, behavioral sequence analysis is performed on the creator's browsing time, number of modifications, types of modified content, and deletion operations within a preset time window; Identify whether there are contradictory or inconsistent operation patterns in the sequence of behaviors. The contradictory or inconsistent operation patterns include repeated modification of the same visual element, modification followed by deletion, or deletion followed by re-uploading. Based on the contradictory or inconsistent operating patterns, and in conjunction with the visual difference, the degree of hesitation of the creator's intention is calculated; Based on the degree of hesitation, determine whether the visual changes in the batch are due to hesitation or repetition of true intentions, or to random fluctuations in system interaction or user habits.

9. The digital creativity cultivation method based on image and video analysis according to claim 8, characterized in that, Within a preset time window, the creator's sequence of operations on each visual element is timestamped and the operation type is recorded. Identify the degree of temporal overlap between operations targeting different visual elements in the operation sequence; When the overlap of operation times of multiple visual elements exceeds a preset threshold, these overlapping operations are classified as parallel exploration mode. When the operation of a single visual element occurs repeatedly in a short period of time, and the operation types have subtle semantic differences, these operations are classified as single-element hesitation mode. Based on the parallel exploration mode and the single-element hesitation mode, the complex, intertwined iterative modification patterns are decoupled and identified.

10. A digital creative incubation system for performing the digital creative incubation method based on image and video analysis as described in any one of claims 1-9.