Picture and video material collection method and system based on artificial intelligence

By using artificial intelligence technology to acquire user-input material and project information, perform demand analysis and path planning, and optimize the collection scheme and parameters, the problem of low material collection efficiency and difficulty in guaranteeing quality in existing technologies is solved, achieving efficient and accurate material collection results.

CN121722964APending Publication Date: 2026-03-24FUTONG (SHANGHAI) INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing methods for collecting materials require filtering through massive amounts of images and videos, resulting in low collection efficiency and difficulty in accurately matching complex needs, leading to low quality.

Method used

By using artificial intelligence-based methods to acquire user-input material and project information, analyze material requirements, determine collection weights and path planning, and optimize collection schemes and parameters through multimodal data analysis and adjustment, we can achieve efficient and accurate material collection.

Benefits of technology

It significantly improves the efficiency and quality of material collection, reduces the cost of invalid collection and storage, enhances material consistency and subsequent processing efficiency, and ensures the accuracy of collection results and the satisfaction of user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121722964A_ABST
    Figure CN121722964A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, in particular to a picture and video material collection method and system based on artificial intelligence. The method comprises the steps of performing material demand analysis on to-be-collected material information and material item information, and determining material demand structure information and a material collection weight; performing acquisition path planning analysis on the material demand structure information by using the material acquisition weight to obtain a data acquisition scheme and data acquisition parameters; carrying out picture material and video material collection on the data collection scheme and the data collection parameters to obtain material initial data; performing multi-modal data analysis on the material initial data, and determining a material qualification rate of the material initial data; and adjusting the data acquisition scheme and the data acquisition parameters by using the material qualification rate, and then acquiring picture materials and video materials to obtain material data. According to the invention, fuzzy demand expression of the user is avoided, and the efficiency and the quality of picture and video material collection are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a picture and video material collection method and system based on artificial intelligence. BACKGROUND

[0002] In the process of digital creation and content production, users need to use picture materials and video materials for information transmission and visual performance enhancement. The existing material collection method mainly collects materials through online search.

[0003] However, in the process of collecting materials, the existing material collection method needs to search by keyword through a search engine or a gallery platform when using online search. After the user inputs the keyword, the demand expression is prone to be ambiguous. This makes the quality of the materials returned by online search low, and it is difficult to accurately match the complex picture material and video material collection requirements. At the same time, the user needs to filter a large amount of materials after keyword search, which makes the collection of picture materials and video materials inefficient and prone to errors. SUMMARY

[0004] The present application aims to provide a picture and video material collection method and system based on artificial intelligence, which solves the problem of low collection efficiency caused by the need to filter a large amount of materials when collecting picture materials and video materials by the existing material collection method.

[0005] To achieve the above-mentioned purpose, the present application provides a picture and video material collection method based on artificial intelligence, comprising: Obtaining user inputted material information to be collected and material project information, the material information to be collected including text description information, voice description information, picture reference information and video reference information; Analyzing the material demand of the material information to be collected and the material project information, determining the material demand structure information and the material collection weight; Using the material collection weight, analyzing the collection path of the material demand structure information, obtaining the data collection scheme and data collection parameters; Collecting picture materials and video materials according to the data collection scheme and data collection parameters, obtaining material initial data; Analyzing the multi-modal data of the material initial data, determining the material qualified rate of the material initial data; Using the material qualified rate, adjusting the data collection scheme and data collection parameters before collecting picture materials and video materials, obtaining material data containing pictures and videos.

[0006] The step of performing material demand analysis on the to-be-collected material information and the material item information, and determining material demand structure information and a material collection weight comprises the following steps: Perform initial analysis on the material demand structure of the to-be-collected material information and the material item information by using a preset multi-modal fusion model, and determine original material demand structure information; Determine a picture information qualification rate and a video information qualification rate based on the original material demand structure information; Adjust the preset multi-modal fusion model by using the picture information qualification rate and the video information qualification rate, and obtain an adjusted preset multi-modal fusion model; Perform material demand structure analysis on the to-be-collected material information and the material item information by using the adjusted preset multi-modal fusion model, and determine material demand structure information.

[0007] The step of determining a picture information qualification rate and a video information qualification rate based on the original material demand structure information comprises the following steps: Extract picture data demand information and video data demand information from the original material demand structure information; Perform completeness verification on the picture data demand information and the video data demand information respectively, and obtain verified picture data demand information and verified video data demand information; Perform data information analysis on the verified picture data demand information and the verified video data demand information respectively, and obtain a picture information qualification rate and a video information qualification rate.

[0008] The step of adjusting the preset multi-modal fusion model by using the picture information qualification rate and the video information qualification rate, and obtaining an adjusted preset multi-modal fusion model comprises the following steps: Adjust the preset multi-modal fusion model by using the picture information qualification rate and the video information qualification rate, and obtain an initially adjusted preset multi-modal fusion model; Perform model verification on the initially adjusted preset multi-modal fusion model, and obtain a verified adjusted preset multi-modal fusion model; If the initially adjusted preset multi-modal fusion model is not verified, then re-determine a picture information qualification rate and a video information qualification rate, and then adjust the preset multi-modal fusion model, and obtain an initially adjusted preset multi-modal fusion model.

[0009] The step of performing collection path planning analysis on the material demand structure information by using the material collection weight, and obtaining a data collection scheme and data collection parameters comprises the following steps: Determine all material data sources based on the material demand structure information; Using the weights of all the material collections, each material data source is prioritized to obtain a sorted material data source. Using each sorted material data source, a data acquisition path planning analysis is performed on the material demand structure information to obtain a data acquisition scheme and data acquisition parameters.

[0010] The step of prioritizing each material data source using the collection weights of all the materials to obtain the sorted material data source includes: Determine the proportion of each of the aforementioned material data sources in the total data volume of all material data sources for each data source; By using the weights of all the collected materials and the proportion of each data source, each of the material data sources is prioritized to obtain the sorted material data sources.

[0011] The step of performing multimodal data analysis on the initial material data to determine the material qualification rate of the initial material data includes: Multimodal data analysis was performed on the initial material data to determine each item type and each complete tag of the initial material data; Perform pass rate analysis on each item type of the initial data of the material, and determine the pass rate threshold for each item type; Based on the pass rate threshold for each project type and each complete tag of the initial material data, the material pass rate of the initial material data is determined.

[0012] The step of determining the material pass rate of the initial material data based on the pass rate threshold for each project type and each complete tag of the initial material data includes: Each complete tag of the initial material data is matched with each project type to obtain all complete tags corresponding to each project type; Compare all complete tags corresponding to each project type with the pass rate threshold for each project type to determine the pass rate of the initial material data.

[0013] The steps for obtaining user-inputted information about the materials to be collected and the material project information include: Obtain the original information of the materials to be collected and the original information of the material items input by the user; The original information of the materials to be collected and the original information of the material items are preprocessed to obtain the preprocessed original information of the materials to be collected and the original information of the material items. The preprocessed original information of the materials to be collected and the original information of the material items are verified to obtain the information of the materials to be collected and the information of the material items.

[0014] This invention also provides an artificial intelligence-based image and video material collection system, applicable to any of the above-mentioned artificial intelligence-based image and video material collection methods, comprising: The information acquisition module is used to acquire the information of the materials to be collected and the information of the materials project input by the user. The information of the materials to be collected includes text description information, voice description information, image reference information and video reference information. The information determination module is used to perform material demand analysis on the material information to be collected and the material project information, and to determine the material demand structure information and material collection weight; The planning and analysis module is used to perform data acquisition path planning and analysis on the material acquisition weights to obtain data acquisition schemes and data acquisition parameters. The data acquisition module is used to acquire image and video materials according to the data acquisition scheme and data acquisition parameters to obtain initial material data; The data analysis module is used to perform multimodal data analysis on the initial data of the material to determine the material qualification rate of the initial data. The data adjustment module is used to adjust the data acquisition scheme and data acquisition parameters based on the material qualification rate before acquiring image and video materials, thereby obtaining material data containing images and videos.

[0015] The present invention provides a method and system for collecting image and video materials based on artificial intelligence, which has the following advantages: This invention enhances the flexibility of user expression and avoids information omissions by acquiring user-inputted information on materials to be collected and material items. It analyzes the material requirements based on this information to determine the material requirement structure and collection weights, transforming vague requirements into quantifiable indicators to guide subsequent collection. Weight allocation ensures that key requirements are prioritized, improving collection efficiency. Simultaneously, using the material collection weights, it analyzes the collection path based on the material requirement structure to obtain a data collection scheme and parameters. This reduces invalid collection and lowers storage and bandwidth costs; parameter standardization improves material consistency and facilitates subsequent processing. Next, image and video materials are collected using the data collection scheme and parameters to obtain initial material data. This rapidly accumulates raw materials, providing a data foundation for subsequent analysis. Multimodal data analysis is then performed on the initial material data to determine the material qualification rate. Low-quality materials are filtered out, reducing subsequent processing costs; the quantified qualification rate provides feedback on collection effectiveness, guiding subsequent adjustments. Finally, using the material qualification rate, the data collection scheme and parameters are adjusted before collecting image and video materials again, resulting in material data containing both images and videos. Through feedback on pass rates and iterative optimization, the final quality of the materials is significantly higher than that of traditional manual collection. This not only avoids situations where users express their needs vaguely, the collection efficiency is low, and the quality is difficult to control, but also significantly improves the efficiency and quality of image and video material collection. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0017] Figure 1 This is a flowchart of the first embodiment of the present invention.

[0018] Figure 2 This is a flowchart illustrating the process of determining the material requirement structure information in step S200 of the second embodiment of the present invention.

[0019] Figure 3 This is a flowchart of step S300 in the fifth embodiment of the present invention.

[0020] Figure 4 This is a flowchart of step S500 in the seventh embodiment of the present invention.

[0021] Figure 5 This is a flowchart of step S100 in the ninth embodiment of the present invention.

[0022] Figure 6 This is a functional block diagram of an artificial intelligence-based image and video material collection system according to the present invention.

[0023] In the diagram: Information acquisition module, 610; Information determination module, 620; Planning and analysis module, 630; Data acquisition module, 640; Data analysis module, 650; Data adjustment module, 660. Detailed Implementation

[0024] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, but should not be construed as limiting the present invention.

[0025] The first embodiment of this application is as follows: Please see Figure 1 An AI-based method for collecting image and video materials includes the following steps: S100: Obtain the user-inputted information on materials to be collected and material item information. The material information to be collected includes text description information, voice description information, image reference information, and video reference information. Specifically, the material information to be collected is the type of material the user wishes to obtain, such as a technologically advanced city night scene image or a corporate promotional video clip. The material item information is the context of the material application, such as for short video platform advertising or a film storyboard material library. This step collects data through multimodal input interfaces (such as text boxes, speech-to-text, and image or video uploads) and utilizes NLP to parse text semantics, speech recognition and translation, and image or video feature extraction, such as CNN to extract visual features. This ensures that the user's needs are fully expressed and avoids information omissions; at the same time, multimodal input enhances the flexibility of user expression and lowers the barrier to entry.

[0026] S200: Perform material requirement analysis on the collected material information and material project information to determine the material requirement structure information and material collection weights. The material requirement structure information is a hierarchical classification of requirements, such as a main visual style of cyberpunk, material type of 4K landscape video combined with accompanying static images, and a quantity of 10 sets. Material collection weights are importance coefficients for different dimensions, such as a video clarity weight of 0.6 and a color saturation weight of 0.4. This step uses a semantic analysis model (such as the BERT model) to parse text or voice requirements, combined with feature matching of image or video references, such as style transfer algorithms to determine the style of reference images, to obtain a structured requirement tree; material collection weights are obtained through the AHP hierarchical analysis method or machine learning models (such as random forests). This transforms vague requirements into quantifiable indicators to guide subsequent collection; weight allocation ensures that key requirements are prioritized, improving collection efficiency.

[0027] S300: Utilizing the aforementioned material acquisition weights, perform acquisition path planning analysis on the material demand structure information to obtain a data acquisition scheme and data acquisition parameters. The data acquisition scheme includes acquisition channels, such as public image libraries, user-generated content platforms, and self-shot content; acquisition tools, such as web crawlers or API interfaces; and time and spatial ranges. Data acquisition parameters are specific technical parameters such as resolution, frame rate, color space, and keyword tags. Through the material demand structure information and material acquisition weights, reinforcement learning algorithms (such as Q-learning) are used to optimize the acquisition path, such as prioritizing access to high-weight material sources; parameters are automatically mapped through the demand structure, such as cyberpunk style corresponding to neon tones and futuristic architecture tags. This reduces invalid acquisition, lowers storage and bandwidth costs; simultaneously, parameter standardization improves material consistency and facilitates subsequent processing.

[0028] S400: Collect image and video materials according to the data acquisition scheme and parameters to obtain initial material data. This step uses automated tools (such as the Scrapy web crawling framework or API calls) to collect data according to the scheme; video acquisition needs to consider frame extraction, such as extracting keyframes per second as image materials. This quickly accumulates raw materials, providing a data foundation for subsequent analysis; at the same time, multi-source acquisition improves the diversity of materials.

[0029] S500: Perform multimodal data analysis on the initial material data to determine the material qualification rate. The material qualification rate is the percentage of materials that meet the required standards, such as 90% of video clips meeting the resolution standard. This step uses a multimodal fusion model (such as CLIP cross-modal alignment) to evaluate the matching degree between the materials and the requirements; and employs computer vision techniques such as the National Image Quality Evaluation (NIQE) model or video smoothness detection to quantify quality indicators. This helps to filter out low-quality materials, reducing subsequent processing costs; and provides feedback on the acquisition effect through the qualification rate quantification, guiding subsequent adjustments.

[0030] S600: Using the aforementioned material pass rate, adjust the data acquisition scheme and parameters before acquiring image and video materials, obtaining material data containing images and videos. This step, based on the pass rate analysis results, adjusts weight allocation or parameters if the pass rate of a certain channel is low. Weight allocation adjustments may include reducing the weight of that channel, or parameter adjustments may include increasing resolution requirements. The acquisition and analysis process is then re-executed. Iterative optimization improves material quality and acquisition efficiency; ultimately, the material data highly meets user needs, reducing manual screening costs.

[0031] In this embodiment, by acquiring user-inputted information on materials to be collected and material item information, the flexibility of user expression is improved, and information omissions are avoided. Material demand analysis is performed on the material demand information and material item information to determine the material demand structure information and material collection weights, thereby transforming vague demands into quantifiable indicators to guide subsequent collection. Weight allocation ensures that key demands are prioritized, improving collection efficiency. Simultaneously, using the material collection weights, collection path planning analysis is performed on the material demand structure information to obtain a data collection scheme and data collection parameters. This reduces invalid collection and lowers storage and bandwidth costs; parameter standardization improves material consistency and facilitates subsequent processing. Next, image and video materials are collected using the data collection scheme and data collection parameters to obtain initial material data. This quickly accumulates raw materials, providing a data foundation for subsequent analysis. Multimodal data analysis is then performed on the initial material data to determine the material qualification rate. By filtering low-quality materials, subsequent processing costs are reduced; the quantified qualification rate provides feedback on the collection effect, guiding subsequent adjustments. Finally, using the material qualification rate, the data collection scheme and data collection parameters are adjusted before image and video materials are collected again, resulting in material data containing images and videos. Through feedback on pass rates and iterative optimization, the final material quality is significantly higher than that of traditional manual collection; it avoids situations where users have vague requirements, low collection efficiency, and difficulty in controlling quality, while significantly improving the efficiency and quality of image and video material collection.

[0032] The second embodiment of this application is as follows: Please see Figure 2 Based on the first embodiment, in step S200 of this embodiment: performing material demand analysis on the material information to be collected and the material project information to determine the material demand structure information and the material collection weight, the step of determining the material demand structure information includes: S210 utilizes a pre-set multimodal fusion model to perform an initial analysis of the material requirement structure of the collected material information and material project information, determining the original information of the material requirement structure. The pre-set multimodal fusion model is a pre-trained cross-modal understanding model, such as the CLIP model or the multimodal BERT model, capable of mapping text, speech, images, and videos to a unified semantic space, achieving feature fusion. The original information of the material requirement structure is a preliminary structured requirement description output by the model, such as the main visual style being cyberpunk; the material type being 4K landscape video combined with accompanying static images; and the quantity being 10 sets. The pre-set multimodal fusion model is pre-constructed based on the historical information of the collected material and material project as input, and the original information of the material requirement structure as output. The collected material information and material project information are used as input to the pre-set multimodal fusion model. The pre-set multimodal fusion model extracts features from each modality through an encoder; text is encoded using BERT, images using ResNet, and videos using 3D-CNN to extract spatiotemporal features. These features are then fused through a cross-modal alignment layer (such as an attention mechanism), outputting structured requirement labels as the original information of the material requirement structure. Cross-modal fusion improves the accuracy of understanding requirements and avoids information loss in single-modal scenarios.

[0033] S220 determines the image information pass rate and video information pass rate based on the original information of the material requirement structure. The image information pass rate is the proportion of image materials that meet the requirement structure, and the video information pass rate is the proportion of video materials that meet the requirement structure. Specifically, an image quality assessment model (such as the NIQE model) is used to detect image sharpness and color matching; a video quality model (such as VMAF) is used to detect frame rate stability and dynamic range. The pass rate is determined by a threshold; for example, if the pass rate threshold is 85%, anything not lower than the threshold is considered passable. This provides a clear optimization direction for adjusting the preset multimodal fusion model.

[0034] S230 uses the image information pass rate and video information pass rate to adjust the preset multimodal fusion model, obtaining the adjusted preset multimodal fusion model. This step uses gradient descent to adjust the model weights. For example, if the image information pass rate is lower than a threshold, the weight of the image feature extraction layer is increased; if the video information pass rate is low, the attention weight of the video spatiotemporal feature fusion module is adjusted. This dynamically optimizes model performance and improves the accuracy of subsequent analysis.

[0035] S240 utilizes the adjusted preset multimodal fusion model to perform material requirement structure analysis on the material information to be collected and the material project information, determining the material requirement structure information. In this step, the adjusted model reprocesses the input information, outputting a fine-grained requirement structure, such as subdividing cyberpunk style into three sub-dimensions: neon tones, futuristic architecture, and dynamic lighting. This generates high-precision requirement structure information to guide subsequent data acquisition path planning, reduce requirement ambiguity, and improve material acquisition efficiency. Furthermore, the quality of the requirement analysis is ensured through quantitative evaluation of the pass rate.

[0036] The third embodiment of this application is as follows: Based on the second embodiment, the steps in this embodiment to determine the image information pass rate and video information pass rate based on the original information of the material demand structure include: This step extracts the image and video data requirement information from the original information of the material requirement structure. Image data requirement information refers to the user's specific parameter requirements for image materials, such as resolution ≥ 4K, color mode GB / CMYK, style tag cyberpunk or minimalism, and quantity ≥ 10 images. Video data requirement information refers to the user's dynamic parameter requirements for video materials, such as frame rate ≥ 60fps, encoding format H.265, duration 15-30 seconds, and dynamic range HDR / SDR. This step uses a multimodal feature decoder (such as CLIP's text and image alignment module) to parse the image / video related descriptions in the original information and extract structured parameters. For example, the parameters parsed from the text information requiring 100 cyberpunk-style 4K landscape images are: resolution 4K, style cyberpunk, quantity 100, and aspect ratio 16:9. This allows for precise positioning of the required dimensions for images or videos, avoiding interference from irrelevant information.

[0037] The image and video data requirements are each subjected to a completeness check to obtain qualified requirements. Specifically, the completeness check verifies whether the requirement information includes key parameters such as resolution, style, or quantity. Missing items are marked as incomplete and trigger a supplementation process. A rule engine is used to match preset verification rules, such as requiring image requirements to include resolution, style, and quantity. For example, if an image requirement only describes a cyberpunk style but lacks resolution, it is marked as incomplete; similarly, if a video requirement lacks frame rate or encoding format, it is also marked as incomplete. This filters out invalid requirement descriptions, reducing subsequent analysis errors. Simultaneously, feedback on missing items guides users to supplement information, improving the completeness of the requirement expression.

[0038] Data information analysis was performed on the qualified image and video data requirements to obtain the image and video information pass rates. Specifically, the data information analysis used computer vision models (such as the National Image Quality Evaluation (NIQE) model or the Virtual Machine Image Flow (VMAF) model) to quantify the rationality of the requirement parameters. Image analysis used an image quality model to check if the resolution met the standard (e.g., 4K images need to be ≥3840×2160 pixels), and a style transfer algorithm to verify the cyberpunk style matching degree, such as color distribution and texture features. Video analysis used a video encoder analyzer to check if the frame rate was ≥60fps, a dynamic range detector to verify HDR / SDR compatibility, and a duration counter to check the segment length. The pass rate feedback drove the iteration of the requirement structure, improving subsequent acquisition efficiency; it also quantified the feasibility of requirement parameters, identified high-risk requirements, such as the feasibility of 4K resolution in low-bandwidth scenarios, and provided data support for subsequent acquisition parameter adjustments.

[0039] The fourth embodiment of this application is as follows: Based on the second embodiment, this embodiment utilizes the image information pass rate and video information pass rate to adjust the preset multimodal fusion model, and the steps to obtain the adjusted preset multimodal fusion model include: Using the pass rates of the image and video information, the preset multimodal fusion model is adjusted to obtain an initially adjusted preset multimodal fusion model. This initially adjusted preset multimodal fusion model is the first improved model after pass rate feedback optimization, and its parameters have been gradient-corrected based on the image or video pass rates. A backpropagation mechanism using the loss function is employed. Utilizing the original loss of the preset multimodal fusion model, parameters are updated through gradient descent to optimize weight dimensions negatively correlated with the pass rate, such as the visual feature extraction layer. Specifically, the loss function of the initially adjusted preset multimodal fusion model is: ; in, This refers to the loss value adjusted after setting the multimodal fusion model. This refers to the preset original loss value of the multimodal fusion model. Refers to the model learning rate in the preset multimodal fusion model; To ensure the pass rate of image information, The pass rate for video information.

[0040] The initially adjusted preset multimodal fusion model is validated to obtain a validated preset multimodal fusion model. A validated preset multimodal fusion model is an improved model that meets preset performance indicators after cross-validation or test set evaluation, such as accuracy ≥ 90% and recall ≥ 85%. Model validation employs a dual validation mechanism of indicator validation and business validation. Indicator validation calculates the matching degree between the model's output requirement structure and the real labels on an independent test set, such as the F1 score. Business validation simulates a data collection scenario to test whether the requirement structure generated by the model can guide the collection of qualified materials, such as actually collecting 100 images and calculating the pass rate. This ensures the generalization ability of the adjusted model and avoids local optimization; a validated model can stabilize the pass rate of subsequent collected materials above 90%, reducing the cost of manual review.

[0041] If the initial adjusted preset multimodal fusion model fails verification, the pass rates for image and video information are recalculated before further adjustment to obtain the initially adjusted preset multimodal fusion model. In this step, when verification fails, a feedback loop is triggered, i.e., the pass rates for images or videos are recalculated, and new low-quality dimensions are identified, such as a high error rate in style tag recognition. The model loss function is then adjusted accordingly, increasing the weight of the corresponding dimension, such as increasing the style matching weight from 0.3 to 0.5; thus obtaining the initially adjusted preset multimodal fusion model. This dynamically adapts to changing requirements, improving model robustness; after multiple iterations, the model's accuracy in parsing complex requirement structures is significantly improved.

[0042] The fifth embodiment of this application is as follows: Please see Figure 3 Based on the first embodiment, step S300 of this embodiment, which involves using the material acquisition weights to perform acquisition path planning analysis on the material demand structure information to obtain a data acquisition scheme and data acquisition parameters, includes: S310: Based on the material demand structure information, determine all material data sources. Material data sources are channels where materials matching the demand can be obtained, including public image libraries (such as Shutterstock), user-generated content platforms (such as YouTube), self-shot devices (such as 4K cameras), and CGI generation tools (such as Blender). This step uses demand feature matching algorithms (such as knowledge graph-based association reasoning) to parse keywords in the material demand structure information, such as associating cyberpunk with science fiction or futurism tags, and establishes a mapping relationship by combining data source metadata (such as image library tag systems and platform content classifications). For example, the TF-IDF algorithm is used to extract demand keywords, and cosine similarity is used to match data source tags; thus determining all material data sources. By constructing a complete mapping network between demand and data sources, potential high-quality sources are avoided; the coverage of data sources is improved, ensuring comprehensive collection.

[0043] S320: Utilizing the collection weights of all the aforementioned materials, prioritize each of the material data sources to obtain a sorted list of material data sources. The material collection weight quantifies the contribution coefficient of each data source to the satisfaction of the demand, determined using the Analytic Hierarchy Process (AHP) or a machine learning model (such as Random Forest). This step employs a sorting algorithm (such as Top-K selection) to rank the data sources in descending order based on their weight scores. For example, data sources with a weight score ≥ 0.7 are considered high priority, those with a weight score ≤ 0.3 are low priority, and those between 0.3 and 0.7 are medium priority. A secondary sort is then performed, combining data source accessibility (such as API availability) and cost (such as pay-per-use). This optimizes resource allocation efficiency and increases the proportion of collections from high-priority data sources. The formula for calculating the material collection weight is: ; in, The weight of the i-th material is given, and the sum of the weights of all materials is 1. The weight of the j-th demand dimension. This can be obtained using the AHP method, such as a style matching weight of 0.4 and the sum of the weights of all requirement dimensions being 1. The matching score (0-1 scale) of the i-th data source in the j-th dimension is given, such as 0.8 indicating an 80% match; n refers to the total number of requirement dimensions.

[0044] S330: Using each sorted material data source, perform a data acquisition path planning analysis on the material demand structure information to obtain a data acquisition scheme and data acquisition parameters. This step uses a reinforcement learning algorithm (such as Q-learning) to optimize the acquisition path. Define the state space as the data source priority sequence, the action space as the acquisition order selection, and the reward function as the acquisition efficiency (such as material quantity or time). Update the state values ​​using the Bellman equation to ultimately generate a Pareto optimal path (i.e., a solution balancing efficiency and quality). This shortens the acquisition path time and improves the quality of the materials.

[0045] The sixth embodiment of this application is as follows: Based on the fifth embodiment, this embodiment utilizes the weights of all material acquisitions to prioritize each material data source, and the steps for obtaining the sorted material data source include: Determine the proportion of data volume from each of the aforementioned material data sources relative to the total data volume from all material data sources. This step involves using distributed crawlers or API aggregation interfaces to statistically analyze the real-time data volume of each data source, employing a MapReduce framework for parallel computation of the proportions to avoid single-point computation bottlenecks. This quantifies the value of data source scale, providing an objective basis for weighted fusion; simultaneously, it avoids collection biases caused by differences in data source scale, improving ranking fairness.

[0046] Using the weights of all collected materials and the proportion of each data source, each material data source is prioritized to obtain a sorted list. This step uses cosine similarity to match requirement tags with data source tags. For example, the tag matching score for "cyberpunk" is 0.9. Data sources are sorted in descending order of priority scores to generate a priority sequence. For instance, a data source with a priority score of 0.8 is prioritized over a data source with a priority score of 0.6; this balances subjective needs with objective scale, improves the scientific nature of the sorting, and reduces invalid data collection. The formula for calculating the priority score is: ; in, The priority score of the i-th data source; This refers to the weighted fusion coefficient (0-1 scale) of the source material data. This refers to the percentage of the i-th data source.

[0047] The seventh embodiment of this application is as follows: Please see Figure 4 Based on the fifth embodiment, step S500 of this embodiment, which involves performing multimodal data analysis on the initial material data to determine the material qualification rate of the initial material data, includes: S510: Perform multimodal data analysis on the initial material data to determine each item type and each complete label of the initial material data. Item type refers to the modality and technical specifications classification of the material, such as 4K landscape images, 1080p portrait videos, and HDR video clips, defined by parameters such as resolution, aspect ratio, and encoding format. Complete labels are multimodal feature fusion labels; for example, cyberpunk style is defined by color distribution combined with texture features, dynamic range of HDR10, and frame rate of 60fps, generated through cross-modal alignment using the CLIP model. This step uses a multimodal pre-trained model (such as ViT-B / 32 for image processing or 3D-ResNet for video processing) to extract visual features, combines this with NLP modules to parse labels in text descriptions, and generates structured labels through feature fusion layers (such as attention mechanisms). For example, image segmentation can be used to identify future architectural elements, and color histograms can be used to match neon hues. This achieves accurate positioning of material attributes and avoids misjudgments of pass rates due to missing labels.

[0048] S520: Perform pass rate analysis on each item type of the initial material data to determine the pass rate threshold for each item type. Specifically, the pass rate threshold is the minimum acceptable pass rate for each item type; for example, the threshold for 4K video is set to 85%. This step uses Statistical Process Control (SPC) theory to monitor the fluctuation range of the pass rate for each item type through control charts. If the pass rate of a batch of 4K videos is lower than 85%, an anomaly alarm is triggered, requiring adjustments to the acquisition parameters, such as changing the data source or optimizing the crawler strategy. Threshold analysis improves the accuracy of identifying unqualified batches and reduces the risk of misjudgment.

[0049] S530: Based on the pass rate threshold for each project type and each complete label of the initial material data, determine the material pass rate of the initial material data. By using the pass rate threshold for each project type and each complete label of the initial material data, the weight of a project type is only included in the total pass rate when its pass rate is ≥ the threshold, thereby quantifying the overall quality of the material and providing direct feedback for adjusting the data collection plan.

[0050] The eighth embodiment of this application is as follows: Based on the seventh embodiment, this embodiment's steps for determining the material pass rate of the initial material data based on the pass rate threshold for each project type and each complete tag of the initial material data include: Match the complete tags of the initial data of the material with each project type to obtain all the complete tags corresponding to each project type. This step is implemented using a knowledge graph association algorithm to map tags to types. For example, construct an ontology library of tags and types, and match the HDR10 tag to the type of HDR video clip through a SPARQL query engine; calculate the matching degree between the tag vector and the type vector using cosine similarity, and a threshold ≥ 0.8 is regarded as a strong association. Furthermore, ensure that the tag evaluation dimension strictly corresponds to the project type to avoid invalid comparison; make the accuracy of tag and type matching higher and reduce the quality evaluation deviation.

[0051] Compare all the complete tags corresponding to each project type with the qualified rate threshold of each project type to determine the material qualified rate of the initial data of the material. This step uses a siamese network to determine the cosine similarity between the tag and the requirement. For example, the matching degree between the cyberpunk style and the requirement tag is 0.92. Then, perform a threshold judgment on each tag dimension. Only when all dimensions are ≥ the threshold, the material is regarded as qualified. To further improve the quality of the material.

[0052] The ninth embodiment of this application is: Please refer to Figure 5 , on the basis of the first embodiment, step S100 of this embodiment: The step of obtaining the information of the material to be collected and the material project information input by the user includes: S110: Obtain the original information of the material to be collected and the original information of the material project input by the user. Among them, the original information of the material to be collected is the unprocessed material requirement directly input by the user, such as shooting a 15-second minimalist product display video or uploading a reference picture. The original information of the material project is the context information of the material application, such as being used in the e-commerce main picture and the movie storyboard material library, implicitly containing constraints such as format, size, and style. This step collects data through a multimodal input interface (such as a text box, microphone, and file upload). The voice input is converted to text by ASR (Automatic Speech Recognition), and the picture or video is parsed for visual features through a feature extractor (such as ResNet), and the text is embedded into the semantic space by BERT. Furthermore, it supports users to express their needs flexibly, reduces the usage threshold; and reduces information loss.

[0053] S120: The original information of the materials to be collected and the original information of the material items are preprocessed to obtain preprocessed original information of the materials to be collected and original information of the material items. Specifically, preprocessing includes data cleaning, format conversion, and feature standardization. For example, text is filtered for stop words, images are scaled to a uniform resolution, and speech-to-text is corrected. For example, text processing uses TF-IDF or BERT-Tokenization for word segmentation, removes punctuation and stop words, and generates standardized word vectors. Visual processing uses OpenCV to adjust the image size to the 4K standard (3840×2160), and video frame extraction uses FFmpeg to extract keyframes. Speech processing uses Waveform-GAN to enhance speech clarity and reduce ASR translation errors. Data standardization improves the efficiency of subsequent analysis and increases processing speed.

[0054] S130: Perform data validation on the preprocessed original information of the materials to be collected and the original information of the material items to be collected, to obtain the information of the materials to be collected and the information of the material items to be collected. Data validation checks data integrity, format validity, and logical consistency. For example, it verifies whether the image resolution is ≥4K, whether the video length is 15-30 seconds, and whether the text requirements include style, size, and quantity elements. Specifically, integrity validation uses a rule engine to check required fields (e.g., material type cannot be empty), and missing fields trigger a supplementary process. Format validation uses regular expressions to validate formats such as email addresses and phone numbers, and image header information is checked to ensure it is a real image (not a screenshot). Logical validation uses knowledge graphs to detect contradictions (e.g., conflicts between minimalism and cyberpunk styles), and corrects them through ontology reasoning. After validation, data validity is improved, reducing errors in subsequent processes; logical validation identifies contradictory user requirement descriptions, guiding users to correct them in advance, reducing rework costs. Thus, validation ensures data quality, ultimately yielding the information of the materials to be collected and the information of the material items to be collected that meet the requirements analysis.

[0055] Please see Figure 6 The present invention also provides an artificial intelligence-based image and video material collection system, which is used in any of the above-mentioned artificial intelligence-based image and video material collection methods, including an information acquisition module 610, an information determination module 620, a planning and analysis module 630, a data acquisition module 640, a data analysis module 650, and a data adjustment module 660.

[0056] The information acquisition module 610 is used to acquire the material information to be collected and the material project information input by the user. The material information to be collected includes text description information, voice description information, image reference information and video reference information.

[0057] The information determination module 620 is used to perform material demand analysis on the material information to be collected and the material project information, and to determine the material demand structure information and material collection weight.

[0058] The planning and analysis module 630 is used to perform data acquisition path planning and analysis on the material acquisition weights to obtain data acquisition schemes and data acquisition parameters.

[0059] The data acquisition module 640 is used to acquire image and video materials based on the data acquisition scheme and data acquisition parameters to obtain initial material data.

[0060] The data analysis module 650 is used to perform multimodal data analysis on the initial data of the material to determine the material qualification rate of the initial data.

[0061] The data adjustment module 660 is used to adjust the data acquisition scheme and data acquisition parameters based on the material qualification rate before acquiring image and video materials to obtain material data containing images and videos.

[0062] In this step, the information acquisition module 610 acquires the user-inputted information on the materials to be collected and the material project information. The information includes text descriptions, voice descriptions, image references, and video references; this enhances the flexibility of user expression and avoids information omissions. The information determination module 620 performs material requirement analysis on the information on the materials to be collected and the material project information, determining the material requirement structure and material collection weights; this transforms vague requirements into quantifiable indicators to guide subsequent collection. The planning and analysis module 630 uses the material collection weights to perform collection path planning analysis on the material requirement structure information, obtaining a data collection plan and data collection parameters. This reduces invalid collection and lowers storage and bandwidth costs; at the same time, parameter standardization improves material consistency and facilitates subsequent processing. The data collection module 640 collects image and video materials based on the data collection plan and data collection parameters, obtaining initial material data. This quickly accumulates raw materials, providing a data foundation for subsequent analysis. The data analysis module 650 performs multimodal data analysis on the initial material data to determine the material qualification rate of the initial data. By filtering out low-quality materials, subsequent processing costs are reduced; and by quantifying the pass rate, feedback on the collection effect is provided to guide subsequent adjustments. The data adjustment module 660 uses the material pass rate to adjust the data collection plan and parameters before collecting image and video materials, resulting in material data containing images and videos. Through pass rate feedback and iterative optimization, not only are situations where users' requirements are vaguely expressed, collection efficiency is low, and quality is difficult to control avoided, but the efficiency and quality of image and video material collection are also significantly improved.

[0063] The above-disclosed embodiments are merely one or more preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art can understand that all or part of the processes for implementing the above embodiments and equivalent changes made in accordance with the claims of this application still fall within the scope of this application.

Claims

1. A method for collecting image and video materials based on artificial intelligence, characterized in that, include: Obtain user-inputted information on materials to be collected and material project information, wherein the information on materials to be collected includes text description information, voice description information, image reference information, and video reference information; Perform material requirements analysis on the material information to be collected and the material project information to determine the material requirements structure information and material collection weights; Using the material acquisition weights, the acquisition path planning analysis is performed on the material demand structure information to obtain a data acquisition scheme and data acquisition parameters; The data acquisition scheme and data acquisition parameters are used to acquire image and video materials to obtain initial material data; Multimodal data analysis was performed on the initial data of the materials to determine the material qualification rate of the initial data. Using the material qualification rate, the data acquisition scheme and data acquisition parameters are adjusted before collecting image and video materials to obtain material data containing images and videos.

2. The method for collecting image and video materials based on artificial intelligence as described in claim 1, characterized in that, In the step of performing material demand analysis on the material information to be collected and the material project information to determine the material demand structure information and the material collection weight, the step of determining the material demand structure information includes: Using a preset multimodal fusion model, an initial analysis of the material demand structure is performed on the material information to be collected and the material project information to determine the original information of the material demand structure; Based on the original information of the material demand structure, determine the pass rate of image information and the pass rate of video information; The preset multimodal fusion model is adjusted using the image information pass rate and video information pass rate to obtain the adjusted preset multimodal fusion model. Using the adjusted preset multimodal fusion model, the material demand structure analysis is performed on the material information to be collected and the material project information to determine the material demand structure information.

3. The method for collecting image and video materials based on artificial intelligence as described in claim 2, characterized in that, Based on the original information of the material demand structure, the steps for determining the pass rate of image information and the pass rate of video information include: Extract the image data requirement information and video data requirement information from the original information of the material requirement structure; The completeness of the image data requirement information and the video data requirement information is checked separately to obtain the image data requirement information and video data requirement information that pass the inspection. Data analysis was performed on the qualified image data requirements and video data requirements to obtain the image information qualification rate and video information qualification rate.

4. The method for collecting image and video materials based on artificial intelligence as described in claim 2, characterized in that, The steps for adjusting the preset multimodal fusion model using the image information pass rate and video information pass rate to obtain the adjusted preset multimodal fusion model include: Using the image information pass rate and video information pass rate, the preset multimodal fusion model is adjusted to obtain the initially adjusted preset multimodal fusion model; The initially adjusted preset multimodal fusion model is validated to obtain a validated adjusted preset multimodal fusion model; If the initial adjusted preset multimodal fusion model fails the verification, the pass rates of image information and video information are re-determined before adjusting the preset multimodal fusion model again to obtain the initial adjusted preset multimodal fusion model.

5. The method for collecting image and video materials based on artificial intelligence as described in claim 1, characterized in that, The steps for using the material acquisition weights to perform acquisition path planning analysis on the material demand structure information to obtain the data acquisition scheme and data acquisition parameters include: Based on the aforementioned material demand structure information, all material data sources are determined; Using the weights of all the material collections, each material data source is prioritized to obtain a sorted material data source. Using each sorted material data source, a data acquisition path planning analysis is performed on the material demand structure information to obtain a data acquisition scheme and data acquisition parameters.

6. The method for collecting image and video materials based on artificial intelligence as described in claim 5, characterized in that, The steps for prioritizing each material data source using the collection weights of all the materials to obtain the sorted material data source include: Determine the proportion of each of the aforementioned material data sources in the total data volume of all material data sources for each data source; By using the weights of all the collected materials and the proportion of each data source, each of the material data sources is prioritized to obtain the sorted material data sources.

7. The method for collecting image and video materials based on artificial intelligence as described in claim 1, characterized in that, The steps for performing multimodal data analysis on the initial material data to determine the material qualification rate of the initial material data include: Multimodal data analysis was performed on the initial material data to determine each item type and each complete tag of the initial material data; Perform pass rate analysis on each item type of the initial data of the material, and determine the pass rate threshold for each item type; Based on the pass rate threshold for each project type and each complete tag of the initial material data, the material pass rate of the initial material data is determined.

8. The method for collecting image and video materials based on artificial intelligence as described in claim 7, characterized in that, The steps for determining the material pass rate of the initial material data based on the pass rate threshold for each project type and each complete tag of the initial material data include: Each complete tag of the initial material data is matched with each project type to obtain all complete tags corresponding to each project type; Compare all complete tags corresponding to each project type with the pass rate threshold for each project type to determine the pass rate of the initial material data.

9. The method for collecting image and video materials based on artificial intelligence as described in claim 1, characterized in that, The steps to obtain user-inputted information about the materials to be collected and the material project information include: Obtain the original information of the materials to be collected and the original information of the material items input by the user; The original information of the materials to be collected and the original information of the material items are preprocessed to obtain the preprocessed original information of the materials to be collected and the original information of the material items. The preprocessed original information of the materials to be collected and the original information of the material items are verified to obtain the information of the materials to be collected and the information of the material items.

10. An artificial intelligence-based image and video material collection system, used in conjunction with the artificial intelligence-based image and video material collection method according to any one of claims 1 to 9, characterized in that, include: The information acquisition module is used to acquire the information of the materials to be collected and the information of the materials project input by the user. The information of the materials to be collected includes text description information, voice description information, image reference information and video reference information. The information determination module is used to perform material demand analysis on the material information to be collected and the material project information, and to determine the material demand structure information and material collection weight; The planning and analysis module is used to perform data acquisition path planning and analysis on the material acquisition weights to obtain data acquisition schemes and data acquisition parameters. The data acquisition module is used to acquire image and video materials according to the data acquisition scheme and data acquisition parameters to obtain initial material data; The data analysis module is used to perform multimodal data analysis on the initial data of the material to determine the material qualification rate of the initial data. The data adjustment module is used to adjust the data acquisition scheme and data acquisition parameters based on the material qualification rate before acquiring image and video materials, thereby obtaining material data containing images and videos.