An efficiency improvement method, system, device and medium for obtaining and downloading advertising materials
By automating the segmentation and processing of ad materials through metadata analysis and high-dimensional feature extraction, the method addresses the inefficiencies of manual handling, improving speed and flexibility in ad material reuse and management.
Patent Information
- Application Number
- CN202411006214.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-07-24
AI Technical Summary
The acquisition and processing of creative materials in the prior art lacks automation and intelligence, which leads to cumbersome and inefficient manual segmentation, and it is difficult to quickly and accurately segment and reorganize material elements.
By analyzing the metadata information of the material, determining whether the material meets the requirements and downloading it, and segmenting and aggregating the material, including obtaining the target outline, expanding the area based on pixel similarity, extracting high-dimensional feature vectors and mapping them to low-dimensional space, and building a multi-level category tree for clustering.
It realizes intelligent segmentation and management of materials, improves processing speed and reuse rate of materials, supports content-based image retrieval, and improves the convenience and intelligence level of material management.
Smart Images

Figure CN119007204B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital media technology, and in particular to a method, system, device and medium for improving the efficiency of obtaining and downloading advertising materials. Background Art
[0002] With the booming development of the digital advertising industry, the scale of advertising materials has shown explosive growth. During the creative process, designers need to frequently download a large number of advertising materials, which include multimedia files such as pictures, videos, and audios. These materials are usually distributed in various online platforms, cloud storage services, and internal systems.
[0003] After obtaining the materials, traditional material processing methods often lack automation and intelligence, and require manual and cumbersome editing and sorting work, which not only increases labor costs, but also easily leads to errors and omissions during the processing. Especially when dealing with complex materials containing multiple elements, how to quickly and accurately segment the required elements and perform effective recombination and reuse is also an urgent problem to be solved. Summary of the Invention
[0004] The present invention aims to solve at least one of the technical problems existing in the prior art. For this purpose, the present invention proposes a method, system, device and medium for improving the efficiency of obtaining and downloading advertising materials.
[0005] In a first aspect, the present application proposes a method for improving the efficiency of obtaining and downloading advertising materials, including: obtaining a link of a material, parsing metadata information of the material through the link, determining whether the material meets the requirements, if so, downloading the material, and performing segmentation and aggregation on the material;
[0006] Performing segmentation and aggregation on the material includes:
[0007] Obtaining target contours of several elements in a material image;
[0008] According to the obtained target contours, segmenting the material image to obtain several segmented elements; wherein, segmenting the material image includes: starting from a preset seed point, gradually expanding the area based on pixel similarity until the entire target area is covered, to obtain several segmented elements;
[0009] According to several segmented elements, extracting high-dimensional feature vectors of several elements, mapping the high-dimensional feature vectors to a low-dimensional space, and obtaining feature representations of several elements;
[0010] According to the feature representations of several elements, constructing an index structure, clustering the feature vectors, obtaining a multi-level category tree, and storing original element information at leaf nodes of the category tree.
[0011] More specifically, in the above technical solution, segmenting the material image includes:
[0012] Select seed points according to the target contour, obtain the positions of the segmentation targets in the material image, and determine the coordinates of the seed points;
[0013] Obtain the similarity between each pixel and the seed point based on the gray difference value, and calculate the relative difference between each pixel and the seed point in the material image through the similarity;
[0014] Set a similarity threshold to determine whether adjacent pixels are added to the current growing region. If the gray difference value between the gray value of the adjacent pixel and the seed point in the current region is less than the set threshold, the pixel is incorporated into the growing region. If it is detected that the boundary gradient is greater than the set gradient threshold, it is confirmed that the region growth reaches the target edge, and the expansion of the current region is stopped;
[0015] Based on the region after expansion, measure the area. If the area of the region is less than the set minimum area threshold, adjust the segmentation result by merging regions.
[0016] More specifically, in the above technical solution, extracting the high-dimensional feature vectors of several of the elements and mapping the high-dimensional feature vectors to a low-dimensional space to obtain the feature representations of several of the elements includes:
[0017] Identify and feature-encode several of the elements in the material image through the multi-layer structure of a pre-trained convolutional neural network to obtain high-dimensional feature vectors, and the high-dimensional feature vectors represent the fine-grained information of each element in the material image;
[0018] Use a dimensionality reduction algorithm to map the high-dimensional feature vectors to a low-dimensional space to obtain the low-dimensional feature representations of each of the elements, and classify each of the elements through a clustering algorithm according to the low-dimensional feature representations.
[0019] More specifically, in the above technical solution, constructing an index structure according to the feature representations of several of the elements and clustering the feature vectors includes:
[0020] Obtain the original data of the elements through a data acquisition system, extract the feature information of each element from the original data to generate feature vectors; use data preprocessing technology to standardize the feature vectors; determine the parameters of the clustering algorithm according to requirements, and the parameters of the clustering algorithm include the number of categories, the number of iterations, and the convergence condition, cluster the feature vectors to obtain a preliminary clustering result, and adjust the parameters of the clustering algorithm according to the preliminary clustering result.
[0021] More specifically, in the above technical solution, before segmenting and aggregating the material, the material image is processed by Gaussian filtering, and the gradient magnitude and direction of the material image are obtained by using an edge detection algorithm; the gradient magnitude is processed by non-maximum suppression to retain the pixels with the maximum local gradient; through double-threshold processing, strong edges above the high threshold are screened out, and weak edges between the high threshold and the low threshold are connected to obtain the edge connectivity analysis result. If the edge connectivity analysis result shows that there are broken edges, the broken edges are connected through morphological operations.
[0022] More specifically, in the above technical solution, obtaining the link of the material includes: optimizing the multi-threaded programming framework according to the task scheduling algorithm and concurrently executing multiple crawler tasks; broadcasting the crawler tasks to multiple nodes through a distributed task distribution system and allocating computing resources by using a load balancing algorithm.
[0023] Parsing the metadata information of the material includes: obtaining material parameters by parsing the metadata information, where the material parameters include material format, size, resolution, color mode, and frame rate, and using regular expressions to match the link characteristics of materials in different formats to obtain the material type identifier.
[0024] Determining whether the material meets the requirements includes: setting the material limitation conditions for different platforms based on the material parameters and determining whether the material parameters meet the limitation conditions.
[0025] Downloading the material includes: allocating download tasks by using a distributed task scheduling system, controlling the server load by adjusting the concurrency number, and adjusting the download speed according to the network condition.
[0026] More specifically, in the above technical solution, downloading the material further includes: monitoring the network connection status. If the network is interrupted, the current download task is paused and the breakpoint information is recorded. After the network is restored, the breakpoint position is determined according to the breakpoint information, and the download is resumed from the breakpoint position.
[0027] In a second aspect, the present application also proposes an efficiency improvement system for obtaining and downloading advertising materials, including:
[0028] The first processing module: obtaining the link of the material;
[0029] The second processing module: parsing the metadata information of the material;
[0030] The third processing module: determining whether the material meets the requirements;
[0031] The fourth processing module: downloading the material;
[0032] The fifth processing module: segmenting and aggregating the material;
[0033] The segmentation and aggregation of the said material includes:
[0034] Obtain the target contours of several elements in the material image;
[0035] According to the obtained target contours, segment the material image to obtain several segmented elements; wherein, segmenting the material image includes: starting from a preset seed point, gradually expanding the area based on pixel similarity until the entire target area is covered, to obtain several segmented elements;
[0036] According to several segmented elements, extract high-dimensional feature vectors of several elements, map the high-dimensional feature vectors to a low-dimensional space, and obtain feature representations of several elements;
[0037] According to the feature representations of several elements, construct an index structure, cluster the feature vectors, obtain a multi-level category tree, and store the original element information at the leaf nodes of the category tree.
[0038] In a third aspect, the present application also proposes a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in any one of the above is implemented.
[0039] In a fourth aspect, the present application also proposes a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the above is implemented.
[0040] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0041] By intelligently segmenting the material image, this method can automatically identify and extract key elements in the image, avoiding the cumbersome and inefficient manual segmentation. At the same time, the region expansion technology based on pixel similarity can accurately and quickly complete the segmentation task, significantly improving the processing speed;
[0042] The segmented elements can be processed and applied independently, which means that advertisers can use these elements more flexibly for creative combination and re-creation; for example, product pictures, backgrounds, texts and other elements in an advertising material can be extracted separately, and then recombined according to new advertising needs, greatly improving the reuse rate and utilization rate of the material;
[0043] By extracting high-dimensional feature vectors and performing dimensionality reduction on the segmented elements, this method can convert complex image information into a data format that is easy to manage and retrieve. The constructed multi-level category tree not only facilitates the quick search and positioning of specific elements, but also supports content-based image retrieval, further enhancing the convenience and intelligence level of material management. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0045] Figure 1 is a schematic flowchart of a method for improving the efficiency of obtaining and downloading advertising materials provided by an embodiment of the present invention;
[0046] Figure 2 is a schematic structural diagram of a system for improving the efficiency of obtaining and downloading advertising materials provided by an embodiment of the present invention;
[0047] Figure 3 is a schematic structural diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0049] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0050] It should also be understood that the term " / and" as used in the specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0051] As used in the specification and claims of this application, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0052] In addition, in the description of the specification and claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0053] Reference to "one embodiment" or "some embodiments" or the like described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0054] Please refer to Figure 1 , this application proposes an efficiency improvement method for obtaining and downloading advertising materials, including: S101 obtaining the link of the material, S102 parsing the metadata information of the material through the link, S103 determining whether the material meets the requirements, S104 downloading the material if it meets the requirements, and S105 segmenting and aggregating the material;
[0055] Segmenting and aggregating the material includes:
[0056] Obtaining the target contours of several elements in the material image;
[0057] According to the obtained target contours, segmenting the material image to obtain several segmented elements; wherein, segmenting the material image includes: starting from a preset seed point, gradually expanding the region based on pixel similarity until the entire target region is covered, to obtain several segmented elements;
[0058] According to the several segmented elements, extracting the high-dimensional feature vectors of the several elements, and mapping the high-dimensional feature vectors to a low-dimensional space to obtain the feature representations of the several elements;
[0059] Construct an index structure based on the feature representations of a number of the said elements, cluster the feature vectors to obtain a multi-level category tree, and store the original element information at the leaf nodes of the category tree.
[0060] By performing intelligent segmentation on the material image, this method can automatically identify and extract key elements in the image, avoiding the cumbersome and inefficient manual segmentation. At the same time, the region expansion technology based on pixel similarity can accurately and quickly complete the segmentation task, significantly improving the processing speed.
[0061] The segmented elements can be processed and applied independently, meaning that advertising creators can use these elements more flexibly for creative combination and re-creation. For example, the product pictures, backgrounds, texts and other elements in an advertising material can be extracted separately and then recombined according to new advertising requirements, greatly improving the reuse rate and utilization rate of the materials.
[0062] By extracting and dimension-reducing high-dimensional feature vectors from the segmented elements, this method can convert complex image information into a data format that is easy to manage and retrieve. The constructed multi-level category tree not only facilitates the quick search and positioning of specific elements, but also supports content-based image retrieval, further enhancing the convenience and intelligence level of material management.
[0063] In some embodiments, segmenting the material image includes:
[0064] Select seed points according to the target contour, obtain the position of the segmentation target in the material image, and determine the coordinates of the seed points.
[0065] Obtain the similarity between each pixel and the seed point based on the gray difference, and calculate the relative difference between each pixel and the seed point in the material image through the similarity.
[0066] Set a similarity threshold to judge whether adjacent pixels are added to the current growing region. If the gray difference between the gray value of the adjacent pixel and the seed point in the current region is less than the set threshold, the pixel is included in the growing region. If it is detected that the boundary gradient is greater than the set gradient threshold, it is confirmed that the region growth reaches the target edge and the expansion of the current region stops.
[0067] Perform area measurement based on the expanded region. If the area of the region is less than the set minimum area threshold, adjust the segmentation result by merging regions.
[0068] Select appropriate seed points according to the target contour information, obtain the exact positions of the expected segmentation targets in the image, and determine the specific coordinates of the seed points in an automated manner; use the gray difference as a measure of pixel similarity to obtain the similarity between each pixel and the seed points, and through this method, the relative differences between each pixel and the seed points in the image can be accurately calculated; set a similarity threshold to determine whether adjacent pixels are added to the current growing region. If the gray difference between the gray value of the adjacent pixel and the seed point in the current region is less than the set threshold, then include this pixel in the growing region; through connectivity requirements, ensure that each expansion during the region growing process maintains the spatial continuity of the target region and avoid generating isolated pixel blocks; if it is detected that the boundary gradient is greater than the set gradient threshold, determine that the region growing reaches the target edge and stop the expansion of the current region; measure the area of the expanded region. If the area of the region is less than the minimum area threshold, adjust the segmentation result by merging small regions; use boundary smoothing technology to optimize the region edge, make the segmented boundary smoother, and improve the visual effect of the image; if there are multiple target regions, repeat the above steps to perform region growing on each target region separately until all target regions are successfully segmented; merge the segmentation results of each target region to obtain the final image segmentation output, ensuring that each target element such as Logo, text, main body, etc. is accurately segmented.
[0069] Exemplarily, during the process of automatic image segmentation, first, appropriate seed points are automatically identified and selected through image processing algorithms. For example, after preprocessing the image, edge detection algorithms such as the Canny algorithm are used to identify significant edges, and seed points are automatically selected in the low-gradient regions near the edges. The gray value of the seed point is used as a benchmark, and the similarity is evaluated by calculating the gray difference between each pixel point in the image and the seed point. A similarity threshold is set, for example, the gray difference does not exceed 10. If the gray difference of adjacent pixels is less than this threshold, these pixels are considered to belong to the same region as the seed point. Next, using the region growing algorithm, starting from the seed point, adjacent pixels that meet the similarity conditions are gradually added to the growing region. During this process, a connectivity judgment is added to ensure that the newly added pixels are spatially continuous with the original region and prevent the appearance of isolated pixels. When the region expands to the edge, that is, the boundary gradient of adjacent pixels exceeds the set threshold (such as the gradient is greater than 15), the expansion of this region stops because this usually means that it has touched other regions or the target edge. For each completed region growth, its area is detected. If it is found that the area of the region is less than a certain minimum value (for example, the area is less than 50 pixel points), an attempt will be made to re-grow by merging surrounding small regions or adjusting parameters to ensure the accuracy and integrity of the result. In addition, to further improve the visual effect of the image, the region edges will be optimized through boundary smoothing techniques, such as using a Gaussian filter to smooth the boundary, making the segmentation result more natural. Finally, the above process is repeated for each expected segmentation target in the image until all important elements in the entire image, such as Logos, texts, etc., are accurately separated. After merging the segmented target regions, the final image segmentation output is formed, ensuring that each element in the image is clearly identified and displayed. Through this series of automated processes, efficient and accurate image segmentation is achieved.
[0070] In some embodiments, extracting high-dimensional feature vectors of several of the elements and mapping the high-dimensional feature vectors to a low-dimensional space to obtain feature representations of several of the elements includes:
[0071] Identifying and feature-encoding several of the elements in the material image through a multi-layer structure of a pre-trained convolutional neural network to obtain high-dimensional feature vectors, where the high-dimensional feature vectors represent the fine-grained information of each element in the material image;
[0072] Using a dimensionality reduction algorithm to map the high-dimensional feature vectors to a low-dimensional space to obtain low-dimensional feature representations of each of the elements, and classifying each of the elements through a clustering algorithm according to the low-dimensional feature representations.
[0073] According to the project requirements, select a suitable pre-trained convolutional neural network to extract features from the images, and identify and encode the key elements in the images through the multi-layer structure of the network; if the convolutional neural network processing is completed, obtain high-dimensional feature vectors, which represent the fine-grained information of each key element in the image; adopt a dimensionality reduction algorithm, such as principal component analysis (PCA) or t-SNE, to process the high-dimensional feature vectors, so as to map the data into a low-dimensional space; after dimensionality reduction processing, obtain the low-dimensional feature representations of each key element, which capture the most core characteristics of the elements; according to the low-dimensional feature representations, use a clustering algorithm to classify the key elements to identify groups of elements with similar characteristics; if the clustering results are determined, judge the representative features of each group and label their categories for further analysis and application; adopt a supervised learning method, such as support vector machine (SVM) or deep neural network, to train the labeled data to establish an accurate classification model; test the new image data through the trained model to verify the generalization ability and accuracy of the model; according to the test results, adjust the model parameters or optimize the data processing process to improve the effect and reliability of the model in practical applications.
[0074] Exemplarily, in an image recognition project, first, by using a pre-trained convolutional neural network such as ResNet-50 or VGG16, feature extraction is performed on the input image data. These networks effectively identify complex patterns and key elements in the image through deep structures. The output feature vectors usually have a very high dimension, for example, 4096 dimensions. To simplify the model and improve computational efficiency, we adopt the principal component analysis (PCA) algorithm to reduce the dimension of the feature vectors from 4096 to 200. This not only retains most of the variability (e.g., more than 95%), but also significantly reduces the computational complexity. Subsequently, the t-SNE method is used to further map the feature space to a two-dimensional plane for facilitating the observation and analysis of the clustering situation in the dataset. Next, the K-means clustering algorithm is used to classify the low-dimensional feature representations, with the number of clusters set to 10. Through iterative calculations of the algorithm, each cluster center represents a group of image elements with similar features. For each clustering result, we analyze the feature vectors of its center point to identify the representative features of the category and assign corresponding labels such as "architecture", "nature", etc. to it. After that, based on these labeled data, a deep neural network model is constructed for supervised learning, which includes multiple convolutional layers, activation layers, and fully connected layers. When designing, a suitable loss function and optimizer (such as cross-entropy loss and Adam optimizer) are selected, and by setting appropriate learning rates (e.g., 0.01) and batch sizes (e.g., 32), the model is trained. After the model training is completed, an independent test set is used to verify the generalization ability and classification accuracy of the model, and it is evaluated through indicators such as accuracy and recall. Finally, according to the feedback of the test results, if it is found that the model performs poorly in the recognition of certain categories, it may be necessary to return to the data preprocessing or model parameter adjustment stage, such as increasing the diversity of data samples, adjusting the network structure, or optimizing the algorithm parameters. Through these steps, not only can the accuracy of the model be improved, but also its performance and reliability in practical applications can be optimized.
[0075] In some embodiments, according to the feature representations of several of the said elements, an index structure is constructed. Clustering the feature vectors includes:
[0076] Obtaining the raw data of the said elements through a data acquisition system, extracting the feature information of each element from the raw data to generate feature vectors; using data preprocessing techniques to perform normalization processing on the feature vectors; determining the parameters of the clustering algorithm according to requirements, where the parameters of the clustering algorithm include the number of categories, the number of iterations, and the convergence condition, clustering the feature vectors to obtain a preliminary clustering result, and adjusting the parameters of the clustering algorithm according to the preliminary clustering result.
[0077] The original data of elements is obtained through the data acquisition system, the characteristic information of each element is extracted, and the characteristic vector is generated; the data preprocessing technology is used to standardize the characteristic vector to ensure the effect of the clustering algorithm; the parameters of the K-means clustering algorithm, such as the number of categories, the number of iterations, and the convergence condition, are determined according to business needs; the K-means algorithm is used to cluster the standardized characteristic vector to obtain preliminary clustering results; the clustering results are analyzed, the clustering parameters are adjusted, and the clustering results are optimized to ensure the rationality of the categories and the accuracy of the clustering; a category tree structure is created to organize the clustering results into a tree structure according to the hierarchical relationship, with the upper layer representing the broad categories and the lower layer representing the specific subclasses; the cluster center and the number of elements contained in each tree node are obtained as the statistical information of the node to assist in subsequent indexing and querying; the original element information of the corresponding category is stored in the leaf node of the category tree to provide detailed data support; an element update strategy is designed, and when a new element is added, the category to which it belongs is recalculated according to its characteristic vector, and the category tree and node statistics are updated.
[0078] Exemplarily, in the data acquisition system, the raw data of the elements of the equipment are first collected through sensors and interface protocols. These data include temperature, humidity, vibration and other sensor data. The Pandas library in the Python programming language is used to organize and clean these data, extract key features such as average temperature, maximum amplitude, etc., and calculate their feature vectors. In order to ensure the fairness of the comparison between features, the StandardScaler method in the Skl earn library is used to standardize the feature vectors. This process makes the mean of each feature 0 and the standard deviation 1. Next, according to business needs, it is determined to use the K-means clustering algorithm to classify the data. The number of categories is set to 5, the number of iterations is set to 300, and the convergence condition is that the moving distance of the cluster center is less than 0.01. The KMeans module in the Skl earn library is used to implement this clustering. After clustering, the data points will be divided into five categories. The initial effect of clustering is evaluated by calculating the average distance between the data points in each category and their cluster centers. If it is found that the cohesion of some categories is not strong, the number of categories may be adjusted or the iteration parameters may be modified to optimize the clustering results. According to the clustering results, a category tree structure is constructed. The upper layer of the tree represents the broader categories, such as "high temperature", "low vibration", etc., and the lower layer represents the specific subcategories, such as "high temperature-high humidity" and "high temperature-low humidity". Each node stores the cluster center of the category and the number of elements contained. This information is stored and managed through the SQL database to facilitate subsequent query and analysis. When new element data enters through the data acquisition system, the system will automatically recalculate its category based on its feature vector using the nearest cluster center, and update the category tree and node statistics in real time. This dynamic update strategy ensures the real-time and accuracy of data classification, and provides strong data support for subsequent data analysis and business decision-making.
[0079] In some embodiments, before the material is segmented and aggregated, the material image is processed by Gaussian filtering, and the gradient amplitude and direction of the material image are obtained by using an edge detection algorithm; the gradient amplitude is processed by non-maximum suppression, and the pixel points with the largest local gradient are retained; through double threshold processing, strong edges above the high threshold are screened out, and weak edges between the high threshold and the low threshold are connected to obtain edge connectivity analysis results. If the edge connectivity analysis results show that the edges are broken, the broken edges are connected through morphological operations.
[0080] Before performing edge detection on the optimized material image using the Canny operator, the image is processed by Gaussian filtering to reduce the noise in the image and obtain smoother image data. Based on the smoothed image data, the Canny operator is used to calculate the gradient magnitude and direction of the image, enabling the precise extraction of edge information in the image. After obtaining the gradient magnitude, non-maximum suppression is applied to the gradient magnitude to retain the pixels with the maximum local gradient, which further refines the edges. Through double-threshold processing, strong edges above the high threshold are selected, and weak edges between the high and low thresholds are connected to obtain preliminary edge connectivity. If the edge connectivity analysis shows that there are breaks or overly narrow areas in the edges, morphological operations such as dilation are used to connect these broken edges to make the edges more complete. Erosion operations are used to remove small irrelevant edge fragments that may be generated due to dilation, making the edges clearer and more accurate. The processed edge data is obtained, and the complete target contour is extracted based on the connectivity analysis. Further analysis is performed on the extracted target contour to determine features such as the shape and size of the target, obtaining a detailed description of the target contour. Based on the above complete target contour information, the targets in the image are finally confirmed and classified to obtain the final image processing result.
[0081] Exemplarily, in image processing, first, the input sketch image is smoothed by Gaussian filtering. Assuming a 5x5 Gaussian kernel is used with a standard deviation set to 4, this can effectively reduce the noise interference in the image. Next, the Canny operator is used to calculate the gradient magnitude and direction of the image. The gradient magnitude is usually calculated by the Sobel operator, which helps to accurately capture the edge information in the image. After the calculation, non-maximum suppression is performed. This step will check each pixel point in the image and only retain the points with the local maximum gradient value, and other non-local maximum points will be suppressed to refine the edges. Subsequently, the double-threshold technique is used to determine the true edges. Specifically, the points with a gradient value higher than the high threshold (such as 09) are marked as strong edges, and the points with a gradient value between the low threshold (such as 05) and the high threshold are marked as weak edges. Only when the weak edges are connected to the strong edges will these weak edges be finally retained. This step can effectively connect the edge information and enhance the edge connectivity of the image. If there are still problems such as broken or too narrow edges, morphological operations such as dilation can be used to strengthen the connectivity of these edges. Generally, a 3x3 square structuring element is used for dilation. After dilation, in order to avoid unnecessary noise caused by excessive edge expansion, an erosion operation is also required. Usually, an erosion is performed using the same-sized structuring element as dilation, which helps to remove the redundant edge segments and make the edges clearer and more accurate. Finally, through these series of operations, the target contour in the image is clearly extracted. Further, the shape, size, and other information of the target can be determined through contour feature analysis, providing a basis for the final image classification. This series of processing procedures fully ensures the accuracy of image edge detection and the integrity of the target contour, providing a solid foundation for subsequent image analysis and understanding.
[0082] In some embodiments, obtaining the link of the material includes: optimizing the multi-threaded programming framework according to the task scheduling algorithm, concurrently executing multiple crawler tasks to improve the information collection efficiency; broadcasting the crawler tasks to multiple nodes through a distributed task distribution system, and using a load balancing algorithm to allocate computing resources to improve the parallel processing ability and overall anti-pressure of the system, and obtaining more efficient and stable information collection and processing performance;
[0083] Parsing the metadata information of the material includes: obtaining the material parameters by parsing the metadata information, where the material parameters include the material format, size, resolution, color mode, and frame rate, and using regular expressions to match the link characteristics of materials in different formats to obtain the material type identifier;
[0084] Judging whether the material meets the requirements includes: based on the material parameters, setting the material limit conditions for different platforms, and judging whether the material parameters meet the limit conditions;
[0085] Downloading the said materials includes: allocating download tasks using a distributed task scheduling system, controlling server load by adjusting the concurrency number, and adjusting the download speed according to the network condition.
[0086] In some embodiments, downloading the said materials further includes: monitoring the network connection status, if the network is interrupted, pausing the current download task and recording breakpoint information, after the network resumes, determining the breakpoint position according to the said breakpoint information, and resuming the download from the breakpoint position.
[0087] Optionally, optimizing the multi-threaded programming framework according to the task scheduling algorithm, and concurrently executing multiple crawler tasks includes: dynamically scheduling multiple crawler tasks using the minimum heap scheduling algorithm according to the priority and resource requirements of the tasks; by constructing a task priority queue, putting tasks with high priority and small resource requirements at the head of the queue, taking out the task at the head of the queue from the task priority queue, and allocating it to an idle thread for execution, if there is no idle thread, re-adding the task to the priority queue; when the crawler task is executed, controlling the number of concurrent threads through the semaphore mechanism to prevent the number of threads from exceeding the system resource limit; dynamically adjusting the upper limit of the number of threads for each task according to the web crawling depth and the number of web pages of the task; during the execution of the crawler task, adopting the producer-consumer mode to coordinate the concurrent operations of the URL queue and web page parsing; the URL queue is concurrently added with URLs by multiple producer threads, and the consumer thread takes out the URLs from the queue and concurrently requests the web page content; web page content parsing adopts the finite state machine model, splitting the web page parsing into multiple states, each state is executed concurrently, during the parsing process, controlling the mutual exclusion access of shared resources through the monitor model to ensure thread safety; storing the parsed structured data into a concurrent queue, and then concurrently writing it into the database by multiple consumer threads; database writing adopts the connection pool mechanism to reuse database connections and reduce connection creation overhead; regularly checking the execution progress of the crawler task through a timer, and performing timeout processing on tasks that have not been completed for a long time; setting the retry interval in combination with the exponential backoff algorithm according to the task failure reason, and re-adding the task to the priority queue.
[0088] Specifically, during the crawler task scheduling process, the minimum heap algorithm is used to construct a task priority queue; the priority value range is from 0 to 10, and the smaller the value, the higher the priority; at the same time, considering the resources required by the task, such as the number of threads and memory occupancy, etc., the resource requirements are quantified into values from 0 to 1, and the smaller the value, the less the resource requirements; the weighted sum of the task priority and resource requirements is used as the comprehensive priority of the task, and the weights are 7 and 3 respectively; the task with the smallest comprehensive priority is maintained at the head of the queue through the minimum heap algorithm; the system sets the threshold of the number of idle threads to 20% of the total number of threads. When the number of idle threads is higher than the threshold, the task at the head of the task queue is taken out and assigned to the idle threads for execution; when the crawler task is executed, according to the concurrent request limit of the website and the server performance, the maximum number of threads for each task is set; the semaphore mechanism is used to control the number of concurrent threads not to exceed the maximum value; for example, if a task sets the maximum number of threads to 10 and there are currently 6 threads executing, the semaphore allows 4 new threads to be created; during the task execution process, the URL queue is concurrently added by 5 producer threads, and 10 consumer threads concurrently request web pages; the web page parsing adopts a finite state machine model, and the parsing process is divided into states such as tag matching, attribute extraction, and content extraction, and each state is concurrently executed by multiple threads; the parsing results are stored in a blocking queue with a capacity of 1000; 3 consumer threads concurrently take data from the queue and write it into the MySQL database, and the database connection is reused through a connection pool; the initial size of the connection pool is 10, the maximum is 20, and the connection is automatically released when the idle time exceeds 5 minutes; the system checks the task execution progress every 10 minutes and performs timeout processing on tasks that have not been completed for more than 30 minutes; the timeout tasks are classified according to the failure reasons, such as server unresponsiveness, network interruption, etc., and the retry intervals are set to 5 minutes and 10 minutes respectively; the retry interval adopts an exponential backoff algorithm, and the interval time between each retry failure doubles until the upper limit of 1 hour is reached; the timeout tasks are re-added to the priority queue and wait for scheduling and retry; through the above multi-level scheduling strategy, while ensuring the timely execution of high-priority tasks, the resource utilization rate and task throughput of the crawler are improved.
[0089] Optionally, the crawler task is broadcast to multiple nodes through a distributed task distribution system. The load balancing algorithm for allocating computing resources includes: According to the design principle of the distributed task distribution system, the crawler task is broadcast to multiple nodes using the distributed hash algorithm, and the tasks are evenly distributed to each node through the consistent hashing algorithm. The computing resource usage of the nodes is obtained to determine whether there is a load imbalance. If there is a load imbalance, the least connection number algorithm is used to dynamically adjust the task distribution strategy, and the tasks are reallocated to the nodes with lower load until the load of all nodes is balanced. During the task execution process, the heartbeat detection algorithm is used to monitor the running status of each node in real time, obtain the resource usage of the node such as CPU, memory, network bandwidth, etc., and determine whether there is a node failure or resource bottleneck; if a node failure is found, the failover algorithm is used to automatically migrate the tasks on the failed node to other normal nodes to ensure the continuous execution of the tasks; if a resource bottleneck is found, the vertical scaling algorithm is used to dynamically adjust the computing resource configuration of the node, increase the allocation of resources such as CPU and memory, and improve the processing ability of the node. During the data collection process, the incremental crawling algorithm is used to only crawl new or updated pages, avoid repeated collection, and reduce data redundancy. The URL is de-duplicated through the Bloom filter algorithm to avoid repeated crawling. After obtaining the page content, the regular expression algorithm is used to extract key information, and the natural language processing algorithm is used for text tokenization, part-of-speech tagging, named entity recognition, etc. to obtain structured data. In the data storage stage, the distributed storage algorithm is used to distribute the data to multiple storage nodes, and the consistent hashing algorithm is used to determine the storage location of the data to achieve load balancing of the data. At the same time, the data redundancy backup algorithm is used to perform multi-copy backup of the data to ensure the reliability and availability of the data. Finally, the data analysis algorithm is used to perform statistical analysis on the collected data, the TF-IDF algorithm is used to calculate the weights of keywords, the PageRank algorithm is used to calculate the importance of web pages, the clustering algorithm is used to aggregate similar web pages, and the classification algorithm is used to divide the topics of web pages to obtain valuable information and knowledge.
[0090] Specifically, in a distributed task distribution system, the consistent hashing algorithm is adopted to evenly distribute crawler tasks to each node. For example, when 100 crawler tasks need to be distributed to 5 nodes, by calculating the hash value of the task URL, the hash value is mapped to the circular space from 0 to 2^32 - 1, and then the hash value of the node is also mapped to this space. The task will be distributed to the nearest node in the clockwise direction, thus achieving the even distribution of tasks. The system obtains the resource usage conditions such as CPU, memory, and network bandwidth of the nodes every 5 minutes. When it is found that the CPU usage rate of a certain node exceeds 80%, it is judged that there is load imbalance. The least connection number algorithm is adopted to redistribute some tasks on this node to other nodes with lower loads until the CPU usage rates of all nodes are below 60%. At the same time, the system sends heartbeat packets to each node every 30 seconds through heartbeat detection. If the response of the node is not received continuously for 3 times, it is judged that the node has failed, and the tasks on it are automatically migrated to other normal nodes. When the memory usage rate of the node exceeds 90%, the vertical scaling algorithm dynamically adjusts the memory configuration of the node, increasing the memory from the original 8GB to 16GB to improve the processing ability of the node. During the data collection process, the incremental crawling algorithm only crawls the pages newly added or updated in the recent 24 hours, and the URL is de-duplicated through the Bloom filter, with the de-duplication rate reaching over 95%. The content of the crawled pages extracts key fields through regular expressions, and then uses natural language processing algorithms such as jieba Chinese word segmentation, NLTK part-of-speech tagging, and Stanford NER named entity recognition to process and extract structured data. The data storage adopts the distributed storage system HBase, distributes the data to 5 storage nodes through the consistent hashing algorithm, and adopts a 3-replica backup strategy to ensure the reliability of the data. Finally, the TF-IDF algorithm is used to calculate the keyword weights, the PageRank algorithm is used to calculate the importance of web pages, the K-means clustering algorithm is used to aggregate similar web pages, and the SVM classification algorithm is used to divide the topics of web pages to mine valuable information and knowledge from the massive data.
[0091] Parsing the metadata information of the material specifically includes: obtaining the response data of the link of the material through an HTTP request, analyzing the Content-Type field in the response header information to determine the basic file format of the material. If the Content-Length field is included in the response header information, directly read the value of this field as the size of the material. If there is no such field, the size needs to be obtained through the file properties after downloading the file; after obtaining the file, use a file parsing library to read the file header information to obtain the specific format of the material, such as whether it is a picture or a video. According to the determined material format, select the corresponding parsing tool or library to process the file. If the material is a picture, use an image processing library to obtain the resolution and color mode. If it is a video, use a video processing library to obtain the resolution and frame rate; after the image processing library analyzes the picture file, obtain the resolution and color mode data of the picture, and after the video processing library analyzes the video file, obtain the resolution and frame rate data of the video; by analyzing the suffix name of the file link and the file format information parsed above, use regular expression matching rules to verify and determine the final type of the material. If the file format is inconsistent with the link suffix name, perform an in-depth match using regular expressions again based on the file header information and possible embedded metadata to finally determine the material type; organize all the obtained metadata (file format, size, resolution, color mode, frame rate, material type) for subsequent data storage or processing processes.
[0092] Exemplarily, during the process of analyzing and processing the material file, first, an HTTP request is sent to the link of the material to obtain the response data. At this stage, by analyzing the Content-Type field in the response header, we can initially judge the basic format of the file, such as "image / jpeg" or "video / mp4". At the same time, if the Content-Length field is included in the response header, we can directly obtain the size of the file; if this field is missing, it is necessary to download the file and then read the actual size of the file through the API of the file system. Next, a file parsing library such as libmagic is used for further analysis. By reading the header information of the file, libmagic can determine the specific format of the file and distinguish whether it is a picture or a video. For picture files, we use an image processing library such as Pillow to read its resolution and color mode. For example, Pillow can provide the width and height data of the picture, as well as information such as color depth. For video files, a video processing library such as FFmpeg is used to obtain the resolution and frame rate information of the video. FFmpeg can provide detailed encoding information of the video stream, including the number of frames transmitted per second and the video resolution. In addition, by analyzing the suffix name of the file link and combining the analysis results of the actual content of the file, we use regular expressions to verify the type of the material. If the actual type of the file does not match the link suffix name, it may be necessary to further analyze the file header information and the embedded metadata to ensure the accuracy of the file type. For example, for a file with an extension of.jpg, if it is actually analyzed by libmagic as a png format, this situation needs to be further verified. Finally, all the analyzed data - including file format, size, resolution, color mode, and frame rate - are sorted out and stored for subsequent data processing and applications. This coherent analysis and processing process ensure our full understanding of the material and provide accurate data support for subsequent use.
[0093] Determining whether the material meets the requirements specifically includes: obtaining material parameters, including material type, format, size, dimensions, duration, and content; setting corresponding material type restriction conditions according to the requirements of different platforms. For example, a certain platform only accepts images in jpg and png formats. If the material type meets the platform requirements, continue to check whether the material format matches. If the material format meets the requirements, then verify whether the file size is within the range allowed by the platform. If the file size verification passes, further check whether the dimensions of the material meet the platform's resolution standard. If the material dimensions meet the standard, check whether the duration of the material meets the playback time limit for videos or audios. If the material duration also meets the requirements, finally check whether the material content contains any illegal or sensitive information. If there is no illegal information in the material content, conduct a final evaluation of the material to determine whether all conditions are met. If all conditions are met, mark the material as in a downloadable state; otherwise, mark it as not meeting the requirements.
[0094] Exemplarily, in an automated material review system, first, perform type recognition on the uploaded material file. For example, determine whether it is an image or video file by the file extension or MIME type. Suppose a specific social media platform only accepts images in JPG and PNG formats. The system will use regular expressions or file header analysis methods to verify the file format. If the file is in JPG or PNG format, the system will continue to check whether the file size meets the limit. For example, the platform may limit the image size to no more than 5MB. This step can obtain the file size by reading the file properties and compare it with the platform standard. Next, the system will analyze the dimensions of the image to ensure that it meets the platform's resolution requirements, such as a minimum resolution of 1920x1080 pixels. This can be achieved by using image processing libraries such as PIL or OpenCV to obtain and verify the image dimensions. If the image dimensions meet the requirements, for video or audio files, the system also needs to check whether their duration meets the playback time limit specified by the platform, such as no more than 10 minutes. This information can be obtained by reading and analyzing the metadata of the media file through media processing tools such as FFmpeg. Finally, the system will scan the material content to identify and filter out materials containing illegal or sensitive information. This can be achieved by integrating content recognition APIs such as Google Cloud Vision API or self-developed machine learning models to automatically identify and mark materials involving inappropriate content. If the material passes all the aforementioned checks, the system will mark the material as in a downloadable state; if any one of the checks fails, the material will be marked as not meeting the requirements. Through this process, the platform can ensure that all materials uploaded by users meet its publication standards, while automatically processing a large amount of data, improving efficiency and accuracy.
[0095] Downloading the said materials specifically includes: Obtaining material information: According to the list of materials marked as downloadable, extract the ID, download address, and priority of each material, which are used to initialize the task queue; Creating a task queue: Use a task management system to sort the extracted material information by priority and store it in the task queue for assignment; Assigning download tasks: Extract download tasks from the task queue and, through a distributed task scheduling system, assign tasks according to the current load conditions of each download node; Monitoring server load: Dynamically monitor the CPU, memory, and disk I / O usage of each download node to facilitate adjusting the number of concurrent download tasks; Adjusting the concurrency: Dynamically adjust the concurrent download numbers of each download node according to the monitored server load data to prevent overload; Adaptive download speed: Automatically adjust the speed of ongoing download tasks according to real-time network bandwidth and latency data; Tracking download status: Record the progress and status of each download task, and the tracking information includes the amount of data downloaded and the download speed; Exception handling mechanism: Automatically retry or record the error status for tasks that encounter download failures, timeouts, or non-existent resources, and trigger the alarm system; Retrying failed tasks: Obtain the task information of failed downloads from the record and re-join the task queue for re-scheduling according to the original priority.
[0096] Exemplarily, first, through the database query statement "SELECT * FROM material library WHERE download status = downloadable", the ID, download address, and priority information of each material are obtained. Using this data, we can initialize the task queue using a min-heap data structure to ensure that high-priority download tasks are processed first. When implementing specifically, the `heapq` library in Python can be used to manage this priority queue, where each task is a tuple (priority, material ID, download address). Subsequently, according to the distribution of download tasks, the system will call a distributed task scheduling system, such as Celery, in cooperation with RabbitMQ as the message broker, to dynamically allocate tasks to each download node. When allocating download tasks, the current CPU and memory usage of each node will be considered, and these metrics can be obtained through an integrated monitoring tool such as Prometheus to ensure that tasks are not allocated to nodes with high loads. Then, through the Prometheus monitoring system, the CPU, memory, and disk I / O usage of each node are obtained in real time, and the number of concurrent download tasks on each node is dynamically adjusted based on this data. For example, if the CPU usage of a node exceeds 80%, the concurrent download count of that node will be reduced; if the usage is below 30%, the concurrent count can be appropriately increased. This adjustment strategy can be achieved by writing corresponding threshold trigger rules. The system will also monitor the real-time network bandwidth and latency, and use TCP congestion control algorithms such as Cubic or BBR to adaptively adjust the download speed to optimize download efficiency and stability. At the same time, the progress and status of each download task will be continuously tracked, and information including the amount of downloaded data and the current download speed will be recorded. The task status can be updated in real time by writing database update statements. Finally, when a download task fails or times out, the system will automatically mark these tasks and retry them under appropriate conditions. For example, failed tasks can be recorded in a specific database table, and according to the original priority, these tasks will be added back to the download queue at regular intervals (such as 30 minutes). The retry mechanism may adopt an exponential backoff strategy to avoid the burden on the system caused by frequent retries. At the same time, the system will send an alarm through an integrated alarm system, such as Alertmanager, when the number of download failures exceeds the threshold, to promptly notify the system administrator to handle abnormal situations.
[0097] Use a network monitoring tool to continuously detect the network connection status. If it is found that the network connection status is unstable or interrupted, proceed to the next step; obtain the progress information and relevant parameters of the current download task, such as the download address and the local save path. The progress information and parameters will be used in the subsequent steps; according to the obtained progress information, pause the current download task and record these breakpoint information in the local database or file; set a scheduled task to check the network status at regular intervals to determine whether the network has recovered; if the network detection result shows that the network has recovered, retrieve the download task information saved at the last pause from the database or file through the saved breakpoint information; according to the obtained download task information, restart the download task from the breakpoint; after the download is completed, call the MD5 algorithm to perform an integrity check on the downloaded file; if the file check result shows that the file integrity is not damaged, complete this download task; if the file check result shows that the file is damaged, restart the entire download task according to the original download address and repeat the above steps until the file check is successful.
[0098] Exemplarily, when implementing the network status monitoring and automatic download management system, first, network monitoring tools such as Wireshark are deployed to continuously monitor TCP / IP connections. It is set that if the response time of three consecutive Ping commands exceeds 500 milliseconds or the packet loss rate reaches 20%, the network is regarded as unstable. At this time, the system will automatically call the download management API to obtain the progress of the current download task, such as the ratio of the downloaded file size to the total size, as well as information such as the URL of the download resource and the local storage path of the file. These information are saved in the local SQLite database in JSON format, which is convenient for quick retrieval and update of the information for resume from breakpoint. Subsequently, the system sets a Cron scheduled task to automatically execute the network status detection script every 10 minutes. Once the network status is restored (that is, the Ping response time is less than 200 milliseconds and there is no packet loss), the system reads the download task information before the interruption from the database, and requests the server to continue transmitting data from the last interruption point through the HTTP Range Requests header information to implement the resume from breakpoint function. After the download is completed, the system will use the MD5 algorithm to verify the integrity of the file. MD5 is a commonly used hash function that can generate a 128-bit (16-byte) hash value to ensure data integrity. The system calculates the MD5 value of the downloaded file and compares it with the MD5 value provided by the server. If the two are consistent, it means the file is not damaged and the download task is successfully completed. If they are inconsistent, it indicates that the file may have been damaged during transmission, and the system will automatically restart the download from the original download URL until the file passes the MD5 verification. Through this series of automated processes, not only the success rate of the download task is improved, but also the data loss caused by network problems and the manual intervention of users are reduced, enhancing the robustness and user experience of the entire download management system. This method is applicable to scenarios such as large file downloads and software update package downloads, and can effectively solve the problem of download interruption caused by network instability.
[0099] Optionally, the downloaded materials are uniformly saved to the location specified by the user for easy management.
[0100] Please refer to Figure 2 , this application also proposes an efficiency improvement system for obtaining and downloading advertising materials, including:
[0101] The first processing module 201: Obtain the link of the material;
[0102] The second processing module 202: Parse the metadata information of the material;
[0103] The third processing module 203: Judge whether the material meets the requirements;
[0104] The fourth processing module 204: Download the material;
[0105] The fifth processing module 205: segment and aggregate the said material;
[0106] Segmenting and aggregating the said material includes:
[0107] Obtain the target contours of several elements in the material image;
[0108] Segment the material image according to the obtained target contours to obtain several segmented elements; wherein, segmenting the material image includes: starting from a preset seed point, gradually expanding the area based on pixel similarity until the entire target area is covered, to obtain several segmented elements;
[0109] Extract the high-dimensional feature vectors of several elements according to the several segmented elements, map the high-dimensional feature vectors to a low-dimensional space, and obtain the feature representations of several elements;
[0110] Construct an index structure according to the feature representations of several elements, cluster the feature vectors to obtain a multi-level category tree, and store the original element information at the leaf nodes of the category tree.
[0111] It can be understood that, as Figure 1 shown in the content of the embodiment of the method for improving the efficiency of obtaining and downloading advertisement materials, it is applicable to the embodiment of the system for improving the efficiency of obtaining and downloading advertisement materials. The specific functions realized by the embodiment of the system for improving the efficiency of obtaining and downloading advertisement materials are the same as those in Figure 1 the embodiment of the method for improving the efficiency of obtaining and downloading advertisement materials shown, and the beneficial effects achieved are the same as those in Figure 1 the embodiment of the method for improving the efficiency of obtaining and downloading advertisement materials shown.
[0112] It should be noted that the information interaction, execution process, etc. between the above systems, because they are based on the same concept as the method embodiment of the present invention, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and will not be elaborated here.
[0113] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0114] Please refer to Figure 3 , an embodiment of the present invention further provides a computer device 3, including: a memory 302, a processor 301, and a computer program 303 stored on the memory 302. When the computer program 303 is executed on the processor 301, it implements the method for improving the efficiency of obtaining and downloading advertisement materials as described in any one of the above methods.
[0115] The computer device 3 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device 3 may include, but is not limited to, a processor 301 and a memory 302. Those skilled in the art can understand that Figure 3 This is only an example of the computer device 3 and does not constitute a limitation on the computer device 3. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.
[0116] The so-called processor 301 may be a central processing unit (CPU), and the processor 301 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0117] In some embodiments, the memory 302 may be an internal storage unit of the computer device 3, such as a hard disk or memory of the computer device 3. In other embodiments, the memory 302 may also be an external storage device of the computer device 3, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 3. Further, the memory 302 may also include both the internal storage unit and the external storage device of the computer device 3. The memory 302 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as program codes of the computer program. The memory 302 may also be used to temporarily store data that has been output or will be output.
[0118] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the method for improving the efficiency of obtaining and downloading advertisement materials as described in any one of the above methods.
[0119] In this embodiment, if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.
[0120] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. An efficiency improvement method for obtaining and downloading advertising materials, characterized in that, Including: Obtain the link of the material, parse the metadata information of the material through the link, determine whether the material meets the requirements, if it meets, download the material, and perform segmentation and aggregation on the material; Performing segmentation and aggregation on the material includes: Obtain the target contours of several elements in the material image; Segment the material image according to the obtained target contours to obtain several segmented elements; wherein, segmenting the material image includes: starting from a preset seed point, gradually expanding the region based on pixel similarity until the entire target region is covered, to obtain several segmented elements; Extract the high-dimensional feature vectors of several elements according to the segmented elements, map the high-dimensional feature vectors to a low-dimensional space, and obtain the feature representations of several elements; Construct an index structure according to the feature representations of several elements, cluster the feature vectors to obtain a multi-level category tree, and store the original element information at the leaf nodes of the category tree; Wherein, extracting the high-dimensional feature vectors of several elements, mapping the high-dimensional feature vectors to a low-dimensional space, and obtaining the feature representations of several elements includes: Identify and feature-encode several elements in the material image through the multi-layer structure of a pre-trained convolutional neural network to obtain high-dimensional feature vectors, and the high-dimensional feature vectors represent the fine-grained information of each element in the material image; Use a dimensionality reduction algorithm to map the high-dimensional feature vectors to a low-dimensional space to obtain the low-dimensional feature representations of each element, and classify each element through a clustering algorithm according to the low-dimensional feature representations.
2. The method for improving the efficiency of obtaining and downloading advertising materials according to claim 1, wherein Segmenting the material image includes: Select a seed point according to the target contour, obtain the position of the segmentation target in the material image, and determine the coordinates of the seed point; Obtain the similarity between each pixel and the seed point based on the gray difference, and calculate the relative difference between each pixel and the seed point in the material image through the similarity; Set a similarity threshold, determine whether adjacent pixels are added to the current growing region. If the gray difference between the gray value of the adjacent pixel and the seed point in the current region is less than the set threshold, include the pixel in the growing region. If it is detected that the boundary gradient is greater than the set gradient threshold, confirm that the region growth reaches the target edge and stop the expansion of the current region; Perform area measurement based on the expanded region. If the area of the region is less than the set minimum area threshold, adjust the segmentation result by merging regions.
3. The method for improving the efficiency of obtaining and downloading advertising materials according to claim 1, wherein Constructing an index structure according to the feature representations of several elements and clustering the feature vectors includes: Obtain the original data of the element, extract the feature information of each element from the original data, and generate a feature vector; Use data preprocessing techniques to standardize the feature vectors; use a clustering algorithm to cluster the feature vectors to obtain a clustering result, and adjust the parameters of the clustering algorithm according to the clustering result.
4. The method for improving the efficiency of obtaining and downloading advertising materials according to any one of claims 1-3, characterized in that, Before segmenting and aggregating the material, the material image is processed by Gaussian filtering, and the gradient magnitude and direction of the material image are obtained by using an edge detection algorithm; the gradient magnitude is processed by non-maximum suppression to retain the pixels with the maximum local gradient; through double-threshold processing, strong edges above the high threshold are screened out, and weak edges between the high threshold and the low threshold are connected to obtain the edge connectivity analysis result. If the edge connectivity analysis result shows that there are broken edges, the broken edges are connected through morphological operations.
5. The method for improving the efficiency of obtaining and downloading advertising materials according to claim 1, wherein, Obtaining the link of the material includes: optimizing the multi-threaded programming framework according to the task scheduling algorithm and concurrently executing multiple crawler tasks; broadcasting the crawler tasks to multiple nodes through a distributed task distribution system and allocating computing resources by using a load balancing algorithm. Parsing the metadata information of the material includes: obtaining material parameters by parsing the metadata information, where the material parameters include material format, size, resolution, color mode, and frame rate, and using regular expressions to match the link characteristics of materials in different formats to obtain the material type identifier. Judging whether the material meets the requirements includes: based on the material parameters, setting the material limit conditions for different platforms and judging whether the material parameters meet the limit conditions. Downloading the material includes: allocating download tasks by using a distributed task scheduling system, controlling the server load by adjusting the concurrency number, and adjusting the download speed according to the network condition.
6. The method for improving the efficiency of obtaining and downloading advertising materials according to claim 5, characterized in that, Downloading the material further includes: monitoring the network connection status. If the network is interrupted, the current download task is paused and the breakpoint information is recorded. After the network resumes, the breakpoint position is determined according to the breakpoint information, and the download continues from the breakpoint position.
7. An efficiency improvement system for obtaining and downloading advertising materials, characterized in that, Including: The first processing module: obtaining the link of the material; The second processing module: parsing the metadata information of the material; The third processing module: judging whether the material meets the requirements; The fourth processing module: downloading the material; The fifth processing module: segmenting and aggregating the material; Segmenting and aggregating the material includes: Obtaining the target contour of several elements in the material image; According to the obtained target contour, segmenting the material image to obtain several segmented elements; where segmenting the material image includes: starting from a preset seed point and gradually expanding the region based on pixel similarity until the entire target region is covered to obtain several segmented elements; According to the several segmented elements, extracting the high-dimensional feature vectors of the several elements, mapping the high-dimensional feature vectors to a low-dimensional space, and obtaining the feature representations of the several elements; According to the feature representations of the several elements, constructing an index structure, clustering the feature vectors, obtaining a multi-level category tree, and storing the original element information at the leaf nodes of the category tree; Among them, extracting the high-dimensional feature vectors of the several elements, mapping the high-dimensional feature vectors to a low-dimensional space, and obtaining the feature representations of the several elements includes: Identify and feature-encode several of the elements in the material image through the multi-layer structure of a pre-trained convolutional neural network to obtain a high-dimensional feature vector, where the high-dimensional feature vector represents the fine-grained information of each element in the material image; Use a dimensionality reduction algorithm to map the high-dimensional feature vector to a low-dimensional space, obtain the low-dimensional feature representation of each element, and classify each element through a clustering algorithm according to the low-dimensional feature representation.
8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image processing method and apparatus, electronic device, and computer-readable storage medium
CN109360176A
Video target segmentation method and device, storage medium and program product
CN114445750A