Method and system for calculating and automatically selecting the diversity of image subject content

End-to-end feature extraction and similarity calculation are performed through deep learning models, combined with entity and abstract semantic features, the inconsistency and adaptability of the diversity calculation of picture theme content in the prior art is solved, and efficient and accurate picture selection is achieved.

CN116935138BActive Publication Date: 2025-08-08SHANGHAI INTERNATIONAL STUDIES UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310977202.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-04
Publication Date
2025-08-08
Estimated Expiration
2043-08-04

AI Technical Summary

Technical Problem

The existing methods of image theme content diversity calculation and automatic selection have problems such as strong subjectivity, separation of feature extraction and similarity calculations, resulting in inconsistency, difficulty in adapting to new fields and large-scale data processing.

Method used

Deep learning models, especially convolutional neural networks (CNNs), are used to perform end-to-end feature extraction and similarity calculations, and combined entity semantics and abstract semantic features, and comprehensively select pictures through correlation and diversity calculations.

Benefits of technology

The calculation accuracy of the relevance of pictures and topics and content diversity is improved, and the image selection needs of different styles and scenarios are adapted to the needs of picture selection, reducing the needs of manual intervention and feature engineering, and improving selection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935138B_ABST
    Figure CN116935138B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of methods for calculating and automatically selecting the diversity of picture subject content, and specifically to methods and systems for calculating and automatically selecting the diversity of picture subject content, comprising the following steps: data preparation and feature extraction, correlation calculation, obtaining a correlation score, diversity calculation, obtaining a diversity metric, comprehensive scoring and automatic selection. In the present invention, a deep learning model is used to perform end-to-end training using a large amount of data to learn more accurate image feature representations. Feature extraction is performed to obtain feature descriptions with more semantic information, thereby improving the accuracy of calculating the relevance of pictures to themes and content diversity. By combining entity semantic features and abstract semantic features, the semantics of pictures are described more comprehensively. Optimization and improvement are performed by adjusting the network structure, increasing training data, etc. to adapt to specific picture selection requirements. By training on large-scale data, feature representations with more generalization capabilities are learned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of methods for calculating and automatically selecting the diversity of picture subject content, and in particular to a method and system for calculating and automatically selecting the diversity of picture subject content. Background Art

[0002] Methods for calculating and automatically selecting image thematic content diversity calculate and evaluate the diversity of a set of images and automatically select the most appropriate images for a given theme. To calculate diversity, a clustering algorithm is used to group images into categories or clusters to capture differences in their content. This ensures that each category contains images with a certain degree of diversity. Automatic selection methods combine diversity calculations with relevance to a given theme, assigning each image a corresponding score or weight, and determining the final selection result based on this. The scoring function is based on the weighting of relevance and diversity and can be adjusted according to specific needs. These methods provide an automated way to more efficiently and intelligently select images related to a given theme and with high content diversity from a large number of images.

[0003] Existing methods for calculating and automatically selecting image subject content diversity typically employ hand-crafted feature extraction methods, often based on manual experience and heuristic rules. This approach is subject to subjectivity and limitations, failing to fully exploit the semantic information of images. Furthermore, feature extraction and similarity calculation in traditional solutions are typically performed separately, making end-to-end learning and optimization impossible. This separation can lead to feature representation mismatch—inconsistencies between the extracted features and the similarity calculation method—impairing the final selection results. Traditional solutions rely on manually designed features and rules for image selection, often based on specialized domain knowledge and experience. This reliance makes traditional solutions difficult to adapt to new domains or diverse subject matter, requiring significant time and resources for feature engineering and rule design. Due to the limitations of feature representation and reliance on domain knowledge, traditional solutions often fail to fully express the semantic information in complex scenes and large-scale data, and are unable to handle the diversity and variability inherent in large-scale data. Improvements are needed. Summary of the Invention

[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a method and system for calculating and automatically selecting the diversity of picture subject content.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a method for calculating and automatically selecting the diversity of image subject content, comprising the following steps:

[0006] Prepare a picture dataset and perform feature extraction on the picture dataset to obtain a feature vector;

[0007] performing a correlation calculation based on the feature vector to obtain a correlation score;

[0008] Performing diversity calculation based on the feature vector to obtain a diversity metric value;

[0009] The relevance score and the diversity metric are comprehensively scored to obtain a final score, and automatic selection is made based on the final score.

[0010] As a further solution of the present invention, the steps of data preparation are specifically as follows:

[0011] Use the public image library ImageNet to collect or obtain image datasets related to a given topic;

[0012] Make sure the images you choose cover different aspects of a given topic, including different perspectives, scenes, and objects.

[0013] Reduce data redundancy and increase data availability based on preprocessing operations including image scaling, cropping, rotation, normalization, and standardization;

[0014] Data augmentation technology is used to perform random transformations on the original images, including flipping, rotation, translation, scaling, and adjusting brightness and contrast, to generate more training samples and expand the size of the dataset.

[0015] As a further solution of the present invention, the feature extraction step is specifically as follows:

[0016] Select a convolutional neural network (CNN) model to perform feature extraction on the image dataset;

[0017] The images in the image dataset are fed into a convolutional neural network (CNN) model and forward propagation is used. During the forward propagation process, the images are processed through a series of convolution, pooling, and activation functions to gradually extract increasingly abstract and advanced feature representations.

[0018] For mid-level feature extraction, shallower convolutional layers acquire low-level features, including edges and textures, while deeper convolutional layers acquire higher-level features, including object shapes and components. Features from multiple levels are combined to obtain richer information. During feature extraction, an attention mechanism is used to enhance the importance and expressiveness of features, and spatial pyramid pooling techniques are applied to capture multi-scale features.

[0019] The features after the convolution layer are reduced in dimension and compressed through the fully connected layer to obtain a more abstract and advanced feature representation. Classification and detection tasks are performed by extracting the feature vector of the fully connected layer.

[0020] As a further solution of the present invention, the steps of calculating the correlation are specifically as follows:

[0021] Chord similarity is selected as the similarity measure;

[0022] performing similarity calculation on the feature vector of each image in the image dataset and the feature vector of a given theme, wherein during the similarity calculation, a dimension reduction method, specifically principal component analysis, is used to reduce the dimension of the feature vector to improve computational efficiency and reduce the influence of noise;

[0023] Use the relevance score calculation formula to obtain the relevance score, which indicates the relevance of the image to the given topic;

[0024] The correlation score calculation formula is specifically as follows:

[0025] =(t×f)÷||t||||f||

[0026] Where (t×f) represents the inner product of the topic vector t and the feature vector f, ||t|| represents the norm of the topic vector t, and ||f|| represents the norm of the feature vector f.

[0027] The relevance score of each image to a given topic is sorted or classified, where all relevance scores greater than or equal to 0.8 indicate a high relevance to the topic, and all relevance scores less than 0.8 indicate a low relevance to the topic.

[0028] As a further solution of the present invention, the steps of diversity calculation are specifically as follows:

[0029] Calculate the mean of the eigenvectors;

[0030] Calculate the difference vector;

[0031] Calculate the variance and standard deviation.

[0032] As a further solution of the present invention, the mean of the calculated characteristic vector is specifically:

[0033] Among them, mean represents the vector mean, n represents the number of feature vectors, and feature_vector represents a feature vector

[0034] The calculation difference vector is specifically: deviation_vector=feature_vector-mean;

[0035] Among them, deviation_vecto represents the difference vector;

[0036] The difference vector represents the difference between each eigenvector and the mean vector;

[0037] The calculation of variance and standard deviation is specifically as follows:

[0038]

[0039]

[0040] Among them, variance represents variance and standard_deviation represents standard deviation.

[0041] As a further solution of the present invention, the comprehensive scoring step is specifically as follows:

[0042] Normalize the relevance scores and diversity metrics to map their value ranges to a unified range;

[0043] Combining relevance and diversity, a weighted synthesis method is used to perform a weighted summation of the normalized relevance score and diversity metric, balancing relevance and diversity according to the weight setting to obtain a comprehensive score;

[0044] The composite score is applied to each image to obtain the final score for each image.

[0045] As a further solution of the present invention, the automatic selection step is specifically as follows:

[0046] Sort all images according to the final score, arranging them in descending order of score;

[0047] According to the needs, set the threshold to filter out the images that meet the specific requirements. The threshold includes absolute value and relative value;

[0048] Select single or multiple images as the final result according to specific needs.

[0049] The image theme content diversity calculation and automatic selection system is composed of an image acquisition module, an image feature extraction module, an image theme analysis module, an image diversity calculation module, and an image automatic selection module;

[0050] The picture acquisition module includes a data source interface submodule, a keyword search submodule, and a download and cache submodule. The output end of the picture acquisition module is communicatively connected to the input end of the picture feature extraction module;

[0051] The picture feature extraction module includes a low-level visual feature extraction submodule, a high-level semantic feature extraction submodule, and a feature storage and association submodule, and the output end of the picture feature extraction module is communicatively connected to the input end of the picture theme analysis module;

[0052] The picture theme analysis module includes an image label recognition submodule, a text recognition submodule, and a theme description generation submodule. The output end of the picture theme analysis module is communicatively connected to the input end of the picture diversity calculation module.

[0053] The picture diversity calculation module includes a similarity calculation submodule, a clustering grouping submodule, and a diversity index calculation submodule, and the output end of the picture diversity calculation module is communicatively connected to the input end of the picture automatic selection module;

[0054] The automatic picture selection module includes a diversity threshold and weight setting submodule, a representative picture screening submodule, and a sorting and display function submodule.

[0055] As a further solution of the present invention, the data source interface submodule is responsible for interacting with different data sources, including image libraries and API interfaces, to obtain image data;

[0056] The search keyword submodule receives keywords input by the user and is used to retrieve relevant pictures from the data source;

[0057] The download and cache submodule downloads the images obtained from the data source and caches them for subsequent processing and use;

[0058] The low-level visual feature extraction submodule extracts basic features of the image, including color, texture, and shape features;

[0059] The high-level semantic feature extraction submodule uses deep learning and natural language processing technology to extract higher-level semantic features from images, including objects, scenes, and emotions;

[0060] The feature storage and association submodule stores the extracted features in a database or other storage medium, and establishes an association relationship between the features and the image for subsequent retrieval and analysis;

[0061] The image tag recognition submodule uses computer vision technology to identify tags or keywords in the image to describe the content or theme of the image;

[0062] The text recognition submodule converts the text in the image into editable text form through optical character recognition (OCR) technology to facilitate subsequent topic analysis and association;

[0063] The subject description generation submodule generates a description or summary of the subject of the picture based on the features of the picture and the recognition results;

[0064] The similarity calculation submodule calculates the similarity scores between the images by comparing their feature similarities, which are used for subsequent diversity calculation and image selection;

[0065] The clustering submodule clusters images according to their features or themes, placing similar images together to facilitate subsequent diversity calculation and image selection;

[0066] The diversity index calculation submodule calculates the diversity index of the picture set based on the similarity and clustering results, and evaluates the diversity level of the picture set;

[0067] The diversity threshold and weight setting submodule allows users to set the diversity threshold and weight according to their needs to adjust the diversity level of the system when selecting pictures;

[0068] The representative image screening submodule screens out highly representative images in the collection based on diversity calculation and a threshold set by the user;

[0069] The sorting and display function submodule presents the selected images to the user in descending order based on similarity and diversity scores.

[0070] Compared with the prior art, the advantages and positive effects of the present invention are:

[0071] In the present invention, a deep learning model is used to perform end-to-end training using a large amount of data, and a more accurate image feature representation can be learned. By using a convolutional neural network for feature extraction, a feature description with more semantic information can be obtained, thereby improving the accuracy of calculating the relevance of the image to the subject and the diversity of the content. Based on the deep learning model, richer semantic information is learned from the image, including entity semantic features and abstract semantic features. By combining entity semantic features and abstract semantic features, the semantics of the image can be more comprehensively described, and the calculation effect of the relevance to the subject can be improved. Optimization and improvement are performed by adjusting the network structure, increasing training data, etc. to adapt to specific image selection needs. By training on large-scale data, a more generalized feature representation can be learned, which can adapt to the image selection needs of different styles, scenes and themes. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] Figure 1 This is a schematic diagram of the workflow of the method and system for calculating and automatically selecting the diversity of picture subject content proposed by the present invention;

[0073] Figure 2 A flowchart of the steps for calculating and automatically selecting the diversity of picture subject content and system data preparation proposed by the present invention;

[0074] Figure 3The present invention proposes a flowchart of the steps of calculating and automatically selecting the diversity of picture subject content and extracting system features;

[0075] Figure 4 A flowchart of the steps of calculating and automatically selecting the diversity of picture subject content and calculating the system relevance proposed by the present invention;

[0076] Figure 5 The present invention proposes a method for calculating and automatically selecting the diversity of picture subject content and a flowchart of the steps of system diversity calculation;

[0077] Figure 6 A flowchart of the steps of the method for calculating and automatically selecting the diversity of picture subject content and the system comprehensive scoring proposed by the present invention;

[0078] Figure 7 The present invention proposes a method for calculating and automatically selecting the diversity of picture subject content and a flowchart of the steps of automatic selection by the system;

[0079] Figure 8 This is a schematic diagram of the system framework of the method and system for calculating and automatically selecting the diversity of picture subject content proposed by the present invention. DETAILED DESCRIPTION

[0080] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0081] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.

[0082] Example 1

[0083] See also Figure 1 The present invention provides a technical solution: a method for calculating and automatically selecting the diversity of picture subject content, comprising the following steps:

[0084] Prepare image datasets and perform feature extraction on them to obtain feature vectors;

[0085] Perform correlation calculation based on the feature vector to obtain a correlation score;

[0086] Perform diversity calculation based on the feature vector to obtain a diversity metric value;

[0087] The relevance scores and diversity metrics are comprehensively scored to obtain the final score, and the selection is automatically made based on the final score.

[0088] First, a large amount of image data is prepared and organized, and low-level visual features and high-level semantic features are extracted to build a feature library. Then, through correlation calculation, similarities and correlation scores are calculated between images to identify other images that are highly relevant to the user's query or selected topic. Based on this correlation calculation, clustering and diversity index calculation are performed to quantitatively measure the diversity of the image collection. Finally, a comprehensive score is calculated for each image, taking into account both relevance and diversity weights, automatically selecting a representative and diverse image collection. This method can provide accurate, diverse, and representative image selection results, saving users time and effort, improving work efficiency, and meeting user needs.

[0089] See also Figure 2 , the steps of data preparation are as follows:

[0090] Use the public image library ImageNet to collect or obtain image datasets related to a given topic;

[0091] Make sure the images you choose cover different aspects of a given topic, including different perspectives, scenes, and objects.

[0092] Reduce data redundancy and increase data availability based on preprocessing operations including image scaling, cropping, rotation, normalization, and standardization;

[0093] Data augmentation technology is used to perform random transformations on the original images, including flipping, rotation, translation, scaling, and adjusting brightness and contrast, to generate more training samples and expand the size of the dataset.

[0094] First, by using publicly available image libraries such as ImageNet, COCO, and PASCAL VOC, we can collect or acquire large datasets of images related to a given topic, ensuring a wide range of data sources and providing more choices. Second, for topics in specific domains, we can select datasets specifically tailored to that domain, ensuring that the datasets are closely related to the task and scenario, thereby improving the adaptability and performance of the algorithm. Furthermore, by ensuring that the datasets include diverse aspects such as different viewpoints, scenes, and objects, we can improve the system's ability to assess diversity and select images for specific topics. Furthermore, data preprocessing operations such as image scaling, cropping, rotation, normalization, and standardization can be performed to reduce redundancy, increase data availability, and improve the robustness and accuracy of the algorithm. Finally, data augmentation techniques such as flipping, rotation, translation, scaling, and brightness and contrast adjustment can be used to generate more training samples, expand the dataset size, and improve the generalization and robustness of the algorithm.

[0095] See also Figure 3 , the steps of feature extraction are as follows:

[0096] Select the convolutional neural network (CNN) model to extract features from the image dataset;

[0097] The images in the image dataset are input into the convolutional neural network (CNN) model and forward propagation is used. During the forward propagation process, the image is processed through a series of convolution, pooling and activation function operations to gradually extract increasingly abstract and advanced feature representations.

[0098] For mid-level feature extraction, shallower convolutional layers acquire low-level features, including edges and textures, while deeper convolutional layers acquire higher-level features, including object shapes and components. Features from multiple levels are combined to obtain richer information. During feature extraction, an attention mechanism is used to enhance the importance and expressiveness of features, and spatial pyramid pooling techniques are applied to capture multi-scale features.

[0099] The features after the convolution layer are reduced in dimension and compressed through the fully connected layer to obtain a more abstract and advanced feature representation. Classification and detection tasks are performed by extracting the feature vector of the fully connected layer.

[0100] See also Figure 4 , the steps of correlation calculation are as follows:

[0101] Chord similarity is selected as the similarity measure;

[0102] For the feature vector of each image in the image dataset and the feature vector of a given theme, similarity is calculated. During the similarity calculation process, the dimension of the feature vector is reduced using a dimensionality reduction method specifically based on principal component analysis to improve computational efficiency and reduce the impact of noise.

[0103] Use the relevance score calculation formula to obtain the relevance score, which indicates the relevance of the image to the given topic;

[0104] The calculation formula of the correlation score is as follows:

[0105] =(t×f)÷||t||||f||

[0106] Where (t×f) represents the inner product of the topic vector t and the feature vector f, ||t|| represents the norm of the topic vector t, and ||f|| represents the norm of the feature vector f.

[0107] The relevance score of each image to a given topic is sorted or classified, where all relevance scores greater than or equal to 0.8 indicate a high relevance to the topic, and all relevance scores less than 0.8 indicate a low relevance to the topic.

[0108] First, convolutional neural network (CNN) models, such as ResNet, are selected as feature extractors to effectively capture the features of image datasets. Second, through forward propagation calculations, CNN models can gradually extract increasingly abstract and advanced feature representations, including information such as edges, textures, shapes, and object parts, which helps to better understand the image content. During the feature extraction process, selecting convolutional layers of different depths can obtain low-level and high-level features. Combining features from multiple levels can obtain richer information, and applying attention mechanisms can enhance the expressive power of features. At the same time, the use of spatial pyramid pooling technology can capture multi-scale features, improving feature diversity and the ability to perceive objects of different scales. Finally, after the convolutional layer, the features are reduced and compressed through a fully connected layer to obtain a more abstract and advanced feature representation. The feature vectors of the fully connected layer are then extracted to perform classification, detection, and other tasks.

[0109] See also Figure 5 ,Specifically, the diversity calculation uses variance and standard deviation as diversity ,measurement methods. Variance and standard deviation are statistical indicators for measuring the ,degree of discrete data distribution. The larger variance and standard deviation are, the higher the difference and ,diversity between the feature vectors.

[0110] The specific steps of diversity calculation are:

[0111] Calculate the mean of the eigenvectors;

[0112] Calculate the difference vector;

[0113] Calculate the variance and standard deviation.

[0114] The specific calculation of the mean of the eigenvector is,

[0115] Among them, mean represents the vector mean, n represents the number of feature vectors, and feature_vector represents a feature vector

[0116] The specific calculation of the difference vector is: deviation_vector = feature_vector-mean;

[0117] Among them, deviation_vecto represents the difference vector;

[0118] The difference vector represents the difference between each eigenvector and the mean vector;

[0119] The calculation of variance and standard deviation is as follows:

[0120] Among them, variance represents variance and standard_deviation represents standard deviation.

[0121] Variance and standard deviation are statistical metrics used to measure the degree of dispersion in data distributions. They can assess the variability and diversity between feature vectors. The specific steps for calculating diversity include calculating the mean of the feature vectors, the difference vector, and the variance and standard deviation. First, the mean of the feature vectors is calculated by summing all the feature vectors in the image collection and dividing the sum by the number of feature vectors. Next, the difference vectors are calculated by subtracting the mean vector from each feature vector to obtain a vector representing the differences. Finally, the difference vectors are squared to obtain the squared difference vectors. These squared difference vectors are summed and divided by the number of feature vectors to obtain the variance, and the standard deviation is obtained by taking the square root of the variance. Using variance and standard deviation as diversity metrics objectively assesses the diversity of feature vectors, helping the system select image content with greater coverage and representativeness. This helps the system provide a more comprehensive and diverse content selection, making the image selection more diverse and representative.

[0122] See also Figure 6 The steps for comprehensive scoring are as follows:

[0123] Normalize the relevance scores and diversity metrics to map their value ranges to a unified range;

[0124] Combining relevance and diversity, a weighted synthesis method is used to perform a weighted summation of the normalized relevance score and diversity metric, balancing relevance and diversity according to the weight setting to obtain a comprehensive score;

[0125] The composite score is applied to each image to obtain a final score for each image.

[0126] First, the relevance scores and diversity metrics are normalized to map their value ranges to a uniform range, ensuring that they are comparable. Next, a weighted synthesis method is used to combine the normalized relevance scores and diversity metrics, and the importance of relevance and diversity is balanced according to the weight settings to obtain a comprehensive score. Finally, the comprehensive score is applied to each image to measure its quality and relevance to the topic. With this comprehensive score, the system can sort or filter the images and select the images with the higher comprehensive scores as the final selection results. This comprehensive score takes into account relevance and diversity, allowing the system to provide a more comprehensive, diverse and high-quality selection of image content.

[0127] See also Figure 7 , the steps of automatic selection are as follows:

[0128] Sort all images according to the final score, arranging them in order from high to low;

[0129] According to the needs, set the threshold to filter out the images that meet the specific requirements. The threshold includes absolute value and relative value.

[0130] Select single or multiple images as the final result according to specific needs.

[0131] First, all images are sorted based on their final scores, arranged in descending order. This prioritizes images with higher scores by placing them first. Then, thresholds are set based on requirements to filter out images that meet specific criteria. Thresholds can be absolute or relative, allowing images to be filtered to meet specific quality and diversity requirements. Finally, a single or multiple images are selected as the final result based on specific requirements. This automated selection process effectively improves system efficiency and accuracy, automatically selects appropriate images, reduces the need for manual intervention, and allows for flexible customization of system behavior based on specific requirements.

[0132] See also Figure 8 ,The image subject content diversity calculation and automatic selection system is composed of the ,image acquisition module, image feature extraction module, image ,theme analysis module, image diversity calculation module, and image ,automatic selection module;

[0133] The image acquisition module includes a data source interface submodule, a keyword search submodule, and a download and cache submodule. The output end of the image acquisition module is communicatively connected to the input end of the image feature extraction module.

[0134] The image feature extraction module includes a low-level visual feature extraction submodule, a high-level semantic feature extraction submodule, and a feature storage and association submodule. The output end of the image feature extraction module is communicatively connected to the input end of the image theme analysis module.

[0135] The picture theme analysis module includes an image label recognition submodule, a text recognition submodule, and a theme description generation submodule. The output end of the picture theme analysis module is communicatively connected to the input end of the picture diversity calculation module.

[0136] The image diversity calculation module includes a similarity calculation submodule, a clustering grouping submodule, and a diversity index calculation submodule. The output end of the image diversity calculation module is communicatively connected to the input end of the image automatic selection module.

[0137] The automatic image selection module includes a diversity threshold and weight setting submodule, a representative image screening submodule, and a sorting and display function submodule.

[0138] The image acquisition module is responsible for efficiently acquiring and managing large amounts of image data. This is achieved through the collaborative work of the data source interface submodule, the keyword search submodule, and the download and cache submodule. The image feature extraction module utilizes the low-level visual feature extraction submodule and the high-level semantic feature extraction submodule to extract feature information at different levels from images, providing the basis for subsequent analysis and calculation. The image topic analysis module, through the image tag recognition submodule, the text recognition submodule, and the topic description generation submodule, performs topic analysis and description on images, helping users better understand their content. The image diversity calculation module, through the similarity calculation submodule, the clustering grouping submodule, and the diversity index calculation submodule, measures the diversity of image collections and provides diverse and comprehensive selection results. The automatic image selection module, through the diversity threshold and weight setting submodule, the representative image screening submodule, and the sorting and display function submodule, automatically selects qualified images and provides users with the optimal selection results. The collaborative work of these modules enables efficient, accurate, and diverse image selection, optimizes system functionality and performance, and meets the image requirements of different application scenarios.

[0139] See also Figure 8 ,The data source interface submodule is responsible for interacting with different data sources, ,including gallery and API interfaces to obtain image data;

[0140] The search keyword submodule receives keywords input by users and is used to retrieve relevant images from the data source;

[0141] The download and cache submodule downloads the images obtained from the data source and caches them for subsequent processing and use;

[0142] The low-level visual feature extraction submodule extracts the basic features of the image, including color, texture, and shape features;

[0143] The advanced semantic feature extraction submodule uses deep learning and natural language processing techniques to extract higher-level semantic features from images, including objects, scenes, and emotions;

[0144] The feature storage and association submodule stores the extracted features in a database or other storage medium and establishes an association between the features and the image for subsequent retrieval and analysis;

[0145] The image tag recognition submodule uses computer vision technology to identify tags or keywords in images to describe the content or theme of the image;

[0146] The text recognition submodule uses optical character recognition (OCR) technology to convert the text in the image into editable text for subsequent topic analysis and association;

[0147] The topic description generation submodule generates a description or summary of the image topic based on the image features and recognition results;

[0148] The similarity calculation submodule calculates the similarity scores between images by comparing their feature similarities, which is used for subsequent diversity calculation and image selection;

[0149] The clustering submodule clusters images according to their features or themes, placing similar images together to facilitate subsequent diversity calculation and image selection;

[0150] The diversity index calculation submodule calculates the diversity index of the image collection based on the similarity and clustering results, and evaluates the diversity level of the image collection;

[0151] The diversity threshold and weight setting submodule allows users to set the diversity threshold and weight according to their needs to adjust the diversity level of the system when selecting images;

[0152] The representative image screening submodule selects the most representative images in the collection based on diversity calculation and user-set thresholds;

[0153] The sorting and display submodule presents the selected images to the user in descending order based on similarity and diversity scores.

[0154] How it works: First, the system needs to prepare and organize a large amount of image data and perform feature extraction on these images. The low-level visual feature extraction submodule extracts visual features such as color, texture, and shape, while the high-level semantic feature extraction submodule extracts semantic information from the image using techniques such as object recognition and scene understanding. This allows the system to represent the features of each image with numerical values.

[0155] Next, similarity and relevance calculations are performed using the images represented by the features. The system uses the similarity calculation submodule to calculate the similarity between different images and generate similarity scores. Relevance calculations compare the relevance of the images to the user's query or selected topic to generate relevance scores. These scores reflect the degree of relevance between the images and the topic.

[0156] Based on the correlation calculation, the system performs clustering and diversity index calculation. The clustering submodule uses a clustering algorithm to group highly correlated images together, forming different image clusters. The diversity index calculation submodule measures the degree of diversity within and between each group, calculating diversity metrics based on various indicators.

[0157] Finally, the system calculates a comprehensive score for each image, taking into account both relevance and diversity. By assigning appropriate weights, images with higher relevance and diversity receive higher scores. Ultimately, the system automatically selects a representative and diverse collection of images as the final result to display to the user.

[0158] Through the above steps, the method for calculating and automatically selecting image thematic content diversity can efficiently select a diverse image collection that is relevant to user needs. This method can save users time and effort, improve work efficiency, and ensure that the selected image collection is representative and diverse in terms of thematic content.

[0159] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A method for calculating and automatically selecting the diversity of image subject content, characterized in that: The following steps are involved: Data preparation and feature extraction; Correlation calculation to obtain correlation score; Diversity calculation to obtain diversity measurement values; Comprehensive scoring and automatic selection; The steps of calculating the correlation are specifically as follows: Chord similarity is selected as the similarity measure; For the feature vector of each image and the feature vector of a given theme, similarity calculation is performed. During the similarity calculation process, the dimension of the feature vector is reduced using a dimensionality reduction method specifically a principal component analysis to improve computational efficiency and reduce the impact of noise; Use the inner product between the eigenvectors to divide by the norm of the eigenvector to obtain the relevance score, which indicates the relevance of the image to the given topic. Rank or classify each image's relevance score to a given topic, where higher relevance scores indicate greater relevance to the topic and lower relevance scores indicate less relevance to the topic; Specifically, the diversity calculation uses variance and standard deviation as diversity measurement methods. The variance and standard deviation are statistical indicators that measure the degree of dispersion of data distribution. The larger the variance and standard deviation, the higher the difference and diversity between feature vectors. The steps of diversity calculation are specifically as follows: Calculate the mean of the eigenvectors; Calculate the difference vector; Calculate variance and standard deviation; The steps of the comprehensive scoring are specifically as follows: Normalize the relevance scores and diversity metrics to map their value ranges to a unified range; Combining relevance and diversity, a weighted synthesis method is used to perform a weighted summation of the normalized relevance score and diversity metric, balancing relevance and diversity according to the weight setting to obtain a comprehensive score; Apply the composite score to each image to get the final score for each image; The steps of automatic selection are specifically as follows: Sort all images according to the final score, arranging them in order from high to low; According to the needs, set the threshold to filter out the images that meet the specific requirements. The threshold includes absolute value and relative value; Select single or multiple images as the final result according to specific needs.

2. The method for calculating and automatically selecting the diversity of picture subject content according to claim 1, characterized in that: The steps of data preparation are specifically as follows: Use the public image library ImageNet to collect or obtain image datasets related to a given topic; For topics in a specific field, choose a dataset specifically for that field. Specifically, in the field of computer vision, when the dataset focuses on target detection, image segmentation, and object recognition tasks, use the COCO and PASCAL VOC image libraries. Make sure the images you choose cover different aspects of a given topic, including different perspectives, scenes, and objects. Reduce data redundancy and increase data availability based on preprocessing operations including image scaling, cropping, rotation, normalization, and standardization; Data augmentation technology is used to perform random transformations on the original images, including flipping, rotation, translation, scaling, and adjusting brightness and contrast, to generate more training samples and expand the size of the dataset.

3. The method for calculating and automatically selecting the diversity of picture subject content according to claim 1, characterized in that: The steps of feature extraction are specifically as follows: Select the convolutional neural network (CNN) model, specifically ResNet, to perform feature extraction on the image dataset; The images in the dataset are fed into a convolutional neural network (CNN) model and forward propagation is used. During the forward propagation process, the image is processed through a series of convolution, pooling, and activation functions to gradually extract increasingly abstract and advanced feature representations. For mid-level feature extraction, shallower convolutional layers acquire low-level features, including edges and textures, while deeper convolutional layers acquire higher-level features, including object shapes and components. Features from multiple levels are combined to obtain richer information. During feature extraction, an attention mechanism is used to enhance the importance and expressiveness of features, and spatial pyramid pooling techniques are applied to capture multi-scale features. The features after the convolution layer are reduced in dimension and compressed through the fully connected layer to obtain a more abstract and advanced feature representation. Classification and detection tasks are performed by extracting the feature vector of the fully connected layer.

4. The method for calculating and automatically selecting the diversity of picture subject content according to claim 3, characterized in that: Calculating the mean of the eigenvectors specifically includes performing a sum operation on all eigenvectors in the image set to obtain a sum vector, and dividing the sum vector by the number of eigenvectors to obtain a mean vector of the eigenvectors; The calculation of the difference vector is specifically as follows: for each eigenvector, subtract the mean vector of the eigenvector to obtain a difference vector, where the difference vector represents the difference between each eigenvector and the overall mean; The calculation of the variance and standard deviation is specifically as follows: performing a square operation on the difference vector to obtain a square difference vector, performing a sum operation on the square difference vector to obtain a total value, dividing the total value by the number of eigenvectors to obtain the variance, and performing a square root operation on the variance to obtain the standard deviation.

Citation Information

Patent Citations

  • Video description method based on target space semantic alignment

    CN114154016A

  • Text adjusted visual search

    US20220138247A1