Image retrieval method based on big data mining
By combining metadata screening, geographical information clustering and visual feature fusion in image retrieval, the problem of failure to effectively combine geographical and temporal information in the prior art is solved, and the accuracy and efficiency of image retrieval is improved.
Patent Information
- Application Number
- CN202510241659.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art fails to effectively combine geographical and temporal information with visual features in image retrieval, resulting in low accuracy and consistency of the retrieval, especially in image data processing in multiple devices and multiple environments.
The image retrieval method based on big data mining is adopted. By collecting images and metadata, metadata screening and geographical information clustering are performed, visual features are extracted and combined with device types, feature fusion and weighting optimization are carried out, inverted index structure is constructed, and similarity matching query and sorting are performed.
It improves the accuracy and response efficiency of image retrieval, enhances the deep understanding of image content, improves the correlation and sorting rationality of matching results, and ensures the good performance of the retrieval system when the data volume increases.
Smart Images

Figure CN120179844A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image retrieval, and particularly to an image retrieval method based on big data mining. Background Art
[0002] Image Retrieval is an important branch in the fields of computer vision and information retrieval, aiming to find images that match the user's query from a large image database. This technology uses computer vision algorithms, deep learning models, and traditional feature engineering methods to extract features, construct indexes, and perform matching on images to improve the accuracy and efficiency of retrieval.
[0003] In the prior art, the application of geographical and time information is relatively limited. Usually, it is only used as an additional attribute for simple filtering, and it fails to effectively combine visual features for in-depth association, affecting the accuracy of retrieval. In the processing of image data captured by multiple devices in multiple environments, the existing methods fail to fully utilize the association between device features and content features, resulting in relatively low consistency of image matching results captured at different devices and different time periods, affecting the stability and scope of application of retrieval. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the drawbacks existing in the prior art, and to propose an image retrieval method based on big data mining.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions. An image retrieval method based on big data mining includes the following steps:
[0006] Collect images and metadata. The metadata includes shooting time, location, and device type. Screen the metadata fields to determine a candidate image set and obtain a metadata screening result; classify the metadata screening result according to the location and timestamp, and cluster the location through a geographic information system to obtain a geographic clustering result;
[0007] Based on the geographic clustering result, identify objects and scenes in the image, extract visual features, and generate visual feature data; combine the visual feature data with the device type in the metadata for feature fusion, and optimize the feature vector through a weighted algorithm to obtain a fused feature vector;
[0008] Based on the fused feature vector, construct an index structure, use an inverted index to index the feature vector, establish a retrieval path, and obtain an index construction result; perform a similarity matching query on the index construction result, and sort according to the similarity scoring standard to obtain a similarity sorting result;
[0009] Using the similarity sorting result, perform user query response. According to the query conditions input by the user, extract the associated images from the sorting result for result presentation, and generate user query results.
[0010] Preferably, the steps for obtaining the metadata screening result are as follows:
[0011] Perform consistency checks on the shooting time, shooting location, and device type fields in the collected image metadata, determine the valid metadata fields through the consistency check results of the metadata fields, and obtain a set of valid metadata fields;
[0012] According to the set of valid metadata fields, calculate the metadata validity score for each image. The calculation formula is:
[0013]
[0014] Among them, M i represents the metadata validity score of image i, T i represents the consistency check result of the shooting time field of image i, L i represents the consistency check result of the shooting location field of image i, E i represents the consistency check result of the device type field of image i, F i represents the difference between the data of the shooting time field of image i and the standard shooting time, G i represents the spatial position gap between the longitude and latitude coordinates of the shooting location field of image i and the standard coordinates, H i represents the difference in matching between the device type field of image i and the mainstream device type;
[0015] Based on the metadata validity score, screen and select the images to form a metadata screening result.
[0016] Preferably, the steps for obtaining the geographical clustering result are as follows:
[0017] According to the shooting location and shooting timestamp of the metadata screening result, perform spatio-temporal classification, delimit the initial location classification area through the spatial distance of the location coordinates, and form an initial spatial classification result;
[0018] Based on the initial spatial classification result, calculate the spatial clustering fitness score of each image's belonging location and the belonging location classification area. The calculation formula is:
[0019]
[0020] Among them, P j represents the spatial clustering fitness score of image j, X j 、Y jrespectively represent the longitude and latitude coordinate values of image j, X c and Y c respectively represent the longitude and latitude coordinate values of the center point of the location classification area to which image j belongs, t j represents the timestamp value of image j, t c represents the standard timestamp value of the location classification area to which image j belongs;
[0021] Based on the spatial clustering fitness score, re-divide the clustering area of the images to form a geographical clustering result.
[0022] Preferably, the obtaining steps of the visual feature data are as follows:
[0023] According to the geographical clustering result, respectively identify the object category and scene category in each clustered image, determine the texture details, color distribution and edge sharpness features of the object category, and obtain the image object feature information;
[0024] According to the image object feature information, calculate the scene saliency score of the image visual feature, and the calculation formula is:
[0025]
[0026] Among them, Q j represents the scene saliency score of image j, c j represents the color saturation of the object category of image j, s j represents the texture detail sharpness of the object category of image j, e j represents the edge sharpness feature of the object category of image j, h j represents the main color uniformity degree of the object category of image j, b j represents the main edge direction of the object category of image j, b s represents the main edge direction standard of the scene category to which image j belongs;
[0027] According to the scene saliency score, extract the visual features to form the visual feature data.
[0028] Preferably, the obtaining steps of the fusion feature vector are as follows:
[0029] Perform a one-to-one association mapping between the visual feature data and the device type in the image metadata, construct the correspondence between the visual features and the device type, and form a feature mapping set;
[0030] According to the feature mapping set, calculate the fusion tightness of the image feature vector, and the calculation formula is:
[0031]
[0032] Among them, Z j represents the feature vector fusion compactness of image j; u j represents the texture complexity of the visual features of image j; d j represents the color diversity of the visual features of image j; f j represents the edge density of the visual features of image j; m j represents the pixel resolution of the device type corresponding to image j; m d represents the average pixel resolution of all device types;
[0033] Select images according to the feature vector fusion compactness, fuse the visual features and the feature vectors of the device type to generate fused feature vectors.
[0034] Preferably, the steps for obtaining the index construction result are as follows:
[0035] Based on the fused feature vectors, according to the numerical distribution characteristics of the fused feature vectors, divide the numerical intervals of the feature vectors, establish index nodes in each numerical interval of the feature vectors, set unique identifiers for the index nodes, mark and sort the index nodes to form an initial index node set;
[0036] Based on the initial index node set, extract the numerical values of the fused feature vectors one by one, perform the matching between the numerical values and the index node identifiers, and according to the matching results, establish the mapping association between the fused feature vectors and the index nodes, and gradually connect all the mapping associations to generate an inverted index mapping structure;
[0037] Based on the inverted index mapping structure, call the node identifiers of the initial index node set, verify the access order of the fused feature vectors one by one according to the mapping association results between the fused feature vectors and the index nodes, and connect the index nodes in sequence according to the access order to construct a complete retrieval path for the fused feature vectors to obtain the index construction result.
[0038] Preferably, the steps for obtaining the similarity sorting result are as follows:
[0039] Calculate the similarity matching score of the fused feature vectors according to the index construction result;
[0040] Sort the fused feature vectors from large to small according to the similarity matching score, determine the sorting positions of the fused feature vectors one by one, and connect the sorted fused feature vectors in sequence to obtain the similarity sorting result.
[0041] Preferably, the steps for obtaining the user query result are as follows:
[0042] According to the similarity sorting result, call the corresponding index nodes of the sorted fusion feature vectors in the index construction result, locate the associated image identifiers of the index nodes in sequence, extract the set of image identifiers corresponding to the query matching result, and generate an associated image identifier list;
[0043] According to the associated image identifier list, extract the corresponding image data, call the visual feature data and metadata field information of the image data, perform content collation and structural adjustment of the query result, and present the image data in the arrangement order of the query matching result to generate a user query result.
[0044] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0045] In the present invention, during the image retrieval process, by combining image metadata screening, geographic information clustering, visual feature fusion, index construction, and optimized matching strategies, the retrieval accuracy and response efficiency are improved. In the metadata screening link, information such as time, location, and device type is used to screen the data, reducing the interference of irrelevant data and making the candidate image set more targeted. The combination of geographic information and timestamp classification, and the use of geographic information systems for clustering analysis enable the spatial distribution information to play a greater role in the retrieval process, and more accurate results can be obtained for queries in specific regions. The visual feature extraction adopts object recognition and scene analysis methods to enhance the in-depth understanding of the image content. After being fused with the device type information, a more discriminative feature expression is formed. After weighted optimization, the representativeness of the feature vector is strengthened, improving the matching effect. During the index construction process, the introduction of the inverted index makes the data structure more efficient. Combined with feature vector optimization, the query speed in a large-scale data environment is improved. The similarity matching optimization strategy combines weighted calculation and multi-layer screening, making the retrieval result more in line with the user's query intention and improving the relevance and sorting rationality of the retrieval result. Based on the efficient processing logic for large-scale data, the retrieval system can still maintain good performance when the data volume increases and provide more targeted matching results in specific application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0047] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] Please refer to Figure 1 , the present invention provides a technical solution, an image retrieval method based on big data mining, including the following steps:
[0049] Collect images and metadata. The metadata includes the shooting time, location, and device type. Screen the metadata fields to determine the candidate image set and obtain the metadata screening result. Classify the metadata screening result according to the location and timestamp, and cluster the locations through a geographic information system to obtain the geographic clustering result.
[0050] Based on the geographic clustering result, identify the objects and scenes in the images, extract visual features, and generate visual feature data. Combine the visual feature data with the device type in the metadata for feature fusion, and optimize the feature vector through a weighted algorithm to obtain the fused feature vector.
[0051] Based on the fused feature vector, construct an index structure, index the feature vector using an inverted index, establish a retrieval path, and obtain the index construction result. Perform a similarity matching query on the index construction result, and sort according to the similarity scoring criteria to obtain the similarity sorting result.
[0052] Use the similarity sorting result to respond to user queries. According to the query conditions input by the user, extract the associated images from the sorting result for result presentation and generate the user query result.
[0053] The steps to obtain the metadata screening result are as follows: Conduct a consistency check on the shooting time, shooting location, and device type fields in the collected image metadata, and determine the valid metadata fields through the consistency check result of the metadata fields to obtain the set of valid metadata fields.
[0054] According to the set of valid metadata fields, calculate the metadata validity score for each image. The calculation formula is:
[0055]
[0056] where M i represents the metadata validity score of image i, T i represents the consistency check result of the shooting time field of image i. The valid value for the check is 1, and the invalid value is 0. L i represents the consistency check result of the shooting location field of image i. The valid value for the check is 1, and the invalid value is 0. E i represents the consistency check result of the device type field of image i. The valid value for the check is 1, and the invalid value is 0. F i represents the difference between the data in the shooting time field of image i and the standard shooting time. G i represents the spatial position gap between the longitude and latitude coordinates of the shooting location field of image i and the standard coordinates. H i represents the difference in the matching of the device type field of image i with the mainstream device type.
[0057] Based on the metadata validity score, images are screened and selected to form the metadata screening result.
[0058] Specifically, based on the content cited in the consistency check of the shooting time, shooting location, and device type fields in the collected image metadata, it is disassembled and extracted. By reading the shooting time value in each image record and comparing it with the uniformly set time range, for example, comparing the date with the interval between January 1, 2022, and January 1, 2023, then comparing the longitude and latitude coordinates of the shooting location with the defined geographical range, for example, comparing the longitude with the interval between 100° east longitude and 120° east longitude and the latitude with the interval between 20° north latitude and 30° north latitude. For the device type, it is matched with the list of registered device types collected in advance and the matching result is recorded. When the matching result is consistent with the registered device type list, the content of this field is considered to pass the verification. Then, the process of extracting data item by item is executed. By reading the shooting time field value item by item and verifying whether it is within the time range, comparing the shooting location coordinate value with the geographical range and confirming that it does not exceed the boundary, and at the same time checking whether the device type field exists in the registered list. Subsequently, all the records that pass the verification are marked as qualified records and stored in the temporary dataset. In this process, the time range and geographical range are determined by the actual acquisition situation. For example, image shooting work is carried out throughout 2022 in a certain place and the shooting area is within the range of fixed longitude and latitude distribution. The list of registered device types can be summarized through the information of the device models actually put into use and specific device identifiers are given. After performing the above item-by-item inspection, the qualified records are merged. After completing the verification and comparison of all records, an effective metadata field set is obtained.
[0059] The benefit of the formula is that by first calculating the validity flag values of the shooting time field, shooting location field, and device type field, and then combining the difference data between them and the standard time, standard geographical coordinates, and mainstream device types, it comprehensively reflects the overall reliability of the metadata, thus taking into account both the consistency and the quantification result of the difference in the same score.
[0060] T i Represents the consistency check result of the shooting time field of image i. A value of 1 indicates that the content of this field passes the verification, and a value of 0 indicates that it fails the verification. The verification process compares the shooting time of the record with the uniformly maintained standard time range, and the time range is set within the specified interval of a certain year or several years. The specific interval is defined by the acquisition plan. For example, taking 0:00 on January 1 of a certain year as the starting value and the end time of December 31 of a certain year as the ending value. Once the shooting time falls within this interval, it is considered qualified and it is determined that T i = 1, otherwise the record T i = 0.
[0061] Li Indicates the consistency check result of the shooting location field of image i. A value of 1 means the content of this field conforms to the specified geographical area range, and a value of 0 means it does not. It is evaluated by parsing the longitude and latitude coordinates in the actual image record and comparing them with the pre-determined geographical limit intervals. For example, the longitude can be restricted between 100°E and 120°E, and the latitude can be restricted between 20°N and 30°N. These interval values are obtained through research on the range of the shooting locations of the project, and the specific precise coordinate range can be provided by relevant surveying and mapping information. In a certain detection session, the longitude of an image is 106.325° and the latitude is 26.874°. Compare it item by item with the limit intervals. If it is within the upper and lower limits, record L i = 1, otherwise set L i = 0.
[0062] E i Indicates the consistency check result of the device type field of image i, taking 1 or 0. It is judged according to whether the shooting device model corresponding to the image belongs to the pre-listed mainstream device type list. The mainstream device type list is composed of the information of the actually used shooting devices, including mobile phone brands and models, professional camera models, and other registered shooting devices. During the inspection, compare the model field in the image EXIF with the device type list one by one. Once the same model is matched, it is considered that E i = 1, if not matched, it is considered that E i = 0.
[0063] F i Represents the difference between the shooting time field data of image i and the standard shooting time. This difference is obtained by calculating the current image shooting timestamp minus the reference timestamp. The reference timestamp is established by the project plan. For example, record the project start time as 1643644800. Calculate the difference between the two for the record with an image shooting timestamp of 1649244800 to get 5600000 seconds. The numerical range depends on the specific shooting cycle length. When obtaining F i The timestamp needs to be converted to second-level data and subtracted from the reference value.
[0064] G i Represents the spatial position gap between the longitude and latitude coordinates of the shooting location field of image i and the standard coordinates. To achieve numerical processing, first obtain the longitude and latitude corresponding to this image, and then obtain the pre-set reference coordinates. For example, record the reference longitude and latitude as 105.500° and 25.200°. Calculate the approximate value of the planar distance between the two as the value of G i The planar distance can be completed through the Pythagorean theorem operation.
[0065] H iIndicates the difference in the device type field of image i matching the mainstream device type. The calculation steps are as follows: assign numerical numbers to each model in the registered device type list, and then calculate the difference between the image model number and the reference model number. The reference model number is selected from the frequently used models in the device type list. The value can be taken as the first one after sorting according to the model voting results as the reference. For example, set the model number of the first-ranked model as 20. If the image model number is 22, then H i = 2.
[0066] For example, substituting into the formula to calculate M i ≈ 2960000.003. This result indicates that the difference in this image is relatively large in terms of time and location. However, both the shooting time field and the shooting location field are qualified data, but only the device type verification fails. The comprehensive metadata validity score is within this numerical range. The larger the value, the farther away from the reference. If in the subsequent steps, for M i When setting a screening threshold of 100000.000, it means that when M i is greater than this threshold, it is considered that the deviation degree is obvious, otherwise it is considered that the deviation degree is limited.
[0067] Based on the previously obtained metadata validity score, screen the images and select to form the metadata screening result. By reading the score values of each image and referring to a pre-defined threshold, for example, setting the threshold to 50000.000, compare the score with this threshold item by item. When the score is not higher than the threshold, mark it as a recommended selected image in the temporary list and record its image identifier. If the score is higher than the threshold, mark it as an excluded item in the temporary list and record the exclusion reason. During the whole process, it is necessary to ensure that the same precision level is used when comparing the score values with the threshold and execute uniformly under the floating-point data type. The score data is calculated from elements such as the shooting time difference and location difference in the previous text, and can be directly read from the corresponding calculation results and stored sorted by image number. To avoid misjudgment in score comparison, an operation of rounding or retaining decimal places will be performed when loading the score, and then compared with the set threshold in a consistent format. The setting of the threshold needs to be clearly confirmed in combination with the acceptable range of time and location differences in the actual shooting project. For example, when the project time span is long and the shooting area is wide, the threshold can be set larger. If the project time span is limited and the shooting range is concentrated, the threshold can be set smaller. After completing all score comparisons, generate a metadata screening result, which contains the numbers and screening marks of each image. In addition, form a list of the recommended selected image numbers and perform matching and reference in subsequent stages. The whole process is carried out in a way of loading the scores and generating the list for all images at one time, and finally obtain the set of all images that pass the screening.
[0068] The steps for obtaining the geographical clustering results are as follows: Based on the shooting locations and shooting timestamps of the metadata screening results, perform spatio-temporal classification, delimit the initial location classification regions through the spatial distances of the location coordinates, and form the initial spatial classification results;
[0069] Based on the initial spatial classification results, calculate the spatial clustering fitness scores of each image belonging to its location and the location classification region. The calculation formula is as follows:
[0070]
[0071] where P j represents the spatial clustering fitness score of image j, X j , Y j respectively represent the longitude and latitude coordinate values of image j, X c , Y c respectively represent the longitude and latitude coordinate values of the center point of the location classification region to which image j belongs, t j represents the shooting timestamp value of image j, and t c represents the standard timestamp value of the location classification region to which image j belongs;
[0072] Based on the spatial clustering fitness scores, re-divide the clustering regions of the images to form the geographical clustering results.
[0073] Specifically, according to the shooting locations and shooting timestamps contained in the metadata screening results obtained previously, combined with the geographical coordinates and shooting moments corresponding to each image, retrieve the longitude, latitude, and actual time values of each image record in sequence, and compare these values with the spatio-temporal range set in the actual project. Based on the comparison results, divide several initial location classification regions. First, extract all records and read the longitude and latitude data one by one. For example, set the longitude range from 100° east longitude to 120° east longitude, and the latitude range from 20° north latitude to 40° north latitude. This range can be determined according to the geographical information obtained from previous investigations. Subsequently, compare the longitude and latitude values of each image with this range, and mark the records that meet the range in the intermediate list. Then, read the shooting timestamp and compare it with the actual project schedule. For example, compare the timestamp value with an interval from 0:00 on January 1, 2023, to 23:59 on December 31, 2023. By calculating the differences between the timestamp and the start and end values of the interval one by one, all records within the interval are grouped into the same time period set. Then, use the distance value of the geographical coordinates to establish the initial grouping method. To define the distance, first use the longitude and latitude differences and the spherical distance formula in the map coordinate system to calculate the approximate kilometer value between two points. Group the images within a distance of 50 kilometers in the same time period into an initial location classification region, and list the records exceeding 50 kilometers separately in other groups. If there are many images, smaller radius ranges can be further subdivided, such as 20 kilometers or 30 kilometers. Then, set the grouping upper limit of the classification region according to the possible densely photographed areas at each location in the project. If the distribution density of some locations is relatively high, the radius value needs to be reduced and multiple groups need to be added. During the retrieval process, the geographical coordinates and shooting timestamp can be packed together for sorting and paired distance calculation, which is convenient for subsequent segmentation by time in sequence and combining the location distance to delimit the classification region. Each record will be assigned to an initial location classification region. Finally, mark and organize all classification regions and write them into a temporary table as the spatial classification reference data for subsequent steps, forming the initial spatial classification result.
[0074] The benefit of the formula lies in comprehensively weighing the geographical distance factor and the shooting moment difference between the image and the corresponding location classification region. By multiplicatively combining the distance component and the time component, the coupling degree of location and time can be reflected simultaneously in a score.
[0075] X jRepresents the longitude coordinate value of image j, which is obtained by parsing the shooting location saved in the image record or EXIF. When parsing, the shooting longitude data can be extracted from the database record corresponding to the image ID, or directly read from the image file header information. To ensure accuracy, the method of continuously collecting geographical locations is adopted in the early stage. The shooting device automatically records the longitude and latitude during shooting and stores them in the project database. The database will save the longitude information of all images and correspond it one by one with the image ID to obtain X j After that, it needs to be calibrated in degrees. For example, 106.3247° is directly recorded as 106.3247.
[0076] Y j Represents the latitude coordinate value of image j. The acquisition method is similar to that of longitude. The actual shooting latitude is obtained by parsing the EXIF data in the same database or image file, and then this value is extracted into the same clustering analysis table and marked in the row record corresponding to the image, matching the longitude. During the project implementation, a calibrated positioning system will be used to accurately collect the latitude to four decimal places or more to meet the accuracy requirements of the clustering operation for location information. For example, when registered as 29.9875, it indicates that the latitude recorded by the positioning system during the shooting of this image is 29.9875°.
[0077] X c Represents the longitude value of the center point of the location classification area to which image j belongs. The acquisition method of this value is usually based on the already divided location classification area. First, the longitude values of all images within each classification area need to be selected, and then the average value or weighted average value is calculated and determined as the central longitude of this area. If the simple average method is used, the longitudes of all images within this area can be added up and then divided by the number of images to obtain X c , In actual operation, factors such as area or distribution density can also be introduced, but it is necessary to clarify how to quantify these data in the calculation. The acquisition method can be referred to as follows: First, retrieve n images from classification area A, and their longitudes are x1, x2,..., x n , When there is no other weighting requirement, use as X c , The size of the longitude should be consistent with the project geographical information. For example, if the longitudes of the images in area A are mostly around 106°, then X c should also fall around 106°. For the convenience of updating, X may be recalculated every time a new image is added c . For example, when there are 4 images in area A with longitudes of 106.2171, 106.2203, 106.2178, and 106.2195 respectively, the sum of the four values is 424.8747, and then divided by 4 to get 106.218675, which is recorded as the central longitude of this area.
[0078] Y c represents the latitude value of the center point of the location classification area to which image j belongs. The calculation method is the same as that of X c and can be determined using the same averaging method or weighted averaging method. If there are m latitude values of images in the same classification area, the center latitude can be obtained by adding them first and then dividing by m. It can also be corrected during calculation by means of the distribution density within the area or other indicators, and finally the measured value of Y c is obtained. For example, in area A, the latitudes of four images are 29.9943, 29.9912, 29.9978, and 29.9905 respectively. Add these four values to get 119.9738, and then divide by 4. The result is 29.99345, then Y c can be recorded as 29.99345.
[0079] t j represents the timestamp value of the shooting time of image j. Specifically, the recorded value at the second level or millisecond level is taken, which can be determined by the internal clock of the shooting device or the "shooting time" field in the image file EXIF, and then converted into the form of an integer timestamp for comparison. In order to ensure effectiveness in calculation, it is necessary to uniformly record the time reference at the beginning of the project and correct the system time of the shooting device, and convert the obtained shooting moment into the cumulative number of seconds or milliseconds since a certain initial time. The acquisition method is as follows: by reading the shooting date and moment in the image, such as 09:30:10 on April 21, 2023, and then subtracting it from the set starting time 00:00:00 on January 1, 2023, and then converting the difference into seconds or milliseconds and saving it as t j . If the device provides more precise time accuracy, more digits can be retained. In a certain project record, the shooting moment of an image is 09:30:10 on April 21, 2023. After conversion, the timestamp obtained is 2629810 seconds, that is, t j = 2629810.
[0080] t c represents the standard timestamp value of the location classification area to which image j belongs. This value is mainly obtained by statistically analyzing the shooting timestamps of all images within the area, just like X c and Y c . The calculation method can simply take the average, or a weighted method can be adopted according to specific requirements. Assume that a region contains k images, and their timestamps are t1, t2,..., t k . If there is no other weight assignment requirement, can be used to obtain t c . For example, in a region, the shooting timestamps of three images are 1000000, 1010000, and 1020300 respectively. Add them up to get 3030300, and then divide by 3 to get approximately 1010100.
[0081] For example, substituting into the formula for calculation gives 0.003211. This result indicates that the time difference between the image and the center of the classified area is relatively large (a 310 - second difference leads to an extremely small multiplier), the geographical distance is small but the time gap is significant. Therefore, the comprehensive score is only 0.003211. If a threshold of 0.01 is set during the subsequent clustering process, it can be recognized that its clustering fitness is low and optimization is required when re - dividing the area or adjusting the time threshold.
[0082] This result shows that when both the geographical deviation and the time deviation are small, the P j value will be higher. Conversely, if any one of the deviations is too large, the score will be reduced. Through such scoring, the matching areas of the image with location and time elements can be further divided, providing a reference basis for subsequent area integration.
[0083] Based on the spatial clustering fitness scores obtained previously, scan each image one by one to obtain the values when comparing the geographical difference and time difference with the center point of the comparison area, and extract these score results in a clustering analysis table and then sort them. Through sorting, the closer corresponding relationship between the images with higher scores and the areas can be located. List the images with lower scores separately in an additional list and check whether there are obvious deviations in their time distribution or longitude - latitude distribution from the originally divided areas. Subsequently, re - identify the classified areas to which all images belong according to the score size. If the score of an image is lower than the specified threshold, for example, lower than 0.02, then perform the re - division operation. By repeatedly matching the longitude - latitude data of such images with the center point information of neighboring areas and calculating the new fitness scores again, if the new score is significantly improved, then merge the image into the new classified area and update the center coordinates and standard timestamps of this area. If the scores of multiple areas are all low, then generate a new classified area separately to accommodate these records. During the whole process, it is necessary to compare the specific geographical radius and time range to judge the area coverage, and set the corresponding upper and lower limits of the values as clustering indicators. If there are still images with a certain degree of matching with multiple areas but no score is significantly higher, then this image can be recorded as an independent point in the clustering analysis table and merged uniformly at the end. After this round of operations, the updated area list and image subordination relationship can be generated according to the new fitness assignment relationship. By writing the final result into the clustering result table and recording the new area numbers corresponding to each image, the geographical clustering result is formed.
[0084] The steps for obtaining visual feature data are as follows: According to the geographical clustering result, for each image after clustering, identify the object category and scene category within the image respectively, determine the texture details, color distribution, and edge sharpness characteristics of the object category, and obtain the image object feature information;
[0085] Calculate the scene saliency score of the image visual features based on the image object feature information. The calculation formula is as follows:
[0086]
[0087] Among them, Q j represents the scene saliency score of image j, c j represents the color saturation of the object category of image j, s j represents the texture detail clarity of the object category of image j, e j represents the edge clarity feature of the object category of image j, h j represents the main color uniformity of the object category of image j, b j represents the main edge direction of the object category of image j, b s represents the main edge direction of the standard of the scene category to which image j belongs;
[0088] Extract the visual features according to the scene saliency score to form visual feature data.
[0089] Specifically, according to the geographical clustering results obtained previously, retrieve the clustering labels corresponding to each image and the object category and scene category information registered in the database. By centrally analyzing the association between these labels and the registered information, read the object category code and scene category identifier of each image one by one, and disassemble and extract the definition content related to texture details, color distribution, and edge sharpness from the pre-established category description file. Combine the texture gradient value, chromaticity histogram curve peak value, and edge sharpness measurement value recorded during the feature detection process of the corresponding image. Identify the main objects and their texture feature trends in the image through quantitative comparison. Subsequently, associate with the corresponding scene definition based on the scene category identifier to confirm the scene environment type of the image, and determine that the image may correspond to scenes such as outdoor perspectives, urban landmarks, or natural environments. When analyzing the texture gradient value, it is necessary to first normalize the gradient intensity and determine a calibration value between 0 and 1. Then, compare the distribution width of the color channels for the chromaticity histogram curve peak value. All detection processes are judged with a unified threshold and measurement standard. For example, when matching the texture gradient value with the scene category, set a threshold range of 0.2 to 0.5. Any texture gradient value corresponding to an image exceeding 0.5 is regarded as having a high texture feature, and less than 0.2 is regarded as having a low texture feature. For the edge sharpness measurement value, it is divided from 0 to 100. Images with a value greater than 60 are recorded as having high edge sharpness, those between 30 and 60 are recorded as medium, and those less than 30 are recorded as low. After comparing the texture gradient values, color channel peak values, and edge sharpness of all images, these values will be recorded in the image feature analysis table and indexed according to the combination key of the image ID and object category and scene category. Finally, summarize the texture details, color distribution, and edge sharpness feature values of the image in the index result to form the image object feature information.
[0090] The advantage of the formula lies in integrating color saturation, texture detail sharpness, and edge features, as well as the uniformity of the main color tone and the difference in edge direction. Through the joint participation of multiple parameters, it improves the quantization and discrimination accuracy of the scene saliency.
[0091] c j Represents the color saturation of the object category of image j. The value range can be set between 0 and 100 in combination with the actual acquisition results. The acquisition method is to perform color model conversion on the image, extract the chroma value of each pixel, and then take the average. The calculation formula can be set as: where p αrepresents the chroma value of the α-th pixel point, n represents the total number of pixels. During the acquisition process, it is necessary to ensure consistent lighting conditions to maximize the avoidance of light flicker interference. The acquired original pixel information can be automatically recorded at the shooting device end and written into the meta-information structure of the image file, and then read and processed uniformly. For example, in one operation, after a certain image undergoes color model conversion, the chroma value of each pixel falls within a certain range, and the sum of all is 480,000. If the total number of pixels is 6,000, then according to this formula, c j = 80.
[0092] s j represents the texture detail sharpness of the object category of image j, and its value range can be set between 0 and 100. It is obtained by performing local gradient detection and neighborhood contrast operation on the image. Each time, the local contrast is calculated in blocks within the object main area, and then the average value is taken and normalized. The resulting numerical formula is: where d β represents the gradient contrast value of the β-th block, m represents the number of blocks. During acquisition, it is necessary to accurately extract the pixel matrix in the object main area, and then obtain the local gradient information through a convolution operation, and ensure the comparability of contrast measurement under the same lighting conditions. In actual projects, s j is often compared with the high-resolution display standard to determine whether the details are sharp enough. For example, in one case, the object main area is divided into 20 blocks, and the sum of the gradient contrast values of each block is 1360.
[0093] e j represents the edge sharpness feature of the object category of image j, and its numerical range can also be set between 0 and 100. It is obtained by analyzing the sharpness index of the gray-scale change at the image edge. Usually, the object contour area is first identified, and then the gray-scale change rate is statistically calculated along the contour line, which is used as the main evaluation index for edge sharpness. Specifically, it can be defined as: where g γ represents the gray-scale jump value in the γ-th edge sample, r represents the number of edge sample points. When acquiring this value, it is necessary to first lock the object edge pixel band through a segmentation algorithm, and calculate the gray-scale gradient of adjacent pixels point by point. After summing and then averaging and normalizing to 0 to 100, in an example acquisition, after contour detection of a certain image, 100 edge samples are obtained, and the total gray-scale jump value is 5,000. After normalization, e j is 50.
[0094] h j represents the main color uniformity degree of the object category of image j, usually set between 0 and 100. It is obtained by measuring the dispersion of the color distribution in the main area of the image. The reciprocal of variance can be used to measure the uniformity degree, and the defined formula is: where u δ represents the main chromaticity value of the δ-th pixel in the main body block, is the average of the main chromaticity values of all pixels, k is the number of pixels. By actual measurement, the camera white balance setting can be unified and the exposure parameters can be locked during shooting to ensure the comparability of the main tones. For example, in a certain acquisition, the main body block of the object contains 3,000 effective pixel points, and the calculated main chromaticity variance is 45. Substituting into the formula gives This value is then scaled to the range of 0 to 100 and can be recorded as 14.75.
[0095] b j represents the main edge direction of the object category of image j. This parameter is obtained by voting on the gradient directions of the edge pixels within the main body area of the object, expressed in degrees between 0° and 179°. During acquisition, first perform edge detection on the image and extract the gradient direction of each edge, then count the range of the gradient direction that appears most frequently, which is defined as the main edge direction. When numerical processing of different directions is required in the aggregation calculation, the degrees need to be discretized into equal-step intervals and recorded. An example of obtaining this value is as follows: extract the gradient directions of all effective edges in the main body area of a certain image, record them in an array, and then count every 5°. It is found that the number of occurrences near 125° is the most. Finally, b j is recorded as 125.
[0096] b s represents the main edge direction standard of the scene category to which image j belongs, also in degrees between 0° and 179°. For each scene category, the main edge directions of the sample images are summarized in advance, and the median value of the most common degree segment is taken as the standard value, which is uniformly stored in the scene category definition table. For example, after a large number of image statistics for an outdoor landscape scene, it is determined that the most common main edge direction is 130°, then b s is recorded as 130.
[0097] For example, substituting into the formula for calculation gives 59.11. The scene saliency score of this image is approximately 59.11, which belongs to a medium to high level in the numerical system. When this value is higher than 30, it can be regarded as having a certain visual prominence, and when it is higher than 60, it can be regarded as having a very strong scene impact. In the current case, it is close to 60, indicating that the texture, color, and edge elements of this image have reached a relatively high saliency after synthesis, providing an important quantitative basis for the extraction of subsequent visual features.
[0098] Based on the scene saliency scores obtained previously, retrieve the records corresponding to the image object feature information and the score value item by item, and compare the scores obtained for each image with a pre-set discrimination threshold. For example, in this project, 30 is set as the discrimination threshold. When the scene saliency score is greater than 30, the texture, color, and edge elements of the image are marked as obvious features. When the score value is lower than 30, they are marked as general features. Subsequently, screen the list of images with scores exceeding 30 by traversing the values of all images in the visual feature analysis table, and record the corresponding scene categories and object categories. Then, perform a secondary summary on the records with relatively high texture gradients and color saturations, annotate the hierarchical details in a visual feature result index, and form visual feature data together with the detected main edge direction information. The entire process requires reading the scores item by item and directly comparing them with the threshold, and then classifying them in combination with the actual ranges of basic parameters such as texture and color. In the project, images with scores exceeding 60 will be centrally marked in the high saliency group, images with scores between 30 and 60 will be marked in the medium saliency group, and images with scores less than 30 will be marked in the low saliency group. After grouping, archive the image data within each group, and record the image ID, score value, and group label in the final index to obtain the visual feature data.
[0099] The steps for obtaining the fused feature vector are as follows: perform a one-to-one associative mapping between the visual feature data and the device type in the image metadata, construct the correspondence between visual features and device types, and form a feature mapping set;
[0100] According to the feature mapping set, calculate the fusion compactness of the image feature vector. The calculation formula is:
[0101]
[0102] where, Z j represents the fusion compactness of the feature vector of image j; u j represents the texture complexity of the visual features of image j; d j represents the color diversity of the visual features of image j; f j represents the edge density of the visual features of image j; m j represents the pixel resolution of the device type corresponding to image j; m d represents the average pixel resolution of all device types;
[0103] Select images according to the fusion compactness of the feature vector, fuse the feature vectors of visual features and device types, and generate a fused feature vector.
[0104] Specifically, after obtaining the device type in the visual feature data and image metadata, first read the information such as texture features, color features, and edge features produced in the image recognition stage item by item. At the same time, parse the unique number of the corresponding model from the device type registration form, match these visual features with the model number one by one, and mark them in the same index table for subsequent calls. Then, retrieve the more detailed attribute records such as the model, resolution, and photosensitive element size in the device type registration form, and associate these attributes with the visual features of the aforementioned image to form a feature mapping set. To ensure that there are no duplicate or missing records in the mapping process, it is necessary to scan all images in sequence and verify whether the image ID and the model number are correctly corresponding. If the model numbers corresponding to the image ID are the same, add a row in the mapping set and record the visual features and all related attributes of the model. If it is found that a certain image lacks a model number, data troubleshooting should be carried out during execution to determine its device type information and complete the mapping write. During the entire scanning process, the resolution value or the device type status should be detected. If the resolution exceeds the effective resolution range of the known device, a mark can be added in the mapping table and reviewed item by item later. For example, when the registered resolution of an image exceeds the range between ten million pixels and twenty million pixels, it can be classified into the high-resolution device type row, and the actual model can be determined by comparing the model resolution item in the registration form. After the mapping is completed, sort all the mapping records once and save them in the feature mapping set, and finally obtain the complete mapping index of all images, their visual features, and the corresponding device types.
[0105] The benefit of the formula is that it incorporates the numerical features of the image in terms of texture, color, and edge, as well as the pixel resolution factor of the shooting device, into the same metric. By simultaneously focusing on visual features and device hardware parameters, it improves the rigor of the fusion and the multi-dimensional coupling degree of the feature vectors.
[0106] u j represents the texture complexity in the visual features of image j, which is usually set in the range of 0 to 100 or a larger interval, and is evaluated by detecting the frequency of gradient changes in the image area. The formula can be defined as: where ΔG α represents the number of gradient intensity transitions in the α-th block area, p is the total number of sub-blocks. To make this value comparable among different images, it is necessary to perform gradient calculations under the same lighting conditions or the same preprocessing strategy, and ensure that the shutter speed and sensitivity during shooting are maintained consistent. In the actual process, the image analysis stage can divide the image main body into several sub-blocks, count the gradient jump points for each sub-block, and then add them up to obtain the value of u j In a certain implementation, there are 25 sub-blocks in total, and the cumulative number of gradient jump points is 520, then u j = 520.
[0107] d jRepresents the color diversity in the visual features of image j, which generally can also be taken within the range of 0 to 100, and is measured by detecting the number of main color gamuts and the average dispersion in the image. The specific method is to first convert the image to a color space where color components are more easily separated, then count the color clusters with obvious differences in the image, and calculate the dispersion degree of the distribution of these color clusters. Normalize the result into a numerical form, and the following formula can be established: where C β is the dispersion measure of the β-th color cluster, q is the number of detected color clusters. To obtain more accurate data, it is necessary to avoid superimposed shadow or overexposed areas during image segmentation, and use the same threshold to judge color differences. If 10 main color clusters are detected in an image during a certain acquisition, and the dispersion measures are 4, 6, 6, 8, 5, 7, 9, 7, 6, 4 respectively, then the sum is 62, and dividing by 10 gives d j = 6.2.
[0108] f j Represents the edge density of the visual features of image j, and is evaluated by calculating the number of edges within the range of unit area or unit pixel. It can be defined as: where E total represents the total number of edge pixels within the main body range of the image, A represents the total area of the main body pixels of the image. To obtain a more accurate result, it is necessary to use a consistent segmentation algorithm in the early stage to identify the main body range of the object and count the number of edge pixels among them. If the main body area of an object in an image occupies 4000 pixels and 200 pixels are detected as edge pixels among them, then
[0109] m j represents the pixel resolution of the device type corresponding to image j, which is mostly registered in units of pixel values or megapixels (MP). For example, a device with a resolution of 4000×3000 pixels can be statistically recorded as a nominal value of 12 million pixels, or directly recorded as 12 million in numbers as m j , and the acquisition method can be directly checked in the device parameter table or the EXIF information of the image. Each device will have an official pixel nominal value when leaving the factory, or it can be calculated based on the width and height pixels of the actually captured image and compared with the list. If the project contains multiple resolution models, the corresponding resolutions of each model will be recorded in a unified table and mapped with the image ID for direct reading during calculation. During a certain operation, the nominal value of a mobile phone is 12 million pixels, then m j = 12.
[0110] m dRepresents the average pixel resolution for all device types, which can be determined by simple arithmetic mean or weighted mean, usually in megapixels (MP). If there are several device types with resolutions of 8 million, 12 million, 20 million, and 32 million pixels respectively, and there are 100 devices each accounting for different quantities, then m can be obtained by multiplying each resolution by the corresponding number of devices and then dividing by the total number of devices. d , for example, there are 20 devices with 8 million pixels, 40 devices with 12 million pixels, 30 devices with 20 million pixels, and 10 devices with 32 million pixels. Then the weighted sum can be calculated first:
[0111] 20×8 + 40×12 + 30×20 + 10×32 = 160 + 480 + 600 + 320 = 1560, and the total number of devices is 20 + 40 + 30 + 10 = 100. Record m d = 15.6.
[0112] For example, substituting into the formula gives 31.13. This result indicates that for an image with relatively high levels of texture complexity, color diversity, and edge density, and a device resolution slightly higher than the average level, it has a relatively high value in the fusion tightness index. When Z j exceeds 20, it can be considered that the fusion tightness is relatively prominent and can be preferentially included in the subsequent screening stage.
[0113] Based on the fusion tightness of the feature vectors obtained previously, traverse the values of all images and record the corresponding fusion tightness sizes in the index table, so as to judge the overall performance of each image at the level of the combination of visual features and device types. Then, extract the images with values exceeding the established threshold into the preliminary selection list. If the tightness exceeds 20, append the image to the high-fusion group and record its texture complexity, color diversity, and device pixel resolution. If the tightness is lower than 10, label the image to the low-fusion group and conduct subsequent supplementary verification. For images between 10 and 20, uniformly classify them into the medium-fusion group and retain them in the intermediate index. After the preliminary grouping, select the images in the high-fusion group and the medium-fusion group, summarize the visual features of these images and the corresponding device attributes in a unified order, then package this information into a fusion feature vector for output, and write it into the fusion result table together with the image numbers to form the final set of fusion feature vectors.
[0114] The steps to obtain the index construction result are as follows: Based on the fusion feature vector, according to the numerical distribution characteristics of the fusion feature vector, divide the numerical range of the feature vector, and establish index nodes in each numerical range of the feature vector. Set the unique identifier of the index node, mark and sort the index nodes to form the initial index node set;
[0115] Based on the initial index node set, extract the numerical values of the fused feature vectors one by one, perform the matching between the numerical values and the index node identifiers, and according to the matching results, establish the mapping association between the fused feature vectors and the index nodes, and gradually connect all the mapping associations to generate an inverted index mapping structure;
[0116] Based on the inverted index mapping structure, call the node identifiers of the initial index node set, verify the access order of the fused feature vectors one by one according to the mapping association results between the fused feature vectors and the index nodes, and connect the index nodes in sequence according to the access order to construct a complete retrieval path for the fused feature vectors, and obtain the index construction result.
[0117] Specifically, based on the fused feature vectors obtained previously, first retrieve the numerical range of all the fused feature vectors in the data registration form. For example, record the minimum value as a1 and the maximum value as a2, and set an interval number n1 according to the distribution observation of the image quality during on-site testing. Subtract a1 from a2 to get a total span, and then divide this span into n1 parts. For example, evenly divide the interval between a1 and a2 into several segments, mark the interval boundaries at the beginning and end of the numerical interval corresponding to each segment, and then create an index node for each interval and assign a unique identifier. The identifier can be in the form of an integer or alphanumeric. When dividing the interval, also record the interval size and compare it with the actual sample quantity to judge whether the distribution quantity of the feature vectors contained in each interval matches the preset range. If it is found that the numerical values of the feature vectors in a certain interval are significantly concentrated and exceed the specified threshold. For example, the threshold can be set to 10% of the total number of records, then this interval can be further subdivided into smaller partitions during the division to balance the sample quantity of each interval. The threshold is calculated from the feature vector distribution data collected in the early stage. For example, by counting the fused feature vectors of all images, it is found that the probability of the concentrated distribution area appearing reaches 15%, so the threshold is set to 10% to ensure that the interval division has discrimination. Subsequently, arrange the index node identifiers corresponding to the intervals in ascending order of the starting numerical values of the intervals and record them in an initial node list. Each node has a sequential number and the actual numerical range. For example, node 1 corresponds to a1 to a1+(total span / n1), node 2 corresponds to a1+(total span / n1) to a1+2*(total span / n1), and so on. Sort all the index nodes in sequence and mark the upper and lower limits of the node. Finally, merge them into a node set arranged in ascending order of the identifiers to form the initial index node set.
[0118] Based on the initial index node set obtained previously, scan the fused feature vector values in the feature vector index table one by one and perform paired matching. Compare the specific value of each fused feature vector with the lower and upper limits of the numerical range in the node set to determine which interval the vector lies in. Then directly record the matching result of the vector and the node in the index analysis table. If the value of a certain vector falls within the overlapping range of multiple intervals, precise comparison will be made according to the interval boundaries to determine the final node attribution. When the overlapping range is less than 0.01 or other preset error thresholds, the vector can be merged into the node that is closer to it. The setting of the error threshold is obtained based on the evaluation of the sampling accuracy and numerical fluctuation. For example, in multiple rounds of experiments, it is found that 0.5% of the vector data will be close to the dividing line. Therefore, the threshold is set to 0.01 to ensure more detailed difference discrimination. After the matching is completed, an associated entry of each vector and the corresponding index node will be established in the mapping record. Through one traversal, all the mappings between vectors and nodes can be recorded, and these mappings will be continuously connected in the matching order to form an inverted index mapping structure. During the process, the number of vectors included in different node intervals will be compared and the quantity statistics will be updated in the node information. If the number of vectors of a certain node is higher than 25 or other empirical values, a special mark will be made for this node in the record table for future re-splitting or hierarchical retrieval. Finally, the generated inverted index mapping structure stores the correspondence between the fused feature vectors and the node identifiers for subsequent operation calls.
[0119] Based on the inverted index mapping structure obtained previously, read the list of node identifiers in the initial index node set and query the corresponding fused feature vector set in ascending order of node identifiers. Check the sequential arrangement of each feature vector in the mapping table, and correspond the index nodes with the vectors to form access paths one by one. To verify the access order, the values of the vectors will be read one by one and compared with the requirements of the predefined retrieval instructions. For example, the retrieval instructions can be divided into several range queries and compared with the node boundaries. When the lower limit to the upper limit of the node meets the retrieval range, the vectors recorded under the current node will be output at once. If the retrieval instructions require a more precise range, the difference calculation will be used during the comparison to determine whether it falls within the specified interval. If the difference between the node numerical range and the retrieval interval is less than a certain empirical threshold, such as 0.05, the corresponding vector will be directly retained. The threshold setting refers to the observed data of the previous retrieval success rate and hit accuracy. For example, in the last retrieval statistics, it is found that 0.05 for the interval difference can balance the needs of identification and error control. Therefore, 0.05 is selected as the empirical threshold. After traversing all the nodes, these access orders will be connected in sequence to form a complete link of the retrieval path and record each node identifier and the corresponding vector set number in a path table, thereby obtaining the index construction result.
[0120] The steps to obtain the similarity sorting result are as follows: According to the index construction result, calculate the similarity matching score of the fused feature vector. The calculation formula is:
[0121]
[0122] where Y u is the similarity matching score of the fused feature vector u, and x u , y u , w u are the texture feature value, color feature value, and edge feature value of the fused feature vector u respectively. x q , y q , w q are the texture feature value, color feature value, and edge feature value corresponding to the feature vector to be queried respectively. m u is the lens pixel value of the device type corresponding to the fused feature vector u, and m q is the lens pixel value of the device type corresponding to the feature vector to be queried;
[0123] According to the similarity matching score, sort the fused feature vectors from largest to smallest, determine the sorting positions of the fused feature vectors one by one, and connect the sorted fused feature vectors in sequence to obtain the similarity sorting result.
[0124] Specifically, the advantage of the formula is that by introducing the comprehensive Euclidean distance value of the texture, color, and edge features of the fused feature vector in the numerator part and combining the difference amount of the lens pixel values in the denominator part for adjustment, the differences in image detail features and the hardware specifications of the shooting device are taken into account during the similarity calculation, which can effectively reflect the matching degree between multi-dimensional features and provide a more discriminative scoring basis for image sorting or retrieval.
[0125] x u represents the texture feature value of the fused feature vector u. To accurately obtain this value, the texture gradient, contrast distribution, etc. will be quantitatively measured in the previous image analysis session, and the measurement results will be summarized into numerical indicators, usually in the range of 0 to 100. The higher the value, the more obvious the image texture complexity or detail sharpness. If the average value of the gradient variance is found to be large after dividing an image into blocks during a certain test, then its texture feature value x u may reach 80 or 90.
[0126] y u represents the color feature value of the fused feature vector u. The specific number can also be in the range of 0 to 100, which is obtained by fusing and calculating indicators such as the main color channel distribution, color saturation, and color dispersion of the image. For example, if the color distribution of an image is relatively full and the detected dispersion value is higher than the normal level, then y u will reach 70 or 80.
[0127] w u represents the edge eigenvalue of the fusion feature vector u, emphasizing the edge sharpness and density of the main body or multiple objects in the image. It is comprehensively evaluated by detecting the gray - level gradient and transition frequency of edge pixels in image analysis. Generally, this value can also be placed in the range of 0 to 100 for unified measurement of images. After edge detection in the early stage and recording information such as the number of pixels and sharpness distribution of the object contour, a numerical value reflecting the overall edge feature is obtained. In practice, it can be compared with existing sample images. For example, for images with extremely high sharpness, it is set between 90 and 100, and for images with more blurred areas, it is set between 0 and 30 to ensure the distinction of images with different qualities.
[0128] x q 、y q 、w q respectively represent the texture eigenvalue, color eigenvalue, and edge eigenvalue of the feature vector to be queried. These three numerical values are related to x u 、y u 、w u with the same meaning, but corresponding to the target vector data used at the query end. Usually, when retrieving, the user provides a reference image or a set of visual feature parameters, and then the system performs the same texture analysis, color statistics, and edge detection processes on this reference image to obtain the query vectors x q 、y q 、w q . The structure of this query vector is consistent with that of the fusion feature vector, so one - to - one difference calculations can be performed. For example, if a user submits an image mainly of an outdoor landscape, after extracting a texture value of about 60, a color value of about 70, and an edge value of about 40, they can be used as x q 、y q 、w q for the next similarity comparison. The whole process is based on the premise that the shooting equipment and lighting environment are as unified or calibratable as possible to ensure that the query values are within the same scale system and match the fusion feature vectors stored in the database.
[0129] m u represents the lens pixel value of the device type corresponding to the fusion feature vector u, which is mainly obtained by sorting out the image device information in the early stage of the project. Each shooting device has a nominal or measured lens pixel value, usually expressed in megapixels or more specific pixel numbers.
[0130] m q represents the lens pixel value of the device type corresponding to the feature vector to be queried, and is the same as m uCorrespondingly, if the query image uploaded by the user comes from a registered device, the pixel values can be directly read from the device registration record, or the model number of the shooting device can be extracted by parsing the image metadata and the matching resolution value can be found in the model library. The usage method is the same as that of m u , after numericalizing the resolution value, perform a difference operation in the denominator part of the formula to balance the image feature deviation caused by hardware differences. In some scenarios, the pixel value of the query device may also be in the range of 10 to 30. For example, if a query image comes from a device with 18 million pixels, it is recorded as m q = 18.
[0131] For example, substituting into the formula to calculate, the similarity score is at the level of 4.082. When a certain threshold is set to 5 or 6 in the project, it can be determined accordingly that the matching degree between the current fused feature vector and the query vector is medium. If it is greater than 6, it means that they are very close in terms of texture, color, edge features and pixel differences. If it is less than 2, it means that the differences are large. These numbers can be set and corrected according to the distribution range of the collected samples.
[0132] According to the similarity matching scores obtained previously, sort the fused feature vectors from largest to smallest, read the score values of all fused feature vectors one by one and record the image identifier and device information corresponding to each score during the sorting process. Then, sort the vectors with similar scores adjacent to each other and separate the vectors with significantly different scores to a farther position. When implementing the sorting, compare the difference between the score of the current vector and the score of the last vector in the existing sorting one by one. If the difference is greater than a certain empirical threshold, insert this vector directly into the position with a lower score. For example, set the threshold to 0.5 and calculate the difference between the previous score and the next score during the comparison. If the difference is 0.8, it is considered that they are not in the same level interval. Then, check sequentially whether a closer score paragraph is found. After the score is determined, arrange all the fused feature vectors associated with the score in sequence and construct the entire list continuously. During the process, the established threshold is also used to prevent local scores from being too dense or too sparse. For the situation where the score differences of multiple consecutive vectors are in the interval of 0.1 to 0.2, these vectors can be grouped into a set with relatively close similarity scores so that the required images can be located more quickly during subsequent retrieval. After all the sorting is completed, the similarity sorting result is obtained.
[0133] The steps to obtain the user query result are as follows: According to the similarity sorting result, call the corresponding index nodes of the sorted fused feature vectors in the index construction result, locate the associated image identifiers of the index nodes in sequence, extract the set of image identifiers corresponding to the query matching results, and generate an associated image identifier list;
[0134] Extract the corresponding image data according to the associated image identifier list, call the visual feature data and metadata field information of the image data, perform content collation and structural adjustment of the query results, and present the image data in the order of the query matching results to generate the user query results.
[0135] Specifically, according to the similarity sorting results obtained previously, obtain the index node information of each fusion feature vector in the sorted list. Read the corresponding node identifiers one by one by referring to the previously established node identifier index table. Then compare these identifiers with the image identifier mapping table to determine which specific image IDs the fusion feature vectors with the highest similarity belong to. During the reading process, first select the first item in the sorted list and retrieve its index node identifier, and then find the image IDs recorded under the same node in the mapping table. Record this ID in a temporary set, and then scan each sorted entry downward in turn. If it is found during recognition that the number of images contained in a certain node exceeds the set threshold, for example, if the set threshold is 20 images, then after scanning, all image IDs under this node need to be grouped once to avoid over-concentration of data. The threshold is determined based on the statistics of the relationship between nodes and images in actual storage. For example, after statistically analyzing the distribution range of the number of images under all nodes, it is found that when it exceeds 20, it will cause fluctuations in the retrieval speed. Therefore, it is used as the splitting reference value. If it is less than or equal to 20, it remains unchanged. Each time the image IDs contained under the node are read, these IDs are written into an associated image identifier list and arranged in ascending order of similarity in the list. Compare the similarity scores of adjacent records to determine whether the list is smoothly connected. If the score difference is less than 0.5, these images are classified into adjacent groups. If the difference is greater than 0.5, they are separated into separate segments. The number 0.5 is set by referring to the distribution data of the image score adjacency in past query results and is confirmed through multiple tests to better distinguish image entries with large differences. After scanning all nodes, the obtained associated image identifier list is sorted and output from the highest score to the lowest score, and check whether there are a small number of images with abnormal score distributions in the last few nodes. If necessary, supplementary verification or removal can be performed, and finally an associated image identifier list is generated.
[0136] According to the obtained associated image identifier list, read the original image data corresponding to each image ID in sequence, including the image file name, the visual feature data recorded during image acquisition, and the metadata fields. For example, query the texture gradient, color saturation, or edge sharpness information of the image in the visual feature table and read the geographical location information and shooting time, etc. in the metadata field information. To ensure the orderly sorting of information, first match the associated image identifier list with the queried image information item by item to obtain an item set, and then sequentially list this set and check whether its order is consistent with the order corresponding to the similarity matching result. If it is found that some images have a status conflicting with the current query requirements in the resolution or other metadata fields (for example, the required resolution is greater than 10MP while this image is only 5MP), it can be marked as non-compliant and temporarily skipped during the matching process. The resolution threshold can be set to 10MP during project research and determined according to the general configuration of the shooting device. After all the image entries that pass the verification are arranged, the query result list can be output. This list presents the visual features and metadata fields of each image in sequence and lists them in descending order of score. If further refinement is required, only several attribute fields associated with the query target, such as texture values, color values, and edge values, as well as the time and location fields in the metadata, can be retained during presentation. Finally, the entire list is displayed centrally or a paged retrieval result is generated in the front-end system to generate the user query result.
Claims
1. An image retrieval method based on big data mining, characterized in that: The following steps are involved: Collect images and metadata, including shooting time, location, and device type, filter the metadata fields, determine the candidate image set, and obtain metadata filtering results; Classifying the metadata screening results according to locations and timestamps, clustering the locations through a geographic information system, and obtaining geographic clustering results; Based on the geographic clustering results, identify objects and scenes in the image, extract visual features, and generate visual feature data; combine the visual feature data with the device type in the metadata, perform feature fusion, optimize the feature vector through a weighted algorithm, and obtain a fused feature vector; Based on the fused feature vector, an index structure is constructed, the feature vector is indexed using an inverted index, a search path is established, and an index construction result is obtained; a similarity matching query is performed on the index construction result, and the result is sorted according to a similarity scoring standard to obtain a similarity sorting result; The similarity ranking result is used to respond to user queries, and according to the query conditions input by the user, related images are extracted from the ranking results for result presentation to generate user query results.
2. The image retrieval method based on big data mining according to claim 1, characterized in that: The steps for obtaining the metadata screening results are: Perform consistency check on the shooting time, shooting location and device type fields in the collected image metadata, determine the valid metadata fields based on the metadata field consistency check results, and obtain a valid metadata field set; According to the valid metadata field set, the metadata validity score of each image is calculated using the following formula: Among them, M i represents the metadata validity score of image i, T i represents the consistency test result of the shooting time field of image i, L i represents the consistency test result of the shooting location field of image i, E i Represents the consistency check result of the device type field of image i, F i Represents the difference between the shooting time field data of image i and the standard shooting time, G i Represents the spatial position difference between the latitude and longitude coordinates of the shooting location field of image i and the standard coordinates, H i Represents the difference between the image i device type field and the mainstream device type match; Based on the metadata validity score, the images are screened and selected to form a metadata screening result.
3. The image retrieval method based on big data mining according to claim 1, characterized in that: The steps for obtaining the geographic clustering results are: Performing spatiotemporal classification according to the shooting location and shooting timestamp of the metadata screening result, defining an initial location classification area according to the location coordinate spatial distance, and forming an initial spatial classification result; Based on the initial spatial classification results, the spatial clustering fitness score between the location to which each image belongs and the location classification area to which it belongs is calculated, and the calculation formula is: Among them, P j represents the spatial clustering fitness score of image j, X j , Y j Represent the longitude and latitude coordinate values of image j, X c , Y c They represent the longitude and latitude coordinate values of the center point of the classification area where image j belongs, t j Represents the timestamp value of image j, t c Represents the standard timestamp value of the location classification area to which image j belongs; Based on the spatial clustering fitness score, the clustering area of the image is re-divided to form a geographic clustering result.
4. The image retrieval method based on big data mining according to claim 1, characterized in that: The steps for obtaining the visual feature data are as follows: According to the geographic clustering results, for each clustered image, the object category and the scene category in the image are respectively identified, and the texture details, color distribution and edge definition characteristics of the object category are determined to obtain image object feature information; According to the image object feature information, the scene saliency score of the image visual feature is calculated, and the calculation formula is: Among them, Q j represents the scene saliency score of image j, c j represents the color saturation of the object category of image j, s j represents the texture detail clarity of the object category in image j, e j represents the edge clarity feature of the object category in image j, h j represents the uniformity of the main color tone of the object category in image j, b j represents the main edge direction of the object category in image j, b s The main edge direction representing the standard of the scene category to which image j belongs; Visual features are extracted according to the scene saliency score to form visual feature data.
5. The image retrieval method based on big data mining according to claim 1, characterized in that: The steps of obtaining the fused feature vector are: Performing one-to-one associative mapping between the visual feature data and the device type in the image metadata, constructing a corresponding relationship between the visual feature and the device type, and forming a feature mapping set; According to the feature mapping set, the image feature vector fusion density is calculated, and the calculation formula is: Among them, Z j represents the fusion density of the feature vector of image j; u j Represents the texture complexity of the visual features of image j; d j represents the color diversity of the visual features of image j; f j represents the edge density of the visual features of image j; m j represents the pixel resolution of image j corresponding to the device type; m d Represents the average pixel resolution across all device types; An image is selected according to the feature vector fusion density, and the feature vector of the visual feature and the device type is fused to generate a fused feature vector.
6. The image retrieval method based on big data mining according to claim 1, characterized in that: The steps for obtaining the index building result are: Based on the fused feature vector, according to the numerical distribution characteristics of the fused feature vector, the numerical interval of the feature vector is divided, and an index node is established in each numerical interval of the feature vector, a unique identifier of the index node is set, and the index nodes are marked and sorted to form an initial index node set; Based on the initial index node set, extract the values of the fused feature vectors one by one, perform matching between the values and the index node identifiers, and establish a mapping association between the fused feature vectors and the index nodes according to the matching results, gradually connect all the mapping associations, and generate an inverted index mapping structure; Based on the inverted index mapping structure, the node identifier of the initial index node set is called, and the access order of the fused feature vector is verified one by one according to the mapping association result between the fused feature vector and the index node, and the index nodes are connected in sequence according to the access order to construct a complete fused feature vector retrieval path and obtain the index construction result.
7. The image retrieval method based on big data mining according to claim 1, characterized in that: The steps for obtaining the similarity ranking result are: Calculate the similarity matching score of the fused feature vector according to the index construction result; According to the similarity matching score, the fused feature vectors are sorted from large to small, the sorting positions of the fused feature vectors are determined one by one, and the sorted fused feature vectors are connected in sequence to obtain a similarity sorting result.
8. The image retrieval method based on big data mining according to claim 1, characterized in that: The steps for obtaining the user query results are: According to the similarity sorting result, calling the corresponding index node of the sorted fusion feature vector in the index building result, locating the associated image identifiers of the index nodes in turn, extracting the image identifier set corresponding to the query matching result, and generating an associated image identifier list; According to the associated image identification list, the corresponding image data is extracted, the visual feature data and metadata field information of the image data are called, the content and structure of the query results are sorted and adjusted, the image data is presented in the order of the query matching results, and the user query results are generated.
Citation Information
Cited By
Brain-like AI architecture-based side computing power acceleration method and apparatus, and computer device
CN121233306A
Travel data management method and system based on cloud computing and scene matching, and computing equipment
CN121707499A
Travel data management methods, systems, and computing devices based on cloud computing and scenario matching
CN121707499B
Travel data management method and system based on cloud computing and timestamp, and computing device
CN121788154A