Image management method and device

By using feature dimensionality reduction, label recognition and tag fusion algorithms in cloud disk services, high-precision image tags are generated, which solves the problem of insufficient image classification and labeling accuracy in existing cloud disk services, and significantly improves the user's retrieval experience.

CN119938959APending Publication Date: 2025-05-06CHINA MOBILE INTERNET CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510021378.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing cloud disk services have problems such as insufficient accuracy and insufficient semantic expression in image classification and labeling, resulting in poor user experience when retrieving images.

Method used

By using the preset feature dimensionality reduction algorithm to extract the feature set of the target image, combining the preset tag recognition algorithm and the tag fusion algorithm, high-precision multi-level tags are generated and associated with the image to store.

Benefits of technology

It improves the efficiency and accuracy of image tag generation and storage, optimizes the correlation and expression capabilities between tags, and significantly improves the user's cloud disk experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938959A_ABST
    Figure CN119938959A_ABST
Patent Text Reader

Abstract

The invention discloses an image management method and device. The method comprises the following steps: in response to an operation of uploading a target image to a cloud disk, performing feature dimension reduction extraction on the target image by using a preset feature dimension reduction algorithm to obtain a target feature set; performing label identification on the target image according to the target feature set by using a preset label identification algorithm to obtain a plurality of labels corresponding to the target image; performing semantic fusion on the plurality of tags by using a preset tag fusion algorithm to obtain a target tag; and in the cloud disk, associatively storing the target label and the target image. According to the method, the target image uploaded by the user is intelligently processed through feature dimension reduction, tag identification and semantic fusion algorithms, the target tag with refined semantics is generated, and the target tag and the image are associated and stored in the cloud disk, so that the management degree of the image in the cloud disk is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of cloud computing technology, and in particular, relates to an image management method and device. Background Art

[0002] Existing cloud disk services have problems with insufficient precision and inadequate semantic expression in image classification and annotation. The generated annotations are often too simple and cannot fully reflect the image content. This defect directly leads to a poor user experience when searching for images, making it difficult for users to find the required content quickly and accurately.

[0003] Therefore, how to design a high-precision, multi-level image classification and annotation solution to improve user retrieval efficiency and experience has become a problem that needs to be solved urgently. Summary of the invention

[0004] The embodiments of the present application provide an image management method and device, which can improve the level of cloud disk image management.

[0005] In a first aspect, an embodiment of the present application provides an image management method, the method comprising:

[0006] In response to the operation of uploading the target image to the cloud disk, a preset feature dimension reduction algorithm is used to perform feature dimension reduction extraction on the target image to obtain a target feature set;

[0007] Using a preset label recognition algorithm, the target image is label recognized according to the target feature set to obtain multiple labels corresponding to the target image;

[0008] Use the preset label fusion algorithm to semantically fuse multiple labels to obtain the target label;

[0009] In the cloud disk, the target label is associated with the target image and stored.

[0010] In a second aspect, an embodiment of the present application provides an image management method, the method comprising:

[0011] Receive an operation to upload the target image to the cloud disk;

[0012] According to the operation, the target image is sent to the cloud disk server so that the cloud disk server executes the method described in the first aspect above.

[0013] In a third aspect, an embodiment of the present application provides an image management device, the device comprising:

[0014] A feature extraction module, configured to, in response to the operation of uploading the target image to the cloud disk, perform feature dimensionality reduction extraction on the target image using a preset feature dimensionality reduction algorithm to obtain a target feature set;

[0015] A label recognition module is used to use a preset label recognition algorithm to perform label recognition on a target image according to a target feature set to obtain multiple labels corresponding to the target image;

[0016] A label fusion module is used to perform semantic fusion on multiple labels using a preset label fusion algorithm to obtain a target label;

[0017] The management module is used to associate the target tag with the target image and store it in the cloud disk.

[0018] In a fourth aspect, an embodiment of the present application provides an image management device, the device comprising:

[0019] A receiving module, used for receiving an operation of uploading a target image to a cloud disk;

[0020] The sending module is used to send the target image to the cloud disk server according to the operation, so that the cloud disk server executes the method described in the first aspect above.

[0021] In a fifth aspect, an embodiment of the present application provides a computer device, comprising: a processor, and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the method of the first aspect above.

[0022] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the method of the first aspect described above is implemented.

[0023] In a seventh aspect, an embodiment of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0024] The technical solution provided by the embodiments of the present application brings at least the following beneficial effects:

[0025] The image management method provided in the embodiment of the present application performs intelligent processing on the target image uploaded by the user through feature dimension reduction, label recognition and semantic fusion algorithms, generates semantically refined target labels, and associates them with the images and stores them in the cloud disk. By extracting key feature sets through feature dimension reduction, the efficiency and accuracy of label recognition are improved; semantic fusion optimizes the correlation and expression ability between labels, making the labels more coherent and semantically rich. This solution can realize the efficient generation and storage of image labels, improve the degree of automation of image management, provide a more accurate and intelligent foundation for subsequent retrieval, classification and recommendation operations, and significantly optimize the user's cloud disk experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application.

[0027] Figure 1 It is a flowchart of an image management method provided in an embodiment of the present application;

[0028] Figure 2 It is a flowchart of an image management method provided in an embodiment of the present application;

[0029] Figure 3 It is a flowchart of an image management method provided in an embodiment of the present application;

[0030] Figure 4 It is a flowchart of an image management method provided in an embodiment of the present application;

[0031] Figure 5 It is a flowchart of an image management method provided in an embodiment of the present application;

[0032] Figure 6 It is a flowchart of an image management method provided in an embodiment of the present application;

[0033] Figure 7 It is a flowchart of an image management method provided in an embodiment of the present application;

[0034] Figure 8 It is a flowchart of an image management method provided in an embodiment of the present application;

[0035] Fig. 9 is a structural diagram of an image management device provided in an embodiment of the present application;

[0036] Fig.10 is a structural diagram of an image management device provided in an embodiment of the present application;

[0037] Fig.11 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to make those of ordinary skill in the art better understand the technical scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.

[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples consistent with some aspects of the present application as detailed in the attached claims.

[0040] As mentioned in the background technology section, existing cloud disk services widely use traditional methods based on feature extraction in image classification and annotation. These technologies can achieve basic classification and label generation, and provide users with preliminary image management and retrieval functions. However, these technologies have obvious shortcomings in practical applications: low classification accuracy, difficulty in accurately distinguishing images that are visually similar but semantically different; simple annotation content, lack of in-depth understanding of image semantics; retrieval results often have low recall and precision rates, and poor user experience.

[0041] In order to solve the above technical problems, the embodiments of the present application provide an image management method and device.

[0042] The image management method provided in the embodiment of the present application is described in detail below through specific embodiments in conjunction with the accompanying drawings.

[0043] Figure 1 A flow chart of an image management method provided by an embodiment of the present application is shown.

[0044] like Figure 1 As shown, the method may include the following steps.

[0045] S110 , in response to the operation of uploading the target image to the cloud disk, performing feature dimensionality reduction extraction on the target image using a preset feature dimensionality reduction algorithm to obtain a target feature set.

[0046] The target image refers to the image that the user uploads to the cloud disk and needs to be classified or labeled.

[0047] It should be noted that the target image can be a static picture or a dynamic image composed of continuous frames in a video.

[0048] Cloud disk refers to an online storage platform that provides storage and management of data such as files and images. On this platform, users can upload, download, manage and share data.

[0049] It is understandable that the execution subject of this embodiment may be a cloud disk server, which is responsible for processing the target image uploaded by the user, performing feature dimensionality reduction extraction and subsequent image classification or labeling tasks. The operation of uploading the target image to the cloud disk can be triggered in a variety of ways.

[0050] In one implementation, the uploading of the target image may be actively performed by the user, for example, the user selects and manually uploads an image file that needs to be classified or labeled.

[0051] In one implementation, the upload of the target image may also be automatically triggered by the system. For example, in some application scenarios, the terminal device may automatically detect a newly generated image or video frame according to preset conditions and upload it to the cloud disk. Such automatic upload operations can occur in scenarios such as image generation, video recording, and real-time monitoring, without user intervention, and automatically perform image upload according to specific rules or events.

[0052] Regardless of whether the upload operation is actively performed by the user or automatically triggered by the system, the cloud disk serves as a data storage and processing platform to receive and manage these target images to ensure that subsequent feature extraction, classification or labeling tasks can be carried out smoothly.

[0053] Feature dimensionality reduction extraction includes two steps: feature extraction and feature dimensionality reduction.

[0054] The target feature set refers to the feature set extracted from the target image and reduced in dimension, which is used as the basis for subsequent classification or labeling.

[0055] In this step, when the user uploads the target image, the cloud disk server uses a preset feature dimensionality reduction algorithm to process the image, extract the significant features of the image, remove irrelevant redundant information, and generate a target feature set.

[0056] In one implementation, this step can be implemented based on a convolutional neural network:

[0057] The target image is processed through a pre-trained convolutional neural network to extract high-dimensional deep features.

[0058] The dimensionality reduction algorithm is used to reduce the dimensionality of high-dimensional depth features, remove redundant information and improve processing efficiency, and obtain a low-dimensional feature set of the target image.

[0059] This implementation method obtains a more abstract and semantic feature representation through a convolutional neural network, and directly generates a target feature set after dimensionality reduction. The process is relatively simple, but it relies on the performance of the deep model and the coverage of the training data.

[0060] In one implementation, this step can be implemented based on principal component analysis and graph clustering methods:

[0061] Perform principal component analysis on the target image and extract the local features of the target image.

[0062] A feature similarity graph is constructed with local features as nodes and the similarities between local features as edge weights.

[0063] Based on the feature similarity graph, the local features are divided into multiple clusters, and representative features are screened from each cluster to obtain the target feature set.

[0064] This implementation uses principal component analysis to reduce feature dimensionality and then extracts representative features through graph clustering, making the target feature set more compact while retaining the core semantic information of the image.

[0065] S120 , using a preset label recognition algorithm, and according to a target feature set, performing label recognition on a target image to obtain a plurality of labels corresponding to the target image.

[0066] The preset label recognition algorithm can be a multi-label classification model, a semantic embedding similarity matching model, etc.

[0067] In one implementation, this step can be implemented based on a multi-label classification model:

[0068] Using the pre-trained multi-label classification model, the target feature set is input into the model to calculate the relevance score between the target image and each label.

[0069] According to the score threshold, multiple labels with higher correlation are selected as the label results of the target image.

[0070] This implementation, through multi-label learning, can simultaneously assign multiple semantically related labels to the target image.

[0071] In one implementation, this step can be implemented based on a method of semantic embedding and similarity matching:

[0072] The target feature set and preset labels are mapped into the same semantic space through semantic embedding technology, and the similarity between the two is calculated.

[0073] Select multiple labels with the highest similarity to the target feature set as the label results of the target image.

[0074] This implementation ensures that the output labels are more consistent with the image content by analyzing the consistency of feature and label semantic representations.

[0075] In one implementation, the identified tags may be classified:

[0076] Based on preset label categories, multiple labels corresponding to the target image are classified to obtain multiple label sets;

[0077] Based on multiple label sets, category labels corresponding to multiple categories to which the target image belongs are obtained.

[0078] Each tag set includes at least one tag. Each tag set corresponds to a preset tag category and can reflect different dimensional features of the target image content.

[0079] Exemplarily, the preset tag categories may include objects, scenes, emotions, and the like.

[0080] For example, when the target image is an image of a kitten, the preset label recognition algorithm will perform refined label recognition on the kitten image and obtain multiple labels, such as "cat", "kitten", "pet", "indoor", "sofa", "carpet", "cute", "warm", "cuddly", etc.

[0081] Furthermore, based on the preset three tag categories of objects, scenes, and emotions, multiple tags are classified to obtain the following multiple tag sets:

[0082] A collection of object labels: "cat", "kitten", "pet".

[0083] Scene label collection: "indoor", "sofa", "carpet".

[0084] A collection of emotional labels: "cute", "warm", and "lovable".

[0085] This implementation method, through hierarchical classification and refined label generation, can more comprehensively characterize the multi-dimensional information of the target image, which can not only meet the user's query needs for specific objects, but also support semantic retrieval based on scenes and emotions, thereby significantly improving the intelligence of cloud disk services and user experience.

[0086] S130: semantically fuse multiple tags using a preset tag fusion algorithm to obtain a target tag.

[0087] It should be understood that the target tag may refer to a tag combination obtained by fusing multiple tags.

[0088] The preset tag fusion algorithm may include a cross-modal tag semantic fusion algorithm.

[0089] The introduction of the cross-modal label semantic fusion algorithm is detailed in the corresponding embodiment below and will not be repeated here.

[0090] In one implementation, a label association network can be used to construct a semantic relationship graph between labels. In the semantic graph, the upper and lower relationships, synonymous relationships, and semantic distances between labels are analyzed, and similar or related labels are merged into a semantically consistent label combination through clustering or path optimization methods.

[0091] In one implementation, a cross-modal semantic embedding model can be used to map image feature labels and text labels into a shared semantic space, in which labels with high correlation are merged into a target label combination through vector clustering, weighted averaging or attention mechanism.

[0092] S140: Store the target tag in association with the target image in the cloud disk.

[0093] The target tags obtained through tag recognition and semantic fusion are associated with the target image and stored in the cloud disk in a structured or semi-structured manner.

[0094] This application solution uses feature dimension reduction, label recognition and semantic fusion algorithms to intelligently process the target images uploaded by users, generate semantically refined target labels, and associate them with the images and store them in the cloud disk. By extracting key feature sets through feature dimension reduction, the efficiency and accuracy of label recognition are improved; semantic fusion optimizes the correlation and expression ability between labels, making the labels more coherent and semantically rich. This solution can achieve efficient generation and storage of image labels, improve the degree of automation of image management, provide a more accurate and intelligent foundation for subsequent retrieval, classification and recommendation operations, and significantly optimize the user's cloud disk experience.

[0095] It is understandable that the existing methods lack the ability to adapt to data characteristics when performing feature analysis on images, and it is difficult to accurately capture the complex and diverse local pattern changes in images. Based on this, the present application constructs a multi-scale adaptive feature dimensionality reduction method to extract image features.

[0096] Figure 2 A flow chart of an image management method provided in an embodiment of the present application is shown.

[0097] It should be understood that Figure 2 The illustrated embodiment can be regarded as an example of step S110, which introduces the process of extracting image features by using a multi-scale adaptive feature dimensionality reduction method.

[0098] like Figure 2 As shown, the method may include the following steps.

[0099] S210: Perform principal component analysis on the target image to extract local features of the target image.

[0100] It is understandable that the target image contains multi-scale patterns from local to global, and the principal component analysis of a single scale is prone to miss important information. Therefore, the present application further extends the adaptive principal component analysis to multi-scale scenarios.

[0101] The implementation methods may include:

[0102] S211, performing gradient down-sampling recursively on the target image to obtain multiple sub-images with different resolutions.

[0103] For example, for the original image I i Perform recursive sampling under the gradient to obtain a set of sub-images Ii with different resolutions (l) l=1 L .

[0104] S212: Divide each sub-image into a plurality of non-overlapping local blocks.

[0105] For example, at each scale l, the sub-image is divided into m×n non-overlapping local blocks Each local block contains a small part of the image.

[0106] S213: Perform principal component analysis on each local block to extract the principal component features of each local block.

[0107] Exemplarily, principal component analysis is performed on each local block to extract K principal component features.

[0108] Traditional principal component analysis methods usually "flatten" the image into a vector first, and then calculate the global mean and covariance matrix, but this method ignores the local information in the image. Therefore, this application proposes an adaptive principal component analysis method, which introduces an adaptive weighting mechanism to make the calculation of the covariance matrix pay more attention to the local pattern of each image. The following is an introduction to the adaptive principal component analysis method:

[0109] The covariance matrix S after introducing the adaptive weighting mechanism a :

[0110]

[0111] N is the number of samples, that is, the total number of samples in the data set.

[0112] x i is the data vector of the i-th sample, representing the coordinates or eigenvalues ​​of the sample point.

[0113] μ is the weighted mean of the samples.

[0114] w i is the weight of the i-th sample, reflecting the importance of the sample in the weighted covariance calculation.

[0115] Among them, the weight w i Determined through the following adaptive learning process:

[0116] Initialize weight w i= 1 / N. This means that at the beginning of the algorithm, no sample is particularly emphasized and all samples contribute equally to the calculation of the mean and covariance.

[0117] Calculate the current weighted covariance matrix S a The maximum eigenvalue λ1 and the corresponding eigenvector v1.

[0118] λ1 is the maximum eigenvalue of the current weighted covariance matrix Sa, which describes the variance of the data in the main direction of the matrix and indicates the degree of change of the sample distribution in the main direction.

[0119] v1 is the eigenvector corresponding to the maximum eigenvalue λ1, which represents the characteristic direction of the data in the main direction (that is, the main axis direction of the covariance matrix).

[0120] Update the weight w of each sample i :

[0121]

[0122] β is used to control the concentration of weight distribution.

[0123] Repeatedly calculate the maximum eigenvalue λ1 and eigenvector v1 of Sa, and update the weight w i , until one of the following conditions is met:

[0124] Weight convergence means that the change of weight value tends to be stable.

[0125] The maximum number of iterations is reached, that is, the maximum number of iterations set.

[0126] The algorithm adaptively adjusts the contribution of each sample through iterative optimization, so that the covariance matrix pays more attention to the intrinsic structure of the data and extracts more discriminative principal component features.

[0127] S214, cascading the principal component features of each local block to obtain local features of the target image.

[0128] For example, local features of all scales are concatenated into a multi-scale local feature representation of the target image.

[0129]

[0130] That is, the principal component features of different scales and different local blocks are combined to form a comprehensive multi-scale feature representation.

[0131] The above implementation extracts image features in a multi-scale manner, avoiding local and global information that may be missed by single-scale principal component analysis. By performing local block processing on images at different resolutions and combining adaptive principal component analysis to extract the principal component features of each local block, a multi-scale image feature representation is finally obtained. This method can better capture local details and global patterns in images, and is suitable for more complex image data analysis, especially when processing images with different scales and different levels of detail, it can provide richer feature expressions.

[0132] S220 , constructing a feature similarity graph using the local features as nodes and the similarities between the local features as edge weights.

[0133] For example, construct a feature similarity graph

[0134] Each feature vector is considered as a node in the graph.

[0135] Edge weights between nodes Reflects feature similarity.

[0136] Exemplarily, the similarity may be calculated using an inner product or a Gaussian kernel function.

[0137] S230, using a spectral clustering algorithm, based on a feature similarity graph, the local features are divided into multiple clusters, representative features are screened out from each cluster, and a target feature set is obtained.

[0138] Calculate the graph Laplacian matrix L:

[0139] L=DS;(4)

[0140] Among them, S is the similarity matrix and D is the degree matrix.

[0141] Perform eigenvalue decomposition on L to obtain the eigenvectors of the matrix. Take the smallest d eigenvectors to form the characteristic matrix Where M is the total number of features.

[0142] Arrange these feature vectors in rows to form a new feature matrix V, and perform K-means clustering on the row vectors of V to obtain d feature clusters.

[0143] For each cluster, the feature closest to the centroid of the cluster is selected as the representative to obtain the final d key features.

[0144] In this step, the spectral theory of graphs is used to automatically discover the most representative and informative feature subsets based on the distribution structure of features in the sample space. This data-driven adaptive feature selection method can effectively reduce feature redundancy and improve downstream task performance.

[0145] In this embodiment, the features of the image are first extracted in a multi-scale manner, avoiding the local and global information that may be missed by single-scale principal component analysis. Secondly, based on graph clustering and spectral theory, all image features are regarded as nodes of the graph, and the nodes are connected by similarity. The structural information of the graph (by calculating the Laplace matrix) is used to analyze the relationship between these features. Through the spectral clustering method, the features are divided into different clusters, each cluster representing an important feature part in the image. The most representative features are selected from each cluster as the final key features. In this way, the most informative and representative features can be automatically selected, redundant features are avoided, and the performance of subsequent tasks is improved.

[0146] It is understandable that based on the problems of complex semantic associations between labels in multi-label image classification tasks and the difficulty of classifiers to fully utilize the co-occurrence rules of labels, this application proposes a multi-label image classifier that can achieve efficient classification and optimized generation of multiple categories of labels in target images, such as objects, scenes, emotions, etc.

[0147] Figure 3 A flow chart of an image management method provided in an embodiment of the present application is shown.

[0148] It should be understood that Figure 3 The illustrated embodiment can be viewed as an example of step S120 , which introduces the process of identifying target image labels using a multi-label image classifier.

[0149] like Figure 3 As shown, the method may include the following steps.

[0150] S310 , using a multi-view classifier, according to a target feature set, identifying label prediction results under multiple view angles corresponding to the target image.

[0151] The multi-view classification model is trained based on multi-view features and label semantic association information.

[0152] The multi-view classifier includes at least two of a global classifier, a local classifier and an attribute classifier.

[0153] It can be understood that in order to fully explore the complementary information of different perspective features, this application designs an innovative perspective fusion strategy based on the idea of ​​multi-kernel learning.

[0154] The process includes:

[0155] Train the global classifier. Use the trained global classifier to output the label prediction result of the target image from the global perspective according to the target feature set.

[0156] For example, using the key feature f k Train a Naive Bayes (NB) classifier based on the RBF kernel (Radial Basis Function Kernel) g (·), f g (·) is the global classifier.

[0157] Train the local classifier. Use the trained local classifier to output the label prediction result of the local perspective corresponding to the target image according to the target feature set.

[0158] For example, the image is divided into blocks, and SIFT (Scale-Invariant Feature Transform) features are extracted from each block, encoded with LLC (Local Linear Embedding, local linear interpolation) and then maximum pooling is performed between blocks to obtain the local feature f l Based on f l Train a polynomial kernel SVM classifier f l (·), f l (·) is a local classifier.

[0159] Train the attribute classifier. Use the trained attribute classifier to output the label prediction result from the attribute perspective corresponding to the target image according to the target feature set.

[0160] For example, a pre-trained CNN (Convolutional Neural Network) is used to extract the attribute features of the image. a Based on f a Train a linear kernel naive Bayes NB classifier f a (·), f a (·) is the attribute classifier.

[0161] S320, integrating the label prediction results from multiple perspectives to obtain a comprehensive label prediction result:

[0162] f(x)=α g f g (f k )+α l f l (f l )+α a f a (fa ); (5)

[0163] The weight α g , α l , α a Optimization is performed through cross-validation to bring out the best performance from different perspectives.

[0164] It is understandable that in multi-label classification tasks, the semantic associations between labels contain rich prior knowledge, and mining the semantic associations of labels can help the classifier make consistent predictions of labels. The following is an implementation method for mining the semantic associations of labels.

[0165] S330, build label association graph Node represents the i-th label, edge e ij The weight of ∈ε reflects the shortest path distance between labels i and j in the vocabulary.

[0166] S340, on this basis, calculate the label similarity matrix S = [s ij ] m×m :

[0167]

[0168] Among them, γ is the scale parameter, dis(v i ,v j ) is v i and v j In the figure The shortest distance in .

[0169] The label similarity matrix describes the semantic correlation between labels and can be used to guide the consistency learning of classifiers.

[0170] S350, based on the multi-view fusion classifier f(·), we further integrate label association information to construct an innovative LS-NB multi-label classification model, which is expressed as follows:

[0171]

[0172] is the weight matrix, W v The weight corresponding to the vth view.

[0173] is the bias vector, b v The bias corresponding to the vth viewing angle.

[0174] is the slack variable matrix, v Slack variables corresponding to the v-th view.

[0175] F v =[f v (x1),…,f v (x n )] T , is the feature matrix of the vth perspective.

[0176] is the label matrix of the vth view.

[0177] L=DS, is the graph Laplacian matrix, and D is the degree matrix of S.

[0178] In this optimization objective, tr(ΞLΞ T ) item uses the label correlation matrix S to smooth the output of the classifier so that related labels tend to obtain consistent prediction results.

[0179] S360: Using the label semantic association information, the comprehensive label prediction result is corrected to obtain multiple labels corresponding to the target image.

[0180] By solving the above optimization problem, the classification decision function under multiple perspectives can be obtained:

[0181]

[0182] Furthermore, the outputs of the three perspectives are integrated to obtain the final multi-label prediction result:

[0183]

[0184] Where y∈0,1 m is the predicted label vector.

[0185] S370: Classify multiple labels of the target image based on preset label categories to obtain multiple label sets.

[0186] After obtaining the output y∈0,1 of the multi-label classifier m After that, it can be further divided into multiple preset label categories, such as labels of three categories: emotion, scene, and object.

[0187] First, according to the predefined label set, the m label numbers are divided into three subsets:

[0188] Emotion label collection:

[0189] Scene tag collection:

[0190] Object label collection:

[0191] Where me +m c +m o =m.

[0192] Then, the binary indicator vectors of the three types of labels are extracted from y:

[0193]

[0194] S380, output a triple (y e ,y c ,y o ) as the prediction result of the emotion, scene, and object labels of the image.

[0195] The classification model provided in this embodiment can comprehensively consider multiple feature perspectives of the image and the semantic associations between labels, thereby generating accurate and consistent label classification results for the image.

[0196] It is understandable that the semantic association, co-occurrence rules and hierarchical relationship between tags are comprehensively considered to generate a more compact, coherent and expressive tag combination for subsequent tag analysis, recommendation and retrieval applications. This application proposes a cross-modal label graph semantic fusion (CLGSF) algorithm to achieve efficient, flexible and expressive tag fusion optimization.

[0197] The core idea of ​​the cross-modal label graph semantic fusion algorithm is: using an external knowledge base to construct a cross-modal label graph to reveal the intrinsic associations of object, scene, and emotion labels in semantic topology; modeling and learning label relationships from the perspectives of node semantics and connection topology through a graph space spectral feature extraction network; on this basis, combining the attention mechanism and conditional random field adapted to this scheme to selectively fuse multi-view labels, while considering global semantic consistency and contextual relevance, and finally obtaining an optimized label combination output.

[0198] Figure 4 A flow chart of an image management method provided in an embodiment of the present application is shown.

[0199] It should be understood that Figure 4 The illustrated embodiment can be viewed as a specific example of step S130, which introduces the process of fusing and optimizing the labels of the target image using the cross-modal label graph semantic fusion algorithm.

[0200] like Figure 4 As shown, the method may include the following steps.

[0201] S410: construct a tag graph according to multiple tag sets and a preset semantic network knowledge base.

[0202] It should be understood that this label map is a cross-modal label map.

[0203] The label graph includes multiple nodes and multiple node connection relationships, each node is used to represent a label, and each node connection relationship is used to represent the semantic relationship between two nodes.

[0204] Specifically, the label graph is constructed by treating the three types of labels, namely, objects, scenes, and emotions, as three types of nodes in the label graph, and using the semantic relationships (such as synonymy, hyponymy, etc.) provided by the preset semantic network WordNet knowledge base to construct semantic connections between nodes. Each connection is given an initial weight, which indicates the semantic relevance between the corresponding labels.

[0205] The edge set of the label graph is denoted by The node set is denoted as N is the total number of tags.

[0206] S420: When there are adjacent nodes in the label graph, the semantic representation of each node is aggregated and updated according to the semantic representations of the adjacent nodes, and the global semantic representation of each node is obtained through continuous iteration.

[0207] In the label graph, the representation of each node is updated by aggregating the semantic information of neighboring nodes to generate a global semantic representation.

[0208] Specifically, the label graph is modeled end-to-end through a graph space spectrum feature extraction network.

[0209] In the graph space spectrum feature extraction network, each label node v i A d-dimensional semantic embedding vector Representation. The graph space spectral feature extraction network updates the node representation by iteratively aggregating the semantic information of neighboring nodes. The formula is:

[0210]

[0211] in, is the node v in the lth layer i The hidden layer representation of

[0212] v i The set of neighbor nodes of

[0213] cij is the normalization constant;

[0214] σ(·) is a nonlinear activation function;

[0215] is the parameter matrix and bias term of the lth layer.

[0216] Finally, after L iterations, a global semantic representation is generated

[0217] S430. According to different regions of the target image, the global semantic representation of each node is weightedly fused using an attention mechanism to generate a regional semantic representation.

[0218] Specifically, the step may include:

[0219] The target image I is decomposed into M key regions (such as objects, background, texture, etc.).

[0220] The attention mechanism is used to weight the fusion of the global semantic representation of the label nodes:

[0221]

[0222] is the fused semantic representation of the mth region;

[0223] αmi is the attention weight of the mth region to the i-th label node, and its calculation formula is:

[0224]

[0225] is the learnable attention parameter vector.

[0226] After fusion, the regional semantic matrix of the target image is obtained

[0227] In this step, the global semantic representation of the label nodes is weighted and fused by using the attention mechanism to dynamically model the semantic features of different regions (such as objects, backgrounds, textures, etc.) in the target image, thereby generating a regional semantic representation. The fusion process adaptively captures the association between labels and regions, and optimizes the weight distribution through learnable parameters to improve feature representation capabilities and classification accuracy.

[0228] S440: Construct an energy function for globally optimizing the label combination according to the global semantic representation or the regional semantic representation of each node.

[0229] The label fusion problem is modeled as an inference problem of Conditional Random Fields (CRF) and an energy function is defined.

[0230] Let y = [y1,…,y M ] T ∈0,1 M The label indicator vector representing the image region, where y m =1 means the label corresponding to the mth region is true, y m=0 means false.

[0231] The energy function of the label is defined as follows:

[0232]

[0233] φ(s m ,y m ) is the single node potential function of region m:

[0234]

[0235] ψ(y m ,y n ,s m ,s n ) is the pairwise potential function of regions m and n:

[0236]

[0237] Where ° represents the Hadamard product. u ,b u ,w p ,b p is the parameter of CRF.

[0238] S450, solving the energy function to find the optimal label combination.

[0239] By minimizing the energy function E(y|S,Θ), the optimal label indicator vector y of the target image is inferred * :

[0240]

[0241] y * The labels corresponding to the non-zero elements in are the final selected label combinations.

[0242] The parameters of the CRF model are $θ=w u ,b u ,w p ,b p $ can be optimized in the training set through methods such as maximum likelihood estimation.

[0243] S460: Determine the optimal tag combination as the target tag.

[0244] In this embodiment, through a multi-stage process including label graph construction, semantic representation update, attention mechanism fusion, and conditional random field inference, the labels of the target image are semantically fused and globally optimized to generate a semantically consistent and accurately expressed label combination for further analysis and processing of image management and application scenarios.

[0245] In one embodiment, in response to the user opening the cloud disk, a tag recommendation algorithm is used to filter out key tags from target tags.

[0246] Among them, key tags are used to display on the front-end interface.

[0247] In this embodiment, the efficiency and convenience of users' image search in the cloud disk are improved through the tag recommendation algorithm. Its working mechanism is: when a user searches for an image in the cloud disk, the tag recommendation algorithm is used to filter out the key tags that best represent the core content of the image or are most relevant to the user's focus from the target tags. These key tags may be the result of user search history, interest preferences, or image semantic analysis. The filtered key tags will then be displayed preferentially on the front-end interface, making it convenient for users to quickly understand the main information of the image or find content of interest, thereby significantly improving the search efficiency and interactive experience.

[0248] This application proposes a personalized tag recommendation based on optimized tag set and contrastive learning (PTR-OTSCL) algorithm to achieve accurate, diverse and efficient personalized tag recommendation.

[0249] Figure 5 A flow chart of an image management method provided in an embodiment of the present application is shown.

[0250] Combine the following Figure 5 , the training method of personalized tag recommendation algorithm based on optimized tag combination and contrastive learning is introduced.

[0251] S510: construct a contrastive learning loss function based on the similarity between the user's historical interested tag set and the candidate tags.

[0252] Construct a high-dimensional portrait feature vector for each user Contains information such as age, gender, preferences, consumption level, etc. These features are extracted from users' daily behavior data and used to characterize users' personalized needs.

[0253] Using the label embedding vector output by the above CLGSF algorithm Directly to the optimized tag set For semantic modeling, there is no need to build a knowledge graph.

[0254] By user u k A collection of historical interest tags With candidate label v i The similarity between them defines the contrastive learning loss function

[0255]

[0256] is the label-user similarity, W is the learned projection matrix, and τ is the temperature parameter.

[0257] By minimizing the contrastive learning loss function, the similarity between the tags of interest and the user can be maximized, thus improving the personalization of recommendations.

[0258] S520: construct a diversity regularization term based on the candidate labels.

[0259] In order to avoid the recommended tags being too concentrated on certain topics, a diversity regularization term is introduced

[0260]

[0261] For user u k The candidate recommendation tag set.

[0262] By minimizing It can reduce the semantic similarity between recommended tags and improve topic diversity.

[0263] S530, comprehensively compare the learning loss function and the diversity regularization term to obtain user u k Personalized recommendation loss function

[0264]

[0265] S540, by minimizing the overall recommendation loss Back propagation updates the parameters and optimizes the personalized recommendation model.

[0266] The above is an introduction to the training method. It should be understood that after the personalized recommendation model training is completed, a prediction method is also included.

[0267] S550, based on the tag-user similarity s(y i ,u k ), select the Top-K tags as the user u k Personalized recommendation results.

[0268] For example, the recommendation list can be divided into themes such as objects, scenes, emotions, etc., supporting diversified display on the front end.

[0269] This embodiment significantly improves the accuracy and diversity of tag recommendations by combining optimized tag semantic embedding and contrastive learning technology. On the one hand, the contrastive learning loss function strengthens the match between user portraits and tags, ensuring that the recommended content is highly relevant to user interests; on the other hand, diversity regularization introduces topic coverage optimization between tags to avoid over-concentration of recommendation results and achieve richness and balance of content. Overall, this solution takes into account both personalized needs and diversified expressions, providing front-end applications with efficient, flexible and user-adaptive tag recommendation services.

[0270] In one embodiment, the method of the present application also proposes a natural language search optimization method, which improves the matching effect of natural language search by using the optimized tag combination generated by the cross-modal tag graph semantic fusion (CLGSF) algorithm. The specific method is as follows:

[0271] Optimize the search space.

[0272] The original search space is usually based on an unoptimized set of tags, while this method replaces the search space with a tag combination optimized by the CLGSF algorithm. The optimized tags are more semantically consistent and contextually relevant, making the search results more accurate.

[0273] Calculate and optimize query-label similarity.

[0274] The label semantic embedding vector y generated by CLGSF i , replacing the original label embedding in traditional NLP algorithms. This optimized embedding vector can more accurately represent the semantic features of the label. When processing natural language queries (Query), the semantic similarity between the Query and the label embedding vector can be calculated by i ) to improve the matching degree of the search.

[0275] Introducing dual learning.

[0276] The dual learning mechanism further improves the training effect of the model by constructing mutually supervised tasks (such as query-to-label matching tasks and label-to-query generation tasks). Dual learning can improve the matching accuracy while enhancing the semantic interpretability of the results, thereby helping users better understand the recommended search results.

[0277] Figure 6 A flow chart of an image management method provided in an embodiment of the present application is shown.

[0278] It should be understood that Figure 6 The embodiment provides a process of matching user query and tag semantics to final result display.

[0279] like Figure 6 As shown, the method may include the following steps.

[0280] S600: User natural language query.

[0281] The user enters a natural language query containing a requirement, such as "find an image of a cozy family scene."

[0282] S610: Input label text corpus.

[0283] S620: Pre-trained language model processing.

[0284] Through pre-trained language models, semantic understanding of user queries is performed, key intent and core semantic information are extracted, and natural language is converted into structured semantic expressions.

[0285] S630, Query-tag semantic matching.

[0286] According to the extracted semantic information, it is matched with the optimized tag combination in the system to filter out the tags most relevant to the query.

[0287] S640, dual learning optimization.

[0288] The dual learning mechanism is used to further enhance the semantic matching relationship between tags and queries, ensuring the accuracy and relevance of the matching.

[0289] S650: Calculate semantic similarity matrix.

[0290] Generate a semantic similarity matrix between the query and the tags to quantify the relevance of the query to each tag.

[0291] S660, tag semantic indexing.

[0292] The tag semantic index structure can be used to quickly retrieve tag combinations related to user queries, significantly improving the retrieval speed.

[0293] S670: sort the images.

[0294] The matching image results are sorted according to semantic similarity to ensure that images with high relevance are displayed to users first.

[0295] S680, front-end search result display.

[0296] The sorted image results are presented to the user, and an intuitive interface is provided so that the user can quickly browse and select images that meet their needs.

[0297] This process achieves an efficient, accurate and personalized experience from user query to image retrieval by combining pre-trained models, semantic matching and optimization mechanisms.

[0298] In this embodiment, natural language search optimization can significantly improve the semantic matching accuracy between the query and the tag, and reduce irrelevant or low-quality matching results. The optimized search results are excellent in accuracy, relevance and interpretability, and can better meet the actual needs of users.

[0299] It is understandable that when traditional cloud disk services process analytical data such as automatic image classification, they usually store the data in an online transaction processing (OLTP) system and then import it into the OLAP system for analysis through regular batch jobs. However, this method has a high data processing delay and cannot meet real-time analysis needs, resulting in a affected user experience. Especially in scenarios where users expect to see personalized classified images immediately after uploading images, the lag of traditional solutions is particularly prominent. Based on this, the present application proposes a real-time metadata management method based on a hybrid transactional and analytical processing (HTAP) architecture with elastic balanced hash indexes and dynamically tuned trees.

[0300] It should be understood that HTAP combines the functions of online transaction processing (OLTP) and online analytical processing (OLAP), and is suitable for application scenarios that require real-time data processing and analysis.

[0301] In one embodiment, in response to a user uploading a target image to the cloud disk, a preset metadata extraction algorithm is used to extract metadata of the target image;

[0302] The extracted metadata is written into the cloud disk, and the cloud disk is equipped with a metadata management engine based on an elastic balanced hash index algorithm and / or a dynamically tuned tree algorithm.

[0303] In other words, this application designs a HTAP metadata management engine that integrates two data structures, Adaptive Elastic Hash Indexing (AEHI) and Dynamic Optimization Tree (DOT), to achieve real-time data writing and millisecond-level query analysis. In addition, through the expansion of Multi-Dimensional Indexing (MDI), it supports complex queries and combined analysis, providing guarantees for high throughput and low latency.

[0304] It should be noted that, as an underlying data structure design, the HTAP metadata management engine provides basic support for HTAP operations for all complex algorithms in the above embodiments of the present application, and is also the basic guarantee for supporting the fast operation of the above complex algorithms.

[0305] The following is an introduction to the three core algorithms included in the HTAP metadata management engine.

[0306] (1) Elastic Balanced Hash Index Algorithm

[0307] Traditional hash indexes require the number of buckets to be defined in advance, which will cause space waste and hash conflicts when data is unevenly distributed. To solve this problem, an elastic balanced hash index (AEHI) structure is designed. The algorithm flow is as follows:

[0308] Initialize a hash table H of size M, each bucket B i Contains a list L of key-value pairs i and a pointer to the overflow table O i Pointer to .

[0309] For any key-value pair (k,v), perform the following insertion operation:

[0310] h←hash(k); (24)

[0311] i←h mod M; (25)

[0312] L i ←L i ∪(k,v); (26)

[0313] If the number of empty buckets in H is lower than the threshold τ, M is doubled and rehashed.

[0314] The query operation for key k is as follows:

[0315] h←hash(k); (24)

[0316] i←h mod M; (25)

[0317] In L i and O i Find the value v with key k in .

[0318] Among them, λ is the bucket overflow threshold, and τ is the empty bucket ratio threshold.

[0319] The above algorithm includes the following steps when executed:

[0320] The amount of metadata stored in each bucket in a hash table and the ratio of empty buckets are monitored, and the hash table is used to store and manage the metadata.

[0321] When the amount of metadata contained in any bucket exceeds a first preset threshold, the metadata in the bucket is moved to a preset overflow table.

[0322] When the empty bucket ratio exceeds a second preset threshold, the hash table is expanded.

[0323] From the above, we can see that AEHI can dynamically adjust the number of buckets according to the data and achieve a balance between query, insertion and space utilization.

[0324] (2) Dynamically tuned tree index algorithm

[0325] For columns that require range queries and statistical analysis, a dynamically tuned tree index is introduced. The DOT tree is a self-optimizing B+ tree variant with better cache locality and concurrency performance. The core idea is to adjust the tree node splitting threshold from 1 / 2 to α (usually 1 / 8 or 1 / 16). When the number of node elements exceeds B·α, a new node is split and propagated to the upper layer in advance to avoid the tree height increasing infinitely.

[0326] Dynamically tune the tree insertion algorithm Insert(k,v), the process is as follows:

[0327] Search for a suitable leaf node N0 from the root node downwards. If N0 is full, split N1 and change the key of N0 in the parent node to the first element of N1. Insert (k,v) into N0. Upgrade(N0).

[0328] Dynamic tuning tree node upgrade algorithm Upgrade(N), the process is as follows:

[0329] Calculate the filling degree f of N and return it if f<α.

[0330] Split N to get new node N ′ .

[0331] N ′ The first element k ′ Insert the parent node P, and recursively split it when it is full.

[0332] if P.isFull()then Upgrade(P)Insert(k ′ ,N ′ ). (27)

[0333] That is, the above algorithm may include the following steps when executed:

[0334] When the node filling degree reaches the split threshold, the node split is triggered, the redundant metadata is moved to the new node, and the first key of the new node is propagated upward to the parent node;

[0335] If the parent node also reaches the split threshold, the split is recursively performed until the adjustment is completed.

[0336] The DOT tree improves random write performance by 2-3 orders of magnitude through the mechanism of delayed splitting and early propagation. In addition, the compact sorted storage optimized by the AVX (Advanced Vector Extensions) instruction set is used inside each node to improve the throughput of retrieval and analysis.

[0337] (3) Metadata Multidimensional Indexing Algorithm

[0338] In order to support multi-dimensional combined query of metadata, a multi-dimensional index is designed based on AEHI and DOT tree. The key of MDI is to map the hash values ​​of multiple dimensions to one dimension, which is converted into the classic space filling curve problem. The Hilbert curve is used as the hash mapping function, which has better locality preservation characteristics than the Z-order curve. The algorithm flow is as follows:

[0339] For d query dimensions A1,…,A d , calculate its range [a i ,b i ],i∈[1,d].

[0340] Normalize each dimension to the interval [0,1], that is

[0341] From the d-dimensional Hilbert curve function H d :[0,1] d →[0,1] defines the following mixed hash:

[0342] H * (v1,…,v d )=H d (h1(v1),…,h d (v d )). (29)

[0343] Using AEHI to detect H * The hash value of the index.

[0344] For the range [l i ,r i ], i∈[1,d], is converted into a one-dimensional interval query of H^ value:

[0345] [L,R]=[H ( l1,…,l d ),H * (r1,…,r d )]. (30)

[0346] MDI is a high-performance and lightweight multidimensional index structure that can provide real-time multidimensional combined query capabilities on a metadata scale of billions. Together with AEHI and DOT tree, it builds the three pillars of HTAP metadata management and provides stable and efficient data services for upper-level businesses.

[0347] Figure 7 A flow chart of an image management method provided in an embodiment of the present application is shown.

[0348] It should be understood that Figure 7 The illustrated embodiment introduces the complete process of the above metadata management.

[0349] like Figure 7 As shown, the method includes the following steps.

[0350] S710: Obtain metadata.

[0351] Extract metadata from the target image to provide basic information for subsequent indexing and query operations.

[0352] S720. Build an elastic balanced hash index.

[0353] Build an AEHI index for metadata and improve the efficiency and flexibility of data access by optimizing hash distribution.

[0354] S730: multi-dimensional index processing.

[0355] During the multidimensional data indexing phase, appropriate index paths are selected based on data characteristics, allowing multidimensional data to be efficiently mapped to query structures.

[0356] S740, one-dimensional hash map.

[0357] The multidimensional query is mapped to one-dimensional space through the Hilbert curve, ensuring that the proximity of the data is maintained after the dimensionality transformation.

[0358] S750, AEHI interval search.

[0359] The AEHI index is used to quickly locate the target data range in one-dimensional space and preliminarily screen out potential data candidate tuples that meet the conditions.

[0360] S760, screening of candidate tuples.

[0361] The candidate tuples found from the AEHI index are further accurately filtered to finally generate a result set that meets the query conditions.

[0362] S770. The result is written back in real time.

[0363] The filtered query results are stored or returned for subsequent use or display.

[0364] S780. Dynamically adjust and optimize tree index.

[0365] During the query and data operation process, new data is synchronized to DOTI in real time to keep the index structure dynamically updated and perform optimally.

[0366] The entire process achieves efficient and accurate multidimensional data management and query by combining multidimensional index construction, one-dimensional mapping, fast interval search and dynamic synchronization mechanism.

[0367] This embodiment uses HTAP technologies such as in-memory computing and intelligent data hierarchical storage to achieve high-throughput real-time processing and low-latency analysis and query of image tag data. This enables the system to respond to users' tagging, retrieval, browsing and other requests in real time, and to update the association analysis and personalized recommendation results of image tags in a timely manner.

[0368] It can be understood that by using the metadata management method based on HTAP technology proposed in the above embodiment to perform image feature extraction, cross-modal label graph fusion, personalized label recommendation and other processing on the target image, the latency speed of image annotation, semantic richness, retrieval accuracy, recommendation personalization and gallery browsing interest can be greatly improved, thereby improving the image management quality and retrieval quality.

[0369] Figure 8 A schematic diagram of the architecture of an image management method provided in an embodiment of the present application is shown.

[0370] from Figure 8 It can be seen that the overall architecture of the present application solution can include five main modules: image upload and metadata extraction 81, HTAP database storage module 82, image feature extraction and classification module 83, multi-label generation and fusion module 84 and label analysis and application module 85.

[0371] Each module is described below.

[0372] Image upload and metadata extraction module 81, which is used to upload images and extract metadata. The steps performed include:

[0373] S810, uploading images. Exemplarily, two modes are supported, namely, a real-time uploading interface and a batch uploading interface, to meet the needs of different user scenarios.

[0374] S811. After the image is uploaded, the system calls a preset algorithm to extract metadata from the image, providing a basis for subsequent analysis and storage.

[0375] HTAP database storage module 82, which uses HTAP technology as its core to achieve efficient data storage and real-time analysis. The steps it performs include:

[0376] S820: Real-time data writing, that is, the uploaded metadata is written into the HTAP database in real time to ensure the real-time processing.

[0377] S821. Column storage, that is, using column storage technology to optimize data access performance and adapt to high-frequency query and analysis needs.

[0378] S822, memory computing engine, which combines memory computing technology to accelerate the computing process and provide support for real-time data analysis.

[0379] S823, i.e. real-time data analysis: through the built-in analysis capabilities of the HTAP database, real-time analysis of the stored data is performed.

[0380] Image feature extraction and classification module 83, which is responsible for in-depth analysis and classification of image content. The steps it performs include:

[0381] S830, Adaptive Principal Component Analysis (APCA), that is, using APCA technology to extract image features, simplify data dimensions and retain main information.

[0382] S831, Least Squares Support Vector Machine (LSSVM), that is, classifying image features based on the LSSVM algorithm and generating preliminary category labels for each image.

[0383] S832. Incremental learning and updating, that is, when the data increases or the features change, the classification model is dynamically optimized through incremental learning to improve the classification performance and adaptability.

[0384] The multi-label generation and fusion module 84 generates multiple types of labels according to image features and performs optimized fusion. The steps performed include:

[0385] S840, generating emotion labels, that is, identifying the emotion information (such as happiness, sadness, etc.) conveyed by the image and generating corresponding emotion labels.

[0386] S841, scene label generation, that is, extracting scene information (such as city, beach, etc.) in the image and generating scene labels.

[0387] S842, object label generation, that is, identifying specific objects (such as animals, vehicles, etc.) in the image through target detection and generating object labels.

[0388] S843, label fusion and optimization, that is, semantically fusing the above labels, removing redundant information and optimizing the accuracy and adaptability of the labels.

[0389] The tag analysis and application module 85 implements in-depth analysis of the generated tags and applies the results to multi-scenario services. The steps performed may include:

[0390] S850, multi-dimensional label analysis, that is, performing multi-dimensional statistics and analysis on the fused labels to extract valuable semantic information.

[0391] S851. Personalized tag recommendation, that is, recommending tags that are highly relevant to user needs based on user preferences and historical behaviors.

[0392] S852, tag search and browsing, that is, supporting users to quickly search and intuitively browse through tags, improving the convenience of gallery management and retrieval.

[0393] This image management method uses HTAP technology and multiple optimization algorithms to effectively improve the efficiency and intelligence of image management and retrieval, significantly reduce processing delays, improve the richness of tag semantics and the accuracy of retrieval recommendations, and enable users to obtain a more efficient, personalized and interesting image management and retrieval experience.

[0394] To illustrate the user experience process, the following takes a user uploading a cute kitten image as an example to show how intelligent algorithms provide real-time, personalized image management and search services, including intelligent tag recognition process, tag optimization and fusion process, personalized tag recommendation process, natural language search process, diversified gallery browsing process and interactive image description process.

[0395] The intelligent label recognition process includes: when a user uploads a kitten image, the intelligent classification algorithm will analyze it in depth and generate multi-dimensional labels. Object labels may include "cat", "kitten", "pet", etc., scene labels may be "indoor", "sofa", "carpet", and emotional labels may identify "cute", "warm", "lovable", etc. The algorithm accurately extracts the most attractive core elements in the image, providing a reliable foundation for subsequent services.

[0396] The label optimization and fusion process includes: Based on the CLGSF cross-modal label graph semantic fusion algorithm, the initially generated scattered labels will be further optimized and integrated. By constructing a cross-modal label graph, the semantic relationship between objects, scenes and emotional labels is deeply explored, and the label semantics is expanded in combination with an external knowledge base, and finally a concise and highly semantically consistent label combination is formed, such as "cute kitten, warm interior, lovable". The optimized labels are not only more expressive, but also convenient for subsequent retrieval and display.

[0397] The personalized tag recommendation process includes: when users browse the image library, the PTR-OTSCL algorithm will dynamically recommend highly relevant and personalized tags based on the user's browsing history and interest preferences. If the user often browses content such as "cats" or "cute pets", the algorithm may highlight and recommend tags such as "cute cats" and "cuddly". The recommendation results are optimized through comparative learning to ensure maximum appeal to user interests. In addition, the diversity regularization mechanism will provide recommendation results with a wide range of topics, covering multiple dimensions such as objects, scenes, and emotions, enhancing users' interest in exploring image content.

[0398] The natural language search process includes: users can directly enter natural language queries, such as "find some cute kitten images". The algorithm first uses the pre-trained language model to parse the user query, extract keywords such as "cute" and "kitten", and quickly locate relevant tag combinations such as "cute kitten, warm interior" by matching the query semantics with the optimized tags. At the same time, the dual learning mechanism is used to enhance the accuracy of semantic understanding. Finally, images matching highly relevant tags will be quickly retrieved and presented in order of relevance.

[0399] Diversified gallery browsing process, including the gallery browsing interface will automatically generate rich image grouping and theme tags based on the semantic information of the tags. For example, the "cute kitten" image can be classified into the "cute pet moment" theme. Users can easily browse related images by selecting the theme tags of interest. Personalized recommendation tags will be dynamically adjusted according to the browsing context, guiding users to discover more theme content, such as "warm" and "loving", to stimulate their interest in exploration.

[0400] The interactive image description process includes: if the user particularly likes a certain kitten image, the algorithm can also automatically generate a vivid image description based on the image content and associated tags, such as "a super cute little kitten, lying on a warm sofa, sleeping soundly, looking adorable." These descriptions can enhance the user's sense of substitution and bring deeper emotional resonance. Users can also like, comment or edit the description to further enhance the human-computer interaction experience.

[0401] Through the above process, fast, efficient and personalized image management and search are achieved, bringing rich experience and higher usage value to users.

[0402] In one embodiment, when the method of the present application is applied to a terminal device, the following steps may also be included:

[0403] Receive the operation of uploading the target image to the cloud disk.

[0404] It is understandable that the operation of uploading the target image to the cloud disk can be triggered in a variety of ways.

[0405] In one implementation, the uploading of the target image may be actively performed by the user, for example, the user selects and manually uploads an image file that needs to be classified or labeled.

[0406] In one implementation, the upload of the target image may also be automatically triggered by the system. For example, in some application scenarios, the terminal device may automatically detect a newly generated image or video frame according to preset conditions and upload it to the cloud disk. Such automatic upload operations can occur in scenarios such as image generation, video recording, and real-time monitoring, without user intervention, and automatically perform image upload according to specific rules or events.

[0407] According to the operation, the target image is sent to the cloud disk server so that the cloud disk server executes the method described in any of the above embodiments.

[0408] The above mainly introduces an image management method of an embodiment of the present application in conjunction with the accompanying drawings. At the same time, it should be understood that although the steps in the flowcharts involved in the embodiments described above are displayed in sequence, these steps are not necessarily executed in sequence in the order shown in the figure. Unless otherwise specified in this article, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the embodiments described above may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps. An image management device of an embodiment of the present application is introduced below in conjunction with the accompanying drawings. For the sake of brevity, appropriate omissions will be made when introducing the device below, and the relevant content can refer to the relevant description in the method above, and will not be repeated.

[0409] Corresponding to the method described in the above embodiment, Fig. 9 A structural block diagram of an image management device 1000 provided in an embodiment of the present application is shown. For the convenience of description, only the part related to the embodiment of the present application is shown.

[0410] like Fig. 9 As shown, the device 1000 may include:

[0411] The feature extraction module 1001 is used to respond to the operation of uploading the target image to the cloud disk, perform feature dimensionality reduction extraction on the target image using a preset feature dimensionality reduction algorithm, and obtain a target feature set.

[0412] The label recognition module 1002 is used to use a preset label recognition algorithm to perform label recognition on the target image according to the target feature set to obtain multiple labels corresponding to the target image.

[0413] The tag fusion module 1003 is used to perform semantic fusion on multiple tags using a preset tag fusion algorithm to obtain a target tag.

[0414] The management module 1004 is used to associate and store the target tag with the target image in the cloud disk.

[0415] Fig.10 A structural block diagram of an image management device 2000 provided in an embodiment of the present application is shown. For the convenience of explanation, only the part related to the embodiment of the present application is shown.

[0416] like Fig.10 As shown, the device 2000 may include:

[0417] The receiving module 2001 is used to receive an operation of uploading a target image to a cloud disk.

[0418] The sending module 2002 is used to send the target image to the cloud disk server according to the operation, so that the cloud disk server executes the method described in any of the above embodiments.

[0419] Fig.11 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present application is shown.

[0420] The computer device may include a processor 7001 and a memory 7002 storing computer program instructions.

[0421] Specifically, the processor 7001 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0422] The memory 7002 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 7002 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In one example, the memory 7002 may include a removable or non-removable (or fixed) medium, or the memory 7002 is a non-volatile solid-state memory. The memory 7002 may be inside or outside the integrated gateway disaster recovery device.

[0423] In one example, the memory 7002 may be a read-only memory (ROM). In one example, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or a flash memory, or a combination of two or more of these.

[0424] The memory 7002 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.

[0425] The processor 7001 reads and executes the computer program instructions stored in the memory 7002 to implement Figure 1 The image management method in the illustrated embodiment.

[0426] In one example, the computer device may further include a communication interface 7003 and a bus 7004. Fig.11 As shown, the processor 7001, the memory 7002, and the communication interface 7003 are connected via a bus 7004 and communicate with each other.

[0427] The communication interface 7003 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0428] Bus 7004 includes hardware, software or both, and the components of online data flow billing equipment are coupled to each other. For example, but not limitation, the bus may include Accelerated Graphics Port (AGP) or other graphics bus, Enhanced Industry Standard Architecture (EISA) bus, Front Side Bus (FSB), Hyper Transport (HT) interconnection, Industry Standard Architecture (ISA) bus, InfiniBand interconnection, Low Pin Count (LPC) bus, Memory bus, Micro Channel Architecture (MCA) bus, Peripheral Component Interconnect (PCI) bus, PCI-Express (PCI-X) bus, Serial Advanced Technology Attachment (SATA) bus, Video Electronics Standards Association Local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 7004 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the present application considers any suitable bus or interconnection.

[0429] In addition, in combination with the image management method in the above embodiment, the embodiment of the present application can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any one of the image management methods in the above embodiment is implemented.

[0430] An embodiment of the present application also provides a computer program product, including a computer program, and when the computer program is processed and executed, any one of the image management methods in the above embodiments is implemented.

[0431] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.

[0432] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), appropriate firmware, plug-in, function card, etc. When implemented in software, the elements of the present application are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or communication link by a data signal carried in a carrier. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (Read-Only Memory, ROM), flash memory, erasable read-only memory (Erasable Read Only Memory, EROM), floppy disks, compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM), optical discs, hard disks, optical fiber media, radio frequency (Radio Frequency, RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0433] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0434] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.

Claims

1. An image management method, characterized in that: include: In response to an operation of uploading a target image to a cloud disk, performing feature dimensionality reduction extraction on the target image using a preset feature dimensionality reduction algorithm to obtain a target feature set; Using a preset label recognition algorithm, according to the target feature set, label recognition is performed on the target image to obtain a plurality of labels corresponding to the target image; Using a preset tag fusion algorithm to semantically fuse the multiple tags to obtain a target tag; In the cloud disk, the target tag is associated with the target image and stored.

2. The method according to claim 1, characterized in that The step of performing feature dimensionality reduction extraction on the target image using a preset feature dimensionality reduction algorithm to obtain a target feature set includes: Performing principal component analysis on the target image to extract local features of the target image; Constructing a feature similarity graph using the local features as nodes and the similarities between the local features as edge weights; Based on the feature similarity graph, the local features are divided into a plurality of clusters, and representative features are screened out from each cluster to obtain the target feature set.

3. The method according to claim 2, characterized in that The performing principal component analysis on the target image to extract local features of the target image includes: Performing gradient recursive sampling on the target image to obtain a plurality of sub-images with different resolutions; Split each sub-image into multiple non-overlapping local blocks; Perform principal component analysis on each local block to extract the principal component features of each local block; The principal component features of each local block are cascaded to obtain the local features of the target image.

4. The method according to claim 1, characterized in that: The method of using a preset label recognition algorithm to perform label recognition on the target image according to the target feature set to obtain a plurality of labels corresponding to the target image includes: Using a multi-view classifier, based on the target feature set, identifying label prediction results under multiple view angles corresponding to the target image; Fusion of the label prediction results from the multiple perspectives to obtain a comprehensive label prediction result; The comprehensive label prediction result is corrected by using label semantic association information to obtain multiple labels corresponding to the target image.

5. The method according to claim 4, characterized in that The multi-view classifier includes at least two of a global classifier, a local classifier, and an attribute classifier; The global classifier is used to output a label prediction result corresponding to the target image under a global perspective according to the target feature set; The local classifier is used to output a label prediction result under a local perspective corresponding to the target image according to the target feature set; The attribute classifier is used to output a label prediction result from an attribute perspective corresponding to the target image according to the target feature set.

6. The method according to claim 1, characterized in that The using a preset tag fusion algorithm to semantically fuse the multiple tags to obtain a target tag includes: Classifying a plurality of labels corresponding to the target image to obtain a plurality of label sets, each of the label sets including at least one label; According to the multiple tag sets and the preset semantic network knowledge base, a tag graph is constructed, wherein the tag graph includes multiple nodes and multiple node connection relationships, each node is used to represent a tag, and each node connection relationship is used to represent a semantic relationship between two nodes; When there are adjacent nodes in the label graph, the semantic representation of each node is aggregated and updated according to the semantic representations of the adjacent nodes, and the global semantic representation of each node is obtained through continuous iteration; According to the global semantic representation of each node, an energy function for global optimization of label combinations is constructed; Solving the energy function to determine the optimal label combination; The optimal tag combination is determined as the target tag.

7. The method according to claim 6, characterized in that The step of constructing an energy function for globally optimizing the label combination according to the global semantic representation of each node includes: According to different regions of the target image, weighted fusion processing is performed on the global semantic representation of each node to generate a regional semantic representation; According to the region semantic representation, an energy function for globally optimizing the label combination is constructed.

8. The method according to any one of claims 1 to 5, characterized in that The method further comprises: In response to the user's operation of opening the cloud disk, a tag recommendation algorithm is used to filter out key tags from the target tags, and the key tags are used to be displayed on the front-end interface.

9. The method according to claim 8, characterized in that The method further comprises: Construct a contrastive learning loss function based on the similarity between the user's historical interested tag set and the candidate tags; Constructing a diversity regularization term according to the candidate labels; Constructing a tag recommendation model according to the contrastive learning loss function and the diversity regularization term; By minimizing the value of the tag recommendation model, the parameters of the tag recommendation model are updated.

10. The method according to any one of claims 1 to 5, characterized in that Before extracting features from the target image using a preset feature dimensionality reduction algorithm to obtain a target feature set, the method further includes: In response to a user uploading a target image to the cloud disk, extracting metadata of the target image using a preset metadata extraction algorithm; The extracted metadata is written into the cloud disk, and the cloud disk is equipped with a metadata management engine based on an elastic balanced hash index algorithm and / or a dynamically tuned tree algorithm.

11. The method according to claim 10, characterized in that In a case where the metadata management engine is optimized based on the elastic balanced hash index algorithm, writing the extracted metadata into the cloud disk includes: Monitoring the amount of metadata stored in each bucket in a hash table and the ratio of empty buckets, wherein the hash table is used to store and manage the metadata; When the amount of metadata contained in any bucket exceeds a first preset threshold, the metadata in the bucket is moved to a preset overflow table; When the empty bucket ratio exceeds a second preset threshold, the hash table is expanded.

12. The method according to claim 10, characterized in that In the case where the metadata management engine performs optimization based on the dynamic tuning tree algorithm, writing the extracted metadata into the cloud disk includes: When the node filling degree reaches the split threshold, the node split is triggered, the redundant metadata is moved to the new node, and the first key of the new node is propagated upward to the parent node; When the parent node also reaches the split threshold, the split is performed recursively until the adjustment is completed.

13. An image management device, characterized in that: include: A feature extraction module, configured to, in response to an operation of uploading a target image to a cloud disk, perform feature dimensionality reduction extraction on the target image using a preset feature dimensionality reduction algorithm to obtain a target feature set; A label recognition module, used to perform label recognition on the target image according to the target feature set using a preset label recognition algorithm to obtain a plurality of labels corresponding to the target image; A label fusion module, used to perform semantic fusion on the multiple labels using a preset label fusion algorithm to obtain a target label; A management module is used to store the target tag in association with the target image in the cloud disk.

14. An image management method, characterized in that: Applied to terminal equipment, including: Receive an operation to upload the target image to the cloud disk; According to the operation, the target image is sent to the cloud disk server so that the cloud disk server executes the method as described in any one of claims 1 to 12.

Citation Information

Cited By

  • Video second-level label labeling method based on deep part multi-label learning

    CN121236665A