Digital archive classification query method based on large model index identification
Through the index recognition method based on the big model, the problem of insufficient accuracy and speed in the existing digital archive query methods is solved, efficient and accurate multimodal data query is realized, and dynamic optimization and expansion are supported.
Patent Information
- Application Number
- CN202510173119.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-17
AI Technical Summary
The existing digital archive query methods have problems with the need to improve accuracy and speed, especially in the query scenarios of non-text modal data. The query effect is poor, and it is difficult to dynamically adjust the index content and optimize the query strategy with static index and rule matching, resulting in slow query speed and poor system scalability.
The index recognition method based on large models is adopted to generate a unified feature representation through the acquisition and processing of multimodal data. The information expression and classification results are optimized by using feature fusion and dynamic verification mechanisms, and the dimensionality reduction embeds the features into the low-dimensional semantic space, and a vector index supporting query and feedback optimization is built.
It improves the accuracy and speed of digital archive query, enhances a deep understanding of user intentions, supports dynamic adjustment of index content and optimizes query strategies, and improves the scalability and response speed of the system.
Smart Images

Figure CN120104853A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a digital archive classification query method based on large model index recognition. Background Art
[0002] With the advent of the information age, a large number of paper documents have been gradually digitized, forming a huge digital archive. These archives cover data of various types and formats, such as text, images, audio, and structured numerical values. Faced with a huge and diverse archive, how to efficiently and accurately manage and query these archives has become an important research topic in the field of archive management.
[0003] As the trend of localization becomes increasingly prominent, various industries are in urgent need of digital management tools with independent intellectual property rights. In terms of digital archive management, domestic related technologies are also constantly being explored and developed.
[0004] However, most traditional digital archive query methods currently rely on keyword matching or pre-set classification tags. Although such methods have certain practicality in the early days, they also have obvious limitations: keyword matching cannot understand semantic relationships, resulting in low query result accuracy; classification tags rely on manual settings, lack flexibility, and are difficult to adapt to dynamically changing user needs and archive data growth.
[0005] The traditional query technology of the existing technology lacks a deep understanding of user intentions, especially in the query scenario of non-text modal data, and the query effect is poor. At the same time, these methods usually use static indexing and rule matching, which makes it difficult to dynamically adjust the index content and optimize the query strategy. In the scenario of large-scale archives, they show problems such as slow query speed and poor system scalability. Summary of the invention
[0006] In view of the shortcomings of the prior art, the present invention provides a digital archive classification query method based on large model index recognition, which solves the problem that the accuracy and speed of digital archive query methods need to be improved.
[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: A digital archive classification query method based on large model index recognition includes the following steps: Generate a unified feature representation through the collection and processing of multimodal data; Optimize the information expression of multimodal data based on feature fusion method; Use dynamic verification mechanisms to improve the consistency of data features and the accuracy of classification results; The optimized features are embedded into the low-dimensional semantic space through the dimensionality reduction method to generate classification labels; Build a vector index that supports query and feedback optimization.
[0008] Preferably, the collection of multimodal data comprises the following steps: Collect text data, and perform word segmentation and embedding extraction on the text data through a pre-trained language model to generate a semantic feature vector; collect image data, perform feature extraction on the image data through a convolutional neural network to generate a high-dimensional visual feature vector; Collect audio data, transcribe it into text data through a speech-to-text model, and then extract semantic feature vectors; Collect structured numerical data and adjust the data range through normalization methods to generate standardized numerical features.
[0009] Preferably, the processing of the multimodal data comprises the following steps: Use language models to extract semantic embeddings from text data; Use convolutional neural networks to extract feature embedding from image data; Use speech recognition models to transcribe audio data and extract semantic features; Normalize structured numeric data.
[0010] Preferably, the feature fusion method optimization comprises the following steps: The high-dimensional features of multimodal data are input into the variational autoencoder model, and the multimodal features are compressed by the encoder to generate latent feature representations; In the latent feature space, the optimization objective is defined as: Where X represents the multimodal features of the input, Z represents the latent feature representation, Y represents the classification target, I(X; Z) represents the mutual information between the multimodal features and the latent features, I(Z; Y) represents the mutual information between the latent features and the classification target, and β is a hyperparameter that balances compression and classification accuracy. Optimize the latent feature representation to maximize the relevance of the latent feature to the classification task, while removing redundant features and reducing the duplication and noise interference between multimodal features; Generate compressed and optimized fusion feature representation for subsequent classification processing.
[0011] Preferably, the step of utilizing the dynamic verification mechanism includes the following: Modal consistency verification: By calculating different modal characteristics Z i and Z j The prediction difference on the target category Y optimizes the consistency between modalities and defines the loss function as: Among them, Z i and Zj represents the feature representation of modalities i and j, f(Z) represents the classification function, ||·|| 2 represents the Euclidean distance; Modal complementarity verification: Evaluate each modal feature Z i The independent contribution to the target classification Y is defined as the complementary loss function: Among them, I(Z i ; Y) represents the modal characteristic Z i The mutual information with the target classification Y, λ is the trade-off factor; Dynamic weight adjustment: According to each mode characteristic Z i The update modal weight related to the target classification Y is updated as follows: in, represents the weight of mode i in the tth iteration, is the weight for the next iteration; Comprehensive optimization objective function: Combining consistency verification, complementarity verification and dynamic weight adjustment, the final optimization goal is: in, represents the comprehensive optimization objective function of the dynamic verification mechanism, represents the loss function of modal consistency verification, which is used to minimize the difference between different modal features in the classification target prediction results. Represents the loss function for modal complementarity verification, which is used to evaluate the independent contribution of each modal feature.
[0012] Preferably, the dimensionality reduction method comprises the following steps: Construct the neighborhood relationship of high-dimensional feature data points and determine the neighborhood of each data point by calculating the similarity between data points: On the basis of maintaining the neighborhood relationship, high-dimensional features are mapped to low-dimensional space to generate low-dimensional embedded features; Optimize the dimensionality reduction process to ensure that the low-dimensional embedded features can preserve the global structure and semantic consistency of the high-dimensional data.
[0013] Preferably, the generation of the classification label comprises the following steps: Input low-dimensional embedded features into the large language model, combine low-dimensional feature representation with contextual semantic information to generate initial classification labels; refine the initial classification labels according to the time dimension, subject dimension and priority dimension of the archival data; Dynamically adjust the classification label content and generate the final classification label based on user query requirements and archive background information.
[0014] Preferably, the construction supports query and feedback optimization and includes the following steps: Collect user feedback on classification query results and input the feedback information into the model as new samples; The dynamic classification label generation module is optimized through feedback samples, and the classification labels are adjusted to more accurately reflect user needs. The feature fusion module is updated using feedback samples to improve the accuracy of multimodal data feature representation and classification effect. The low-dimensional vector index library is dynamically adjusted so that new samples and optimized features can be updated to the query system in real time, ensuring the dynamic and accurate nature of the query results.
[0015] Preferably, the construction of the vector index comprises the following steps: Store low-dimensional embedded features into a vector database and index features based on vector similarity; Use the approximate nearest neighbor search algorithm to optimize the index to ensure low latency and high efficiency of the query process; In response to the update of dynamic classification labels, the feature data in the vector index is adjusted in real time to support multi-dimensional query and matching of classification labels.
[0016] The present invention also provides a digital archive classification query system based on large model index recognition, comprising: Data collection module, used to collect multimodal data such as text, images, audio and structured numerical values; Data preprocessing module, used to standardize the collected multimodal data and generate a unified feature representation; Feature fusion module, used to optimize the information expression of multimodal data and eliminate redundant information; Verification and optimization module, used to verify the consistency of data features and optimize classification results; Dimensionality reduction classification module, used to embed feature data into low-dimensional semantic space and generate dynamic classification labels; Index query module, used to build a vector index library that supports fast query; The user feedback module is used to collect user feedback information and optimize the classification query model.
[0017] The present invention provides a digital archive classification query method based on large model index recognition. It has the following beneficial effects: 1. The present invention realizes unified feature representation of text, image, audio and structured numerical data through the collection and processing of multimodal data, effectively solving the problems of single data source and inconsistent feature expression in the prior art.
[0018] 2. The present invention uses a variational autoencoder to compress and fuse multimodal features, eliminates redundant information, and retains the core features related to the classification task, thereby improving the feature expression capability and solving the problem of poor classification performance caused by information redundancy in the prior art.
[0019] 3. The present invention ensures the collaborative performance of multimodal features in classification tasks through modal consistency verification and complementarity verification, and dynamically adjusts the modal weights, thereby solving the problem of insufficient feature synergy in the prior art and significantly improving the accuracy and robustness of the classification results.
[0020] 4. The present invention maps high-dimensional features to low-dimensional semantic space based on manifold learning technology, thereby reducing computational complexity while retaining global semantic consistency, thus solving the problem of high computational cost of existing high-dimensional features.
[0021] 5. The present invention supports dynamic updating of classification models and index content by introducing an approximate nearest neighbor search algorithm and a user feedback optimization mechanism, thereby achieving rapid response to classification queries and solving the problem that existing static indexes cannot adapt to dynamic scenarios.
[0022] 6. The present invention dynamically integrates user feedback into the index library and classification model optimization to form a closed-loop optimization mechanism, thereby improving the system's adaptability to user needs and solving the problem of insufficient response of existing classification systems in scenarios with dynamically changing needs.
[0023] 7. Through full-link optimization, the present invention provides efficient and accurate digital archive classification query services, solves the deficiency that existing rule-based classification methods are difficult to adapt to complex multimodal data, and provides strong technical support for digital archive management and intelligent applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flow chart of the method of the present invention; Figure 2 It is a system architecture diagram of the present invention. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] Please see attached Figure 1 The embodiment of the present invention provides a digital archive classification query method based on large model index recognition, comprising the following steps: S1. Generate a unified feature representation through the collection and processing of multimodal data; S2, optimize the information expression of multimodal data based on feature fusion method; S3. Use dynamic verification mechanism to improve the consistency of data features and the accuracy of classification results; S4, embed the optimized features into low-dimensional semantic space through dimensionality reduction method to generate classification labels; S5. Build a vector index that supports query and feedback optimization.
[0027] For step S1, this embodiment provides a method for collecting and processing multimodal data to generate a unified feature representation. This step is the basic step of the method of the present invention, which aims to solve the inconsistency problem of multimodal data in structure, scale and semantic representation. Through the effective collection and standardized processing of multimodal data, heterogeneous data is uniformly converted into high-dimensional feature representation, laying the foundation for subsequent feature fusion, dynamic verification and index query.
[0028] As an option, data collection includes data input in four modes: text, image, audio, and structured numerical values. It should be noted that these data can be obtained through sensor devices, database interfaces, or file reading, and feature vectors are generated through a unified processing flow to form standardized inputs.
[0029] The collection and processing of multimodal data includes the following: The collection and processing of text data, which may come from documents, emails, webpage content or text fields in a database. In one implementation, the text data is processed using a pre-trained language model.
[0030] Specifically: First, the text data is segmented to remove stop words and meaningless characters.
[0031] Subsequently, a pre-trained language model is used to generate high-dimensional semantic embeddings, and the embedded representation retains the semantic features of the text.
[0032] For example, assuming that the input text is "archives contain classified information in 2023", the word segmentation results are "archives", "contains", "2023", "classification", "information", and the embedding generated by the pre-trained language model is represented as a 512-dimensional feature vector. The specific embedding form is: E t =Encoder(T) Among them, E t Represents text feature embedding, T represents text data, and Encoder is a pre-trained language model.
[0033] The acquisition and processing of image data includes the following: Image data comes from document scans, photos or key frames of video frames. Specifically, deep convolutional neural networks are used to extract features from image data. Text feature embedding is a high-dimensional dense vector that can capture text semantic information and is compatible with other modality features.
[0034] Alternatively, the image data is first preprocessed, including operations such as resizing, grayscaling, or denoising. Then, a convolutional neural network is used to generate a high-dimensional visual feature vector.
[0035] The formula is: E i =CNN(I) Among them, E i represents image feature embedding, I represents image data, and CNN is a convolutional neural network.
[0036] It should be noted that, in some embodiments, the dimension of image feature extraction can be adjusted according to the specific application scenario. For example, the output layer of ResNet-50 can be truncated to generate a 2048-dimensional high-dimensional feature vector.
[0037] The collection and processing of audio data includes the following: The audio data may come from a voice recording or a multimedia file. In one implementation, the audio data is first transcribed into text using a speech-to-text model, and then a text processing flow is used to generate semantic embeddings.
[0038] Specifically: The audio data is fed into a speech-to-text model to generate a transcribed text.
[0039] The transcribed text is processed by a pre-trained language model to generate high-dimensional semantic features.
[0040] The formula is: E a =Encoder(STT(A)) Among them, E a represents the high-dimensional feature embedding of audio, A represents audio data, and STT represents the speech-to-text model.
[0041] It should be noted that, in some embodiments, audio data can also directly generate feature vectors through an audio feature extraction model. This method can be used without the need for a semantic level.
[0042] The collection and processing of structured numerical data includes the following: The structured numerical data usually comes from sensors, table data or database fields. In this embodiment, the structured data is first normalized to ensure that its value range is within 0, 1 or other standard ranges.
[0043] Specifically: Use the Min-Max normalization method to normalize the data: Among them, E n represents the normalized feature representation, N represents the original structured numerical data, min(N) and max(N) are the minimum and maximum values of the data respectively.
[0044] It is understood that the normalized data can avoid the impact of numerical scale differences on subsequent processing and be unified and compatible with other modal features.
[0045] The unified representation of multimodal features includes the following: After all modal data are processed independently, their feature embeddings are concatenated or fused to form a unified high-dimensional feature representation.
[0046] Specifically: Let the text feature be E t , the image feature is E i , audio feature is E a , the numerical feature is E n , the unified feature is expressed as: X=[E t ,E i ,E a ,E n ] It is understood that this high-dimensional feature representation contains the key information of multimodal data and can provide support for subsequent feature fusion and classification tasks.
[0047] Extended content and technical adaptation include the following: It should be noted that, in some embodiments, other modal data (such as video data, sensor data, etc.) can also be used as input to generate feature representations through a similar preprocessing process.
[0048] Specifically: For video data, image features can be extracted from key frames, or converted into time series features through time series modeling.
[0049] For sensor data, signal features can be extracted directly, or frequency domain features can be extracted through Fourier transform.
[0050] In summary, this step describes in detail the collection and processing methods of multimodal data, from the extraction of features of text, images, audio and structured numerical data to the generation of unified high-dimensional feature representation, ensuring the technical foundation and data compatibility of subsequent steps.
[0051] For step S2, a technical solution for optimizing the information expression of multimodal data based on feature fusion method is used. Based on the collection and processing of multimodal data, this step compresses and fuses the unified high-dimensional feature representation of text, image, audio and structured numerical data, removes redundant information, and improves the effectiveness and expression ability of data features. Through feature fusion, the most useful feature representation for classification tasks is further extracted, providing an accurate and efficient data foundation for subsequent dynamic verification mechanisms and classification tasks.
[0052] It should be noted that this step uses a variational autoencoder to compress and optimize features, maximally retaining the useful information between multimodal data features, while eliminating irrelevant information and noise, to achieve efficient fusion expression of multimodal data.
[0053] Feature fusion methods optimize the information expression of multimodal data, including the following: In one implementation, the high-dimensional feature X of the multimodal data obtained by step S1 is input into the variational autoencoder to achieve feature compression and fusion.
[0054] Specifically, the variational autoencoder consists of two parts: the encoder and the decoder: The encoder is used to map high-dimensional input features to the latent feature space X and generate latent feature representation Z.
[0055] The decoder is used to reconstruct the latent features Z back to the original features Used to verify the quality of feature compression.
[0056] The formula is as follows: Among them, X represents the multimodal features of the input, Z represents the potential feature representation, represents the reconstructed feature representation, Enc and Dec represent the encoder and decoder of the variational autoencoder, respectively.
[0057] It should be noted that the latent feature representation Z extracts the most valuable information from multimodal data through an adaptive compression mechanism while removing redundant and noisy information.
[0058] In the latent feature space, the optimization goal is to maximize the correlation between the latent feature Z and the classification target Y, while minimizing the redundant information between Z and the input data X, thereby achieving efficient feature fusion.
[0059] In one possible implementation, the optimization objective is defined as follows: in, represents the optimization objective function of feature fusion, I(X; Z) represents the mutual information between the input feature X and the latent feature Z, I(Z; Y) represents the mutual information between the latent feature Z and the classification target Y, which is used to retain the information in the latent feature that is useful for the classification task, and β is the balance coefficient, which is used to control the trade-off between compression and classification accuracy.
[0060] Specifically, by minimizing I(X; Z) and maximizing I(Z; Y), effective information compression and optimization of classification objectives can be achieved in the latent feature space.
[0061] In some embodiments, to further improve the effect of feature fusion, the latent feature representation Z can be optimized through regularization constraints to ensure that the distribution of features has better representation ability and generalization.
[0062] Exemplarily, the regularization constraint of the latent feature can be expressed as follows: Among them, D KL represents the Kullback-Leibler divergence, q(Z|X) represents the potential feature distribution generated by the encoder, and p(Z) represents the prior distribution of the potential features, which is usually taken as the Gaussian distribution N(0,I).
[0063] It can be understood that regularization constraints make the latent feature representation Z obey a predefined distribution, which helps to improve the stability and generalization performance of the model.
[0064] It should be noted that, in some embodiments, feature fusion can also be achieved by weighted concatenation, in which the features of each modality are weighted and superimposed according to a certain weight ratio to generate a final fused feature representation.
[0065] The formula is as follows: Among them, Z fuse represents the fused feature representation, E i represents the feature representation of the i-th mode, α i represents the weight coefficient of mode i, and n represents the number of modes.
[0066] It can be understood that through weighted fusion, the features of different modalities can be weighted and adjusted according to their importance, further improving the effect of feature fusion.
[0067] In summary, this step compresses and fuses the high-dimensional features of multimodal data through a variational autoencoder, and the optimization goal is to retain information related to the classification task and eliminate irrelevant information and noise. At the same time, through mutual information optimization and regularization constraints, the effectiveness and stability of feature representation are further improved. In some embodiments, feature fusion can also be achieved by weighted splicing to ensure that the feature representation has stronger expressiveness and generalization. This step provides high-quality input data for subsequent dynamic verification mechanisms and classification tasks.
[0068] For step S3, this embodiment proposes an optimization method based on a dynamic verification mechanism to improve the consistency of multimodal features and the accuracy of classification results. The mechanism works together through modal consistency verification, modal complementarity verification, and dynamic weight adjustment to achieve information coordination and unification between multimodal features and enhance the performance of classification tasks.
[0069] Modal consistency verification: Modal consistency verification aims to reduce the prediction deviation of different modal features on the classification target and ensure the coordinated performance of each modal feature. It is specifically achieved through the following loss function: in, Represents the loss function of modality consistency verification, which is used to measure the prediction difference of different modalities on the classification target, Z i , Z j They represent the feature representation of modality i and modality j, respectively, which are the results of multimodal feature fusion. f(Z) represents the classification function, whose output is the predicted value of the classification target, usually a probability distribution or classification label. 2 Represents the Euclidean distance, which is used to measure the difference between the classification prediction results of different modalities.
[0070] Modal complementarity verification: Modal complementarity verification is used to evaluate the independent contribution of each modal feature to the classification target and enhance the complementarity of modal features through mutual information optimization. The optimization formula is as follows: in, represents the loss function of modal complementarity verification, which is used to evaluate and optimize the independent contribution of each modal feature. λ represents the trade-off coefficient, which is used to control the impact of modal complementarity verification on the overall optimization. I(Z i ; Y) mutual information, representing the modal feature Z i The amount of information between Z and the classification target Y, measuring i Independent contribution to Y.
[0071] Dynamic weight adjustment Dynamic weight adjustment updates the modal weights in real time according to the contribution of each modal feature to the classification task to achieve optimal resource allocation. The weight update rule is: in, represents the weight of modality i in the tth iteration, indicating the current contribution ratio of modality features to the classification task, represents the weight of mode i in the t+1th iteration, which is used for the next optimization, I(Z i ; Y) mutual information, indicating mode Z i Correlation with the classification target Y, ∑ j I(Z j ; Y) represents the total mutual information of all modal features and the classification target Y, which is used to normalize the calculation weights.
[0072] Comprehensive optimization goals Combining modal consistency verification and modal complementarity verification, a comprehensive optimization objective function of the dynamic verification mechanism is formed: in, Comprehensive optimization objective function, used to simultaneously optimize the consistency and complementarity of modal characteristics, The loss function for modal consistency verification reduces the prediction deviation between modal features. Loss function for modal complementarity verification, enhancing the independent contribution of modal features.
[0073] The prediction bias of different modal features on the classification target is reduced through modal consistency verification, and the independent contribution of each modal feature is enhanced through modal complementarity verification. In combination with dynamic weight adjustment, resource optimization allocation is achieved to ensure the coordination and unification of multimodal features in the latent space. Finally, through comprehensive optimization of the objective function, the dynamic verification mechanism effectively improves the collaborative performance of multimodal features and the overall performance of the classification task.
[0074] For step S4, this embodiment provides a technical solution for embedding optimized high-dimensional features into low-dimensional semantic space based on dimensionality reduction method, aiming to reduce the redundancy of high-dimensional data, retain the core information related to the classification target, and reduce the computational complexity. This step is based on manifold learning and linear transformation technology, and optimizes feature representation through high-dimensional to low-dimensional mapping, providing efficient and accurate low-dimensional feature input for the generation of classification labels.
[0075] It should be noted that this dimensionality reduction method further optimizes the compactness between features on the basis of ensuring the global structure and semantic consistency of the data, laying a solid data foundation for subsequent classification steps.
[0076] The optimized high-dimensional feature dimensionality reduction includes the following: Constructing neighborhood relationships of high-dimensional features As an option, this step first constructs the neighborhood relationship of the high-dimensional feature data points to preserve the local structure of the data. Specifically, the neighborhood of each data point is determined by calculating the similarity between the high-dimensional feature data points.
[0077] In one possible implementation, the similarity measure can be Euclidean distance or cosine similarity, which is defined as follows: d ij =||Z i -Z j || 2 Among them, d ij Represents data point Z i and Z j The Euclidean distance between i and Z j Represents two data points in a high-dimensional feature space.
[0078] Dimensionality reduction mapping Specifically, on the basis of constructing neighborhood relationships, the high-dimensional feature Z is mapped to the low-dimensional semantic space Z low , ensuring that the reduced-dimensional data can retain the consistency of the local neighborhood structure. Commonly used dimensionality reduction methods include principal component analysis and manifold learning methods.
[0079] In one implementation, a dimensionality reduction method based on manifold learning is used, and its optimization goal is as follows: in, represents the optimization objective function of dimensionality reduction, Z low,i and Z low,j Represents two data points in the low-dimensional semantic space, w ij Represents high-dimensional feature Z i and Z j The weights between are usually calculated by the Gaussian kernel function: Among them, l ij Represents high-dimensional feature Z i and Z j , σ represents the parameter of the Gaussian kernel, which is used to control the similarity decay rate.
[0080] It should be noted that by minimizing It can ensure that the features after dimensionality reduction maintain a neighborhood structure in the low-dimensional space similar to that in the high-dimensional space.
[0081] Optimizing the dimensionality reduction process In one implementation, in order to ensure the global structural consistency of the low-dimensional embedding features, a global semantic constraint can be introduced, and the optimized dimensionality reduction objective is defined as follows: in, Global optimization objective function, Cov(Z) and Cov(Z low ) represent the covariance matrices of high-dimensional features and low-dimensional features, respectively, and α and β are trade-off parameters used to control the balance between local neighborhood consistency and global structural consistency.
[0082] It is understood that by adding global semantic constraints, the reduced-dimensional features can further enhance the global structural expression ability of the data while retaining the consistency of the local neighborhood.
[0083] Generate low-dimensional embedding features For example, assuming that the input high-dimensional feature is 1024-dimensional, it is mapped to a 128-dimensional semantic space through a dimensionality reduction method, and the reduced-dimensional feature can be expressed as: Z low =DimReduction(Z) Among them, Z low Represents low-dimensional embedding features, and DimReduction represents the dimensionality reduction mapping function, including linear transformation or nonlinear methods.
[0084] It should be noted that the low-dimensional embedding feature Z low , is the key input for the subsequent classification label generation, ensuring the efficiency and accuracy of the classification task.
[0085] This embodiment describes in detail the specific process of reducing the dimension of the optimized high-dimensional features, including neighborhood relationship construction, dimension reduction mapping and optimization process, to ensure that the features after dimension reduction retain the core information in the low-dimensional space and maintain semantic consistency with the high-dimensional space.
[0086] For step S5, this embodiment proposes a solution based on building a vector index mechanism that supports query and feedback optimization to address the problem of efficiency and accuracy in dynamic classification query tasks. By establishing a dynamically updated vector index library and introducing a user feedback optimization mechanism, efficient query and adaptive optimization of the classification query system are achieved. The core of this step is to use the low-dimensional feature Z after dimensionality reduction. low Perform vectorized storage and indexing, and dynamically adjust index content to adapt to data updates and user needs.
[0087] It should be noted that this step supports fast query through vector index technology and combines the user feedback optimization mechanism to ensure that the classification query results can be dynamically adjusted, thereby improving the adaptability and reliability of the system.
[0088] Building a vector index that supports query and feedback optimization includes the following: Building a vector index library As an option, this step can be performed by embedding the low-dimensional feature Z low Perform vectorized storage and indexing to build a vector index library that supports efficient queries. Specifically, the vector index library is implemented based on the approximate nearest neighbor search algorithm.
[0089] In one possible implementation, the low-dimensional embedding feature Z low is stored as a vector set {Z low,1 ,Z low,2 ,…,Z low,n} and indexed by vector similarity (such as cosine similarity or Euclidean distance).
[0090] Exemplarily, the similarity calculation formula based on Euclidean distance is: d ij =||Z low,i -Z low,j || 2 Among them, d ij Represents vector Z low,i and Z low,j The Euclidean distance between low,i and Z low,j is a low-dimensional feature vector.
[0091] It should be noted that the vector index library uses ANN technology, which can significantly reduce the query time complexity while maintaining a high query accuracy.
[0092] User feedback collection and dynamic update Specifically, in order to ensure that the classification query results are consistent with user needs, this step optimizes the index content and classification model by collecting user feedback.
[0093] In one implementation, the optimization mechanism of user feedback includes the following steps: First, user feedback on the classification query results is collected, such as the classification labels annotated by the user or the preferences for the query results.
[0094] Then the feedback data is used as the new sample {Z feedback,1 ,Z feedback,2 ,…} Input index library, which is used to dynamically update the low-dimensional feature set.
[0095] The formula is: Z low,new =Z low,old ∪Z feedback Among them, Z low,newrepresents the updated low-dimensional feature set, Z low,old Represents the feature set in the original index library, Z feedback Indicates new samples generated by user feedback.
[0096] It is understood that the user feedback mechanism enables the classification system to adapt to changes in user needs by dynamically adjusting the content of the index library.
[0097] Feedback-driven classification model optimization As an option, this step further uses user feedback to optimize the classification model, thereby improving the performance of the classification label generation module.
[0098] In one possible implementation, user feedback is considered as a new sample with high confidence and used to retrain the classification model. The objective function of the optimization process is defined as follows: in, represents the optimized objective function, represents the loss function of the original classification model, represents the confidence-weighted loss function of user feedback, and γ is the trade-off coefficient, which is used to control the impact of user feedback on model optimization.
[0099] It should be noted that by introducing a feedback-driven optimization mechanism, the classification model can more accurately reflect user needs and improve the accuracy of classification results.
[0100] This embodiment describes in detail the technical solution for building a vector index that supports query and feedback optimization, including the construction of a vector index library, dynamic update of user feedback, and feedback-driven classification model optimization. Through efficient indexing technology and adaptive optimization mechanism, it is ensured that the classification query results can meet the dynamically changing user needs.
[0101] Please see attached Figure 2 ,The present invention also provides a digital archive classification query system based on large model index recognition, including: a data acquisition module for collecting multimodal data such as text, image, audio and structured numerical value; Data preprocessing module, used to standardize the collected multimodal data and generate a unified feature representation; Feature fusion module, used to optimize the information expression of multimodal data and eliminate redundant information; Verification and optimization module, used to verify the consistency of data features and optimize classification results; Dimensionality reduction classification module, used to embed feature data into low-dimensional semantic space and generate dynamic classification labels; Index query module, used to build a vector index library that supports fast query; The user feedback module is used to collect user feedback information and optimize the classification query model.
[0102] The function of the data acquisition module is to collect multimodal data (such as text, images, audio and structured numerical data) to provide the system with basic raw input data. This module solves the heterogeneity problem of different data modal sources and ensures the diversity and extensiveness of input data.
[0103] Specific content: Text data collection, providing semantic information of archives, such as titles, descriptions or tags; image data collection, providing visual information of archives, such as cover photos or scanned documents; audio data collection, providing voice recording content, such as interview recordings or meeting minutes; structured numerical collection, providing statistical information or timestamps, such as dates, classification numbers or statistical fields.
[0104] Significance: Provide sufficient data input for the multimodal system to ensure that subsequent modules can process and fuse multimodal data.
[0105] The function of the data preprocessing module is to standardize the collected multimodal data and convert the raw data into a standardized feature representation. This module solves the inconsistency of the structure, scale, and distribution of different modal data, ensuring the compatibility of the data in subsequent processes.
[0106] Specific content: text processing, word segmentation, stop word removal and semantic embedding generation; image processing, adjusting resolution, denoising and extracting visual features; audio processing, converting audio data into text and generating feature vectors through the speech-to-text model; numerical data processing, normalization to standardize numerical features.
[0107] Significance: Generate a unified multimodal feature representation for the system, eliminate the differences between modalities, and improve the efficiency and accuracy of subsequent processing.
[0108] The function of the feature fusion module is to compress, fuse and optimize the features of multimodal data to generate an efficient unified feature representation. This module eliminates data redundancy and retains the core information related to the task.
[0109] Specific content: Use deep learning models such as variational autoencoders to jointly learn multimodal features; maximize feature expression capabilities by optimizing mutual information in the latent feature space; weight different modal features and dynamically adjust them according to their task contribution.
[0110] Significance: It solves the problems of information redundancy and representation differences between multimodal data, and provides more compact and effective feature input for classification tasks.
[0111] Verification and Optimization Module: Verify the quality and consistency of multimodal features and optimize the classification results through dynamic adjustment. This module ensures that different modalities can work together in the classification task and improves the robustness of feature expression.
[0112] Specific content: Modal consistency verification: reduce the deviation of different modal features in the prediction results; Modal complementarity verification: evaluate and optimize the independent contribution of each modal feature; Dynamic weight adjustment: dynamically assign weights according to the importance of each modal feature.
[0113] Significance: Improve the classification system's ability to comprehensively utilize multimodal data, optimize feature synergy effects, and enhance classification accuracy and stability.
[0114] The dimensionality reduction classification module embeds high-dimensional features into low-dimensional semantic space, reduces the data dimension while retaining task-related information, and generates dynamic classification labels. This module improves the system's computing efficiency and classification performance.
[0115] Specific content: Map high-dimensional features to low-dimensional semantic space through dimensionality reduction technology; retain the global structure and local consistency of data during the dimensionality reduction process; generate dynamic classification labels that combine multi-dimensional information such as time, topic, priority, etc. to meet the needs of different application scenarios.
[0116] Significance: While reducing data complexity, it improves classification efficiency and accuracy and provides low-dimensional feature input for subsequent indexing and querying.
[0117] The function of the index query module is to build a vector index library that supports fast query and use low-dimensional features to achieve efficient classification query. This module ensures the real-time and high accuracy of the query.
[0118] Specific content: Use approximate nearest neighbor technology to index and store low-dimensional features; achieve fast retrieval through vector similarity (such as cosine similarity or Euclidean distance); support dynamic update of index content to meet data changes and user needs.
[0119] Significance: It provides low-latency classification query capabilities and provides users with fast and accurate query results.
[0120] Function of the user feedback module: Collect user feedback on classification query results, and use this feedback to dynamically optimize classification models and index content.
[0121] This module improves the system's adaptability and user satisfaction.
[0122] Specific content: Dynamically add new samples fed back by users to the index library and update vector features; optimize the loss function of the classification model based on feedback information to enhance the model's adaptability to user needs; adjust the classification label content to ensure that the classification results can dynamically adapt to user preferences.
[0123] Significance: A closed-loop optimization system based on user feedback is built, so that the classification results can be continuously improved and dynamically adapted to changing needs.
[0124] Each module plays a vital role in the system. The complete technical chain from data collection to final query ensures that the system has the ability to process multimodal data, optimize classification tasks, and dynamically adapt to user needs. The functional design of each module complements each other to form an efficient, accurate, and dynamically optimized classification query system.
[0125] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A digital archive classification query method based on large model index recognition, characterized in that: The following steps are involved: Generate a unified feature representation through the collection and processing of multimodal data; Optimize the information expression of multimodal data based on feature fusion method; Use dynamic verification mechanisms to improve the consistency of data features and the accuracy of classification results; The optimized features are embedded into the low-dimensional semantic space through the dimensionality reduction method to generate classification labels; Build a vector index that supports query and feedback optimization.
2. The digital archive classification query method based on large model index recognition according to claim 1 is characterized in that: The collection of the multimodal data comprises the following steps: Collect text data, and perform word segmentation and embedding extraction on the text data through a pre-trained language model to generate a semantic feature vector; collect image data, perform feature extraction on the image data through a convolutional neural network to generate a high-dimensional visual feature vector; Collect audio data, transcribe it into text data through a speech-to-text model, and then extract semantic feature vectors; Collect structured numerical data and adjust the data range through normalization methods to generate standardized numerical features.
3. The digital archive classification query method based on large model index recognition according to claim 1 is characterized in that: The processing of the multimodal data comprises the following steps: Use language models to extract semantic embeddings from text data; Use convolutional neural networks to extract feature embedding from image data; Use speech recognition models to transcribe audio data and extract semantic features; Normalize structured numeric data.
4. The digital archive classification query method based on large model index recognition according to claim 1 is characterized in that: The feature fusion method optimization comprises the following steps: The high-dimensional features of multimodal data are input into the variational autoencoder model, and the multimodal features are compressed by the encoder to generate latent feature representations; In the latent feature space, the optimization objective is defined as: Where X represents the multimodal features of the input, Z represents the latent feature representation, Y represents the classification target, I(X; Z) represents the mutual information between the multimodal features and the latent features, I(Z; Y) represents the mutual information between the latent features and the classification target, and β is a hyperparameter that balances compression and classification accuracy. Optimize the latent feature representation to maximize the relevance of the latent feature to the classification task, while removing redundant features and reducing the duplication and noise interference between multimodal features; Generate compressed and optimized fusion feature representation for subsequent classification processing.
5. The digital archive classification query method based on large model index recognition according to claim 1 is characterized in that: The dynamic verification mechanism comprises the following steps: Modal consistency verification: By calculating different modal characteristics Z i and Z j The prediction difference on the target category Y optimizes the consistency between modalities and defines the loss function as: Among them, Z i and Z j represents the feature representation of modalities i and j, f(Z) represents the classification function, ||·|| 2 represents the Euclidean distance; Modal complementarity verification: Evaluate each modal feature Z i The independent contribution to the target classification Y is defined as the complementary loss function: Among them, I(Z i ; Y) represents the modal characteristic Z i The mutual information with the target classification Y, λ is the trade-off factor; Dynamic weight adjustment: According to each mode characteristic Z i The update modal weight related to the target classification Y is updated as follows: in, represents the weight of mode i in the tth iteration, is the weight for the next iteration; Comprehensive optimization objective function: Combining consistency verification, complementarity verification and dynamic weight adjustment, the final optimization goal is: in, represents the comprehensive optimization objective function of the dynamic verification mechanism, represents the loss function of modal consistency verification, which is used to minimize the difference between different modal features in the classification target prediction results. Represents the loss function for modal complementarity verification, which is used to evaluate the independent contribution of each modal feature.
6. The digital archive classification query method based on large model index recognition according to claim 1 is characterized in that: The dimensionality reduction method comprises the following steps: Construct the neighborhood relationship of high-dimensional feature data points and determine the neighborhood of each data point by calculating the similarity between data points: On the basis of maintaining the neighborhood relationship, high-dimensional features are mapped to low-dimensional space to generate low-dimensional embedded features; Optimize the dimensionality reduction process to ensure that the low-dimensional embedded features can preserve the global structure and semantic consistency of the high-dimensional data.
7. The digital archive classification query method based on large model index recognition according to claim 1 is characterized in that: The generation of the classification label includes the following steps: Input low-dimensional embedded features into the large language model, combine low-dimensional feature representation with contextual semantic information to generate initial classification labels; refine the initial classification labels according to the time dimension, subject dimension and priority dimension of the archival data; Dynamically adjust the classification label content and generate the final classification label based on user query requirements and archive background information.
8. The digital archive classification query method based on large model index recognition according to claim 1 is characterized in that: The construction of support query and feedback optimization includes the following steps: Collect user feedback on classification query results and input the feedback information into the model as new samples; Optimize the dynamic classification label generation module through feedback samples and adjust the classification labels to more accurately reflect user needs; Use feedback samples to update the feature fusion module to improve the accuracy of multimodal data feature representation and classification effect; The low-dimensional vector index library is dynamically adjusted so that new samples and optimized features can be updated to the query system in real time, ensuring the dynamic and accurate nature of the query results.
9. The digital archive classification query method based on large model index recognition according to claim 1 is characterized in that: The construction of the vector index includes the following steps: Store low-dimensional embedded features into a vector database and index features based on vector similarity; Use the approximate nearest neighbor search algorithm to optimize the index to ensure low latency and high efficiency of the query process; In response to the update of dynamic classification labels, the feature data in the vector index is adjusted in real time to support multi-dimensional query and matching of classification labels.
10. A digital archive classification query system based on large model index recognition, characterized in that: include: Data collection module, used to collect multimodal data such as text, images, audio and structured numerical values; Data preprocessing module, used to standardize the collected multimodal data and generate a unified feature representation; Feature fusion module, used to optimize the information expression of multimodal data and eliminate redundant information; Verification and optimization module, used to verify the consistency of data features and optimize classification results; Dimensionality reduction classification module, used to embed feature data into low-dimensional semantic space and generate dynamic classification labels; Index query module, used to build a vector index library that supports fast query; The user feedback module is used to collect user feedback information and optimize the classification query model.
Citation Information
Patent Citations
Multi-modal data fusion method and device based on tensor and mutual information
CN116975776A
Multi-modal recognition method and device based on large language model, electronic equipment and storage medium
CN118410457A
Contract question and answer method based on multiple modes
CN118656469A
Intelligent access control management method and system based on multi-mode identification and Internet of Things technology
CN118968665A
Layered contrast anti-fact learning method and system oriented to visual question and answer model
CN119166795A
Cited By
Large-scale article classification method and system
CN120744652A
Video recommendation method and model training method and device for video recommendation
CN120994871A
Digital archive multi-dimensional label automatic classification method
CN121901862A