Intelligent Retrieval Methods and Systems for Associating Images with Content in PDF Documents

By preprocessing PDF documents and constructing an image-text index structure through an associated indexing mechanism, the problem of the inability to effectively associate images and text content in existing PDF document retrieval technologies is solved, achieving efficient and accurate retrieval results.

CN120407819BActive Publication Date: 2025-10-28BEIJING GUANGLIANDA YUNTU DREAM TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510927429.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-28
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Existing PDF document retrieval technologies cannot effectively link images and text content, resulting in insufficient retrieval efficiency and accuracy.

Method used

By preprocessing PDF documents, extracting image and text content, generating image processing results and content processing results, constructing an image and text index structure based on the association index mechanism, and performing retrieval matching based on this structure.

Benefits of technology

It enables intelligent association retrieval of image and text content, improving retrieval efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407819B_ABST
    Figure CN120407819B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent retrieval method and system for associating images and content in PDF documents, belonging to the field of data processing technology. The method includes: preprocessing the PDF document using a document preprocessing strategy to obtain image processing results and content processing results; performing association analysis on the image processing results and content processing results according to an association indexing mechanism to obtain an image-text index structure; and performing retrieval matching on the target retrieval request of the PDF document based on the image-text index structure to obtain target retrieval information. This invention solves the technical problem that existing PDF document retrieval methods cannot effectively associate images and text content, resulting in insufficient retrieval efficiency and accuracy. It achieves the technical effect of intelligently associating image and text content through the construction of an image-text index structure, thereby improving retrieval efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, specifically to an intelligent retrieval method and system for associating images with content in PDF documents. Background Technology

[0002] With the widespread use of digital office tools and electronic documents, PDF documents are widely used for storing and sharing files containing rich text and images due to their strong compatibility and stable format. However, most existing PDF document retrieval technologies only support text-based keyword searches, with extremely limited capabilities for retrieving image content, and they cannot effectively link the inherent relationship between images and text. This often forces users to manually browse documents when searching for text information associated with a specific image, or to find a corresponding image based on a text description, wasting a lot of time and effort, resulting in low retrieval efficiency and a poor user experience. Summary of the Invention

[0003] This application provides an intelligent retrieval method and system for associating images with content in PDF documents, which solves the technical problem that existing PDF document retrieval methods cannot effectively associate images with text content, resulting in insufficient retrieval efficiency and accuracy.

[0004] The first aspect of this application provides an intelligent retrieval method for associating images and content in a PDF document. The method includes: invoking a document preprocessing strategy to preprocess the PDF document to obtain document processing results, wherein the document processing results include image processing results and content processing results; performing association analysis on the image processing results and the content processing results according to an association indexing mechanism to obtain an image-text index structure; and performing a retrieval matching on a target retrieval request of the PDF document based on the image-text index structure to obtain target retrieval information.

[0005] A second aspect of this application provides an intelligent retrieval system for associating images and content in PDF documents. The system includes: a document preprocessing module, which invokes a document preprocessing strategy to preprocess the PDF document and obtain document processing results, wherein the document processing results include image processing results and content processing results; an image-text association analysis module, which performs association analysis on the image processing results and the content processing results according to an association indexing mechanism to obtain an image-text index structure; and a retrieval matching module, which performs retrieval matching on target retrieval requests in the PDF document based on the image-text index structure to obtain target retrieval information.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0007] This application provides an intelligent retrieval method and system for associating images and content in PDF documents, relating to the field of data processing technology. It extracts images and text content from PDFs through document preprocessing strategies, generates image processing results and content processing results, constructs an image-text index structure based on an association index mechanism, and uses this index structure to match target retrieval requests, outputting corresponding retrieval results. This solves the technical problem that existing PDF document retrieval methods cannot effectively associate images and text content, resulting in insufficient retrieval efficiency and accuracy. It achieves the technical effect of intelligently associating image and text content through the construction of an image-text index structure, thereby improving retrieval efficiency and accuracy. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 A schematic diagram of the intelligent retrieval method for associating images and content in PDF documents provided in this application embodiment;

[0010] Figure 2 This is a schematic diagram of the structure of an intelligent retrieval system that associates images and content in a PDF document, as provided in an embodiment of this application.

[0011] Figure labeling: Document preprocessing module 11, image and text association analysis module 12, retrieval and matching module 13. Detailed Implementation

[0012] This application provides an intelligent retrieval method and system for associating images with content in PDF documents, which solves the technical problem that existing PDF document retrieval methods cannot effectively associate images with text content, resulting in insufficient retrieval efficiency and accuracy.

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0014] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0015] Example 1, as Figure 1 As shown, this application provides an intelligent retrieval method for associating images with content in PDF documents. This method includes:

[0016] P10: Retrieve the document preprocessing strategy to preprocess the PDF document and obtain the document processing result, wherein the document processing result includes image processing result and content processing result.

[0017] Furthermore, step P10 in this embodiment of the application also includes:

[0018] P11: Extract the PDF document using object structure as a discrimination constraint to obtain the extraction result; P12: Process the image objects in the extraction result according to the image strategy in the document preprocessing strategy to obtain the image processing result; wherein, it includes:

[0019] P12-1: Extract the first object and the second object from the image object; P12-2: Perform standardization processing on the first object and the second object sequentially according to the image standardization scheme in the image strategy to obtain the first standard image and the second standard image respectively; P12-3: Calculate the first hash value of the first standard image and the second hash value of the second standard image sequentially; P12-4: Compare the first hash value and the second hash value to obtain the Hamming distance; P12-5: If the Hamming distance does not reach the predetermined threshold, then divide the first object and the second object to obtain the image processing result.

[0020] It should be understood that, in order to construct the basic data source for the subsequent image-text association index, the preset document preprocessing strategy is first invoked to systematically parse the input PDF document, obtaining document processing results including image processing results and content processing results. Among them, the image processing results can be obtained through a multi-level image object extraction and structural comparison process, and the content processing results include text semantic units, paragraph structure, and their page layout information.

[0021] Specifically, to achieve standardized extraction of image processing results, an object structure analysis mechanism is first introduced. This mechanism uses object structure as a constraint to perform preliminary structural analysis of the PDF document. Here, object structure refers to the organization and hierarchical relationships of various elements (such as text, images, tables, etc.) within the PDF document. By analyzing the object structure, image objects and other relevant elements in the document can be accurately identified, thus obtaining preliminary extraction results. This process can be based on the internal format and structural features of the PDF document, effectively avoiding the erroneous extraction of non-target objects and ensuring the accuracy and completeness of the extraction results.

[0022] Next, the image objects in the extracted results are processed according to the image strategy in the document preprocessing strategy to obtain the image processing results. The specific processing flow is as follows:

[0023] First, extract the first and second objects from the image object. These can be different regions within the image object, different types of image elements, or different versions of the same image. For example, in a PDF document containing multiple illustrations, the first object might be the main illustration, while the second object might be a related auxiliary illustration or legend.

[0024] Next, an image standardization scheme is introduced to perform image standardization processing on the first and second objects respectively, resulting in a first standard image and a second standard image with uniform size, grayscale range, and image ratio. The image standardization scheme is a preprocessing method for images, aiming to convert images into a uniform format and quality standard. This scheme may include image normalization operations such as grayscale conversion, edge enhancement, and size standardization (e.g., normalization to 128×128 pixels) to eliminate differences in resolution, brightness, or compression format between the original images. Through standardization processing, the first standard image and the second standard image can be obtained, thereby eliminating format and quality differences between the images and providing a consistent basis for subsequent feature extraction and comparison.

[0025] Next, the first hash value of the first standard image and the second hash value of the second standard image are calculated sequentially. A hash value is a digital digest generated from image data using a specific algorithm, uniquely identifying the content of the image. By calculating the hash value, the similarity between two images can be quickly compared. In this step, the hash value calculation is based on the standardized images; the first hash value and the second hash value represent the feature-compressed representations of the two standard images, respectively.

[0026] Subsequently, the Hamming distance is calculated by comparing the first and second hash values. The Hamming distance refers to the number of distinct bits between two hash values, used to measure their similarity. A smaller Hamming distance indicates greater similarity between the two hash values, and consequently, greater similarity in their corresponding image content. By calculating the Hamming distance, the similarity between the first and second objects can be quickly determined, providing a basis for further image processing.

[0027] If the Hamming distance does not reach a predetermined threshold, the first and second objects are separated to obtain the image processing result. The predetermined threshold is a parameter pre-set based on the actual application scenario and image similarity requirements, used to determine whether two images are sufficiently similar. If the Hamming distance does not reach the predetermined threshold, it indicates that the similarity between the first and second objects is insufficient, and they need to be separated, i.e., processed as independent image objects. This avoids merging image objects with low similarity, thus ensuring the accuracy and effectiveness of the image processing result.

[0028] Through the aforementioned image processing chain, this application enables precise preprocessing and analysis of image objects in PDF documents, providing a high-quality data foundation for subsequent image-content association analysis. This method balances structural analysis and perceptual feature extraction, exhibiting good adaptability and scalability, and is particularly suitable for document processing scenarios with complex text-image mixes and diverse image styles.

[0029] Furthermore, step P10 in this embodiment of the application also includes:

[0030] P13: The text objects in the extracted results are processed according to the text strategy in the document preprocessing strategy to obtain the content processing result; wherein, it includes:

[0031] P13-1: Extract the third object from the text object; P13-2: Perform semantic analysis on the third object according to the text strategy to obtain semantic features; P13-3: Construct a semantic vector based on the semantic features and combine it with the third object to obtain a mapping relationship; P13-4: Form the content processing result based on the mapping relationship.

[0032] Optionally, step P10 includes not only structured preprocessing of image objects in the PDF document, but also semantic feature extraction and semantic structure mapping operations for text objects to form a complete content processing result.

[0033] Specifically, after preprocessing the image objects, the extracted text objects are processed according to the text strategy in the document preprocessing strategy to obtain the content processing results. First, third-party objects are extracted from the text objects. Here, "third-party objects" refer to key text paragraphs or sentences within the text objects. This text content is usually directly related to the image objects, such as the image title, descriptive text, or legend. By analyzing the structural and semantic features of the text objects, these key text paragraphs can be accurately identified. For example, using text segmentation algorithms in Natural Language Processing (NLP) technology, combined with text formatting features (such as font size, bold, italics, etc.) and semantic cues (such as keywords, contextual semantic coherence, etc.), third-party objects related to the image objects are filtered from the extracted text objects.

[0034] Next, semantic analysis is performed on the third object according to the text strategy to obtain semantic features. Semantic analysis refers to the in-depth parsing of text content using natural language processing techniques to extract semantic information. Specific operations include word segmentation of the third object, dividing the text into independent lexical units; part-of-speech tagging to determine the part of speech of each word; and syntactic analysis to analyze the structure and grammatical relationships of sentences. Based on this, semantic understanding algorithms are used to extract the semantic features of the text. For example, word embedding models (such as Word2Vec and BERT) are used to map words to a semantic space, generating semantic vectors for each word. These vectors are then aggregated to obtain the semantic features of the text paragraph. For instance, for a text paragraph describing image content, semantic analysis can extract keywords related to the image topic and their semantic relationships, thus obtaining the semantic features of the text paragraph.

[0035] Next, based on the obtained semantic features, semantic vectors are constructed, and a mapping relationship is obtained by combining them with a third object. A semantic vector is a high-dimensional vector capable of representing the semantic information of text. By converting semantic features into semantic vectors, semantic similarity comparisons between texts can be achieved. In this application, the extracted semantic features can be input into a pre-trained word embedding model to generate corresponding semantic vectors. Then, the third object is associated with the corresponding semantic vector to form a mapping relationship. For example, for the title text of an image, its semantic vector can reflect the semantic content of the title, and the mapping relationship between the title text and the semantic vector provides the foundation for subsequent image-text association analysis.

[0036] Finally, the content processing result is formed based on the above mapping relationship. The content processing result is text data represented in a structured and semantic form, containing not only the original text content but also its semantic features and semantic vector information. In this way, the text object is transformed into a form that facilitates subsequent processing and analysis, enabling effective correlation analysis with the image processing result. For example, the content processing result can be a data structure containing text paragraphs, semantic vectors, and mapping relationships, used for subsequent image-text correlation retrieval.

[0037] By implementing the above steps, semantic information in PDF documents can be extracted, represented, and structured without disrupting the original layout structure, providing an executable processing foundation for subsequent text-image alignment, intelligent indexing, and efficient retrieval.

[0038] P20: Based on the association indexing mechanism, the image processing results and the content processing results are analyzed to obtain the image and text index structure.

[0039] Furthermore, step P20 in this embodiment of the application also includes:

[0040] P21: Obtain any image group from the image processing results; P22: Perform multi-dimensional feature collection on the first arbitrary image in the arbitrary image group to obtain the first arbitrary feature parameter; P23: Analyze the first arbitrary feature parameter to determine the target feature parameter; P24: Activate the label classifier in the association indexing mechanism to classify and analyze the target feature parameter to obtain the arbitrary category label of the arbitrary image group; P25: Traverse the content processing results with the arbitrary category label as the traversal constraint to obtain the traversal result; P26: Establish the image and text index structure based on the traversal result.

[0041] Specifically, based on the image processing and content processing results, a semantic association structure between images and text is constructed, namely, an image-text index structure. This structure is the key supporting data organization form for subsequent intelligent retrieval. Through feature matching and semantic classification analysis based on images and text, semantic coupling of the two types of heterogeneous information can be achieved, thereby completing one-to-one mapping or one-to-many, one-to-one linkage organization of images and text.

[0042] In the specific execution process, an arbitrary image group is first obtained from the image processing results. Here, an arbitrary image group refers to a collection of image objects obtained after preprocessing. These image objects may include the main image, local detail images, or other related image elements. The acquisition of image groups can be achieved by classifying or grouping the image processing results. For example, image objects can be grouped into different groups based on their source, type, or similarity.

[0043] Next, multi-dimensional feature collection is performed on the first arbitrary image in the arbitrary image group to obtain the first arbitrary feature parameters. Multi-dimensional feature collection refers to extracting features from different aspects of the image to comprehensively represent its content and attributes. Specific operations include visual feature extraction, using image processing algorithms (such as SIFT, SURF, ORB, etc.) to extract local feature points of the image. These feature points can reflect the image's texture, shape, and edge information. Simultaneously, the image's color histogram, texture features (such as GLCM), and shape features (such as contour information) are calculated to obtain the image's visual features. Furthermore, semantic feature extraction is also required. Combining the image's contextual information and semantic labels, semantic features of the image are extracted. For example, image recognition techniques (such as deep learning models) are used to classify or label the image to obtain its semantic category (such as people, scenery, charts, etc.) and keyword descriptions. Finally, spatial feature extraction is performed to analyze the image's positional information within the document, including its coordinates, size, and spatial relationship with other elements (such as text). These spatial features can be used for subsequent association analysis to determine the proximity between the image and the text.

[0044] After collecting the first arbitrary feature parameters, these parameters are analyzed to determine the target feature parameters, which are the subset of features that play a key role in representing the image content and semantics. For example, feature selection algorithms (such as those based on information gain or principal component analysis (PCA)) can be used to filter the collected multi-dimensional features, removing redundant or irrelevant features and retaining the features most valuable for image classification and association analysis. For instance, for an image containing people and scenery, the target feature parameters might include facial features of the people, texture features of the scenery, and the overall semantic category of the image.

[0045] Subsequently, the label classifier in the association indexing mechanism is activated to perform classification analysis on the target feature parameters, obtaining arbitrary category labels for any group of images. The label classifier is a pre-trained machine learning model capable of classifying images based on input feature parameters and assigning them corresponding category labels. For example, a classifier built using a convolutional neural network (CNN) can classify images into different semantic categories (such as people, landscapes, charts, etc.) and output corresponding category labels. Through classification analysis, one or more category labels can be assigned to each group of images; these labels will serve as an important basis for subsequent association analysis.

[0046] Next, the content processing results are traversed using arbitrary category labels as traversal constraints to obtain the traversal results. Content processing results refer to the preprocessed text data, including the semantic features and semantic vectors of the text objects. The traversal process involves searching for related text content in the text data based on the image's category label. Specific operations include semantic matching, matching the image's category label with the text's semantic features to find text paragraphs or sentences related to the image category. For example, if the image's category label is "landscape," then the text data is searched for text content containing semantic features related to "landscape." Simultaneously, the spatial location information of the image and text within the document is combined to analyze the proximity between the image and text. If the text content and the image are spatially adjacent (e.g., located near the image or in the same paragraph), they are considered to have a high degree of relevance.

[0047] Finally, an image-text index structure is built based on the traversal results to store the associations between images and text. Specifically, this involves associating the image's category label, feature parameters, and corresponding text content, and storing this association in a structured manner. For example, an index table can be constructed where each entry contains the image's identifier, category label, feature vector, and the identifier and semantic vector of the associated text paragraph. This index structure enables rapid image-text retrieval and matching, improving retrieval efficiency and accuracy.

[0048] Furthermore, step P23 in this embodiment of the application also includes:

[0049] P23-1: Multi-dimensional feature collection obtains the second arbitrary feature parameter of the second arbitrary image in the arbitrary image group; P23-2: Analyze the first arbitrary feature parameter and the second arbitrary feature parameter to determine the target feature parameter; wherein, before analyzing the first arbitrary feature parameter and the second arbitrary feature parameter to determine the target feature parameter, the process includes:

[0050] P23-21a: Perform discrete cosine transform on the first arbitrary image and the second arbitrary image in sequence to obtain the first transform coefficient and the second transform coefficient respectively; P23-22a: Based on the first DC coefficient and the first AC coefficient in the first transform coefficient, construct the first arbitrary feature parameter; P23-23a: Based on the second DC coefficient and the second AC coefficient in the second transform coefficient, construct the second arbitrary feature parameter.

[0051] Optionally, the process of determining the target feature parameters can be further refined to enhance the accuracy and robustness of image feature analysis. In particular, in-depth calculations can be performed on the structural similarity and content feature differences between different images in the image group to generate image feature vectors that contain both local and global information, which can be used for subsequent determination of target feature parameters and preparation of input for image classifiers.

[0052] Specifically, firstly, multi-dimensional feature collection is performed on the second arbitrary image in the arbitrary image group to obtain the second arbitrary feature parameters. Here, the "second arbitrary image" refers to another image object in the image group, which, together with the first arbitrary image, is used to further analyze the features of the image group. The method of multi-dimensional feature collection is similar to the feature collection of the first arbitrary image in step P22, including visual feature extraction, semantic feature extraction, and spatial feature extraction, to comprehensively characterize the content and attributes of the second arbitrary image.

[0053] Next, before analyzing the first and second arbitrary feature parameters and determining the target feature parameters, Discrete Cosine Transform (DCT) is performed sequentially on the first and second arbitrary images. Discrete Cosine Transform is a widely used transformation method in image compression and feature extraction. Its function is to map the image from the spatial domain to the frequency domain, concentrating the main energy of the image in a small number of low-frequency components, facilitating compression and analysis. By dividing the image into fixed-size (e.g., 8×8) blocks, a two-dimensional DCT operation is performed on each block to obtain the first transform coefficients of the first image and the second transform coefficients of the second image. These transform coefficients reflect the energy distribution of the image at different frequency components, where the DC coefficient (DC) represents the average brightness of the image, and the AC coefficient (AC) reflects the detail and texture information of the image.

[0054] Next, based on the first DC coefficient and the first AC coefficient in the first transform coefficients, a first arbitrary feature parameter is constructed. For example, the first DC coefficient is used as the overall brightness feature of the image, while certain key coefficients in the first AC coefficients (such as low-frequency AC coefficients) are used as texture and detail features of the image. Following a preset frequency domain feature construction strategy, the DC value is combined with several of the most significant AC values ​​into an ordered vector to form the first arbitrary feature parameter. In this way, the coefficients after the DCT transform can be combined with the previously collected multi-dimensional feature parameters to further enrich the feature representation of the first arbitrary image.

[0055] Similarly, based on the second DC coefficient and the second AC coefficient in the second transformation coefficients, a second arbitrary feature parameter is constructed. Similar to the first arbitrary image, the second DC coefficient is used as the overall brightness feature of the image, while the key coefficients in the second AC coefficients are used as the texture and detail features of the image. Through this combination, the feature representation of the second arbitrary image is also enhanced. In this process, an energy threshold can be applied to the AC coefficients, retaining the top few terms with a cumulative energy of 95% to reduce redundant computation.

[0056] After constructing the two frequency domain feature parameters, a comprehensive analysis is performed on the first and second arbitrary feature parameters. For example, Euclidean distance, cosine similarity, or projection similarity evaluation methods based on principal component analysis (PCA) are used to calculate the feature correlation and difference between the two sets of features, and then a representative and highly discriminative feature subset is determined as the target feature parameter.

[0057] This process considers not only the visual and semantic features of the image, but also analyzes its frequency characteristics through DCT, enabling the target feature parameters to more comprehensively characterize the image's content and attributes. Ultimately, these target feature parameters will be used for subsequent classification analysis and the construction of an image-text index structure, providing more accurate and efficient support for intelligent image-content association retrieval in PDF documents.

[0058] Furthermore, step P23-2 in the embodiments of this application also includes:

[0059] P23-21: Obtain the first difference between the first DC coefficient and the second DC coefficient; P23-22: Obtain the second difference between the first AC coefficient and the second AC coefficient; P23-23: Perform a variation weighted calculation on the first difference and the second difference to obtain the first deviation index; P23-24: Optimize to obtain the target pair with the first deviation index as the target; P23-25: Take the mean of the first target image group and the second target image group in the target pair to form the target feature parameters.

[0060] In one possible embodiment of this application, in order to further improve the discrimination capability and structural stability of image feature parameters, the determination process of target feature parameters can be further refined. Image pairs with the greatest difference representativeness can be automatically selected from the image group, and the final target feature parameters can be generated by their structural mean.

[0061] First, the first DC coefficient (DC1) and the second DC coefficient (DC2) obtained after performing discrete cosine transform on the first and second arbitrary images are acquired, and their difference is calculated to obtain the first difference. This difference reflects the difference between the two images in terms of overall brightness or grayscale reference, and can be used to evaluate the degree of change in the overall structure or main region morphology of the image. By calculating the difference between the DC coefficients of the first and second arbitrary images, the difference in overall brightness between the two images can be quantified. For example, if the first DC coefficient is DC1 and the second DC coefficient is DC2, then the first difference... It can be represented as .

[0062] Next, the first AC coefficient set (AC1) and the second AC coefficient set (AC2) are obtained, representing the detail changes in the non-DC portion of the two images in the frequency domain, respectively. The differences between the corresponding AC coefficients are calculated one by one to form a difference vector, which serves as the second difference. This difference vector reveals the differences in high-frequency information such as texture and edge structure between the images.

[0063] Next, a variation-weighted calculation mechanism is introduced, in which the first and second differences are used together in the calculation to generate a numerical index that comprehensively measures the degree of difference between the two images, denoted as the first deviation index. The specific implementation of mutation weighting can adopt the following model:

[0064] ; where α and β are weighting coefficients used to adjust the balance between the overall brightness difference and detail variability; The variance or mean squared error of the difference vector represents the degree of structural complexity of the images. The larger the value of this bias index, the stronger the expressive differences between the image pairs, and the better their ability to distinguish categories.

[0065] Subsequently, the target pair is obtained by optimizing the image pair with the largest first deviation index. The target pair is the image pair with the largest deviation index in the image group. By maximizing the first deviation index, the image pair with the greatest difference in brightness and texture can be found. This process can be achieved by traversing all image pairs in the image group and calculating the first deviation index for each pair. Finally, the image pair with the largest first deviation index is selected as the target pair.

[0066] Finally, based on the identified target image pairs, all feature parameters of the first and second target image groups are extracted and averaged dimension-wise. Specifically, for the frequency domain feature vector formed by the combination of DC and AC coefficients, the average of the corresponding values ​​of the two vectors in each dimension is taken, ultimately forming a target feature parameter vector representing the overall characteristics of the image group. This vector not only comprehensively preserves the average structural features among representative images but also avoids the bias caused by outliers in the overall feature representation, enhancing the classification model's ability to express the commonalities in the structure of the image group. These target feature parameters will be used for subsequent classification analysis and the construction of the image-text index structure, providing more accurate and efficient support for intelligent image and content association retrieval in PDF documents.

[0067] Furthermore, step P24 in this embodiment of the application also includes:

[0068] P24-1: The target DC coefficient is classified and analyzed by the label classifier to obtain the target object type; P24-2: The target AC coefficient is classified and analyzed by the label classifier to obtain the target scene type; P24-3: The target object type and the target scene type constitute the arbitrary category label.

[0069] Specifically, to enhance the fine-grained expressive power of image semantic labels, the generation process of arbitrary category labels can be further refined. By classifying and analyzing the target DC coefficient and the target AC coefficient separately, the target object type and the target scene type can be obtained, and the two can be combined into arbitrary category labels.

[0070] Specifically, the target DC coefficient is first analyzed and classified using a label classifier to determine the target object type. The target DC coefficient is the DC coefficient portion extracted from the target feature parameters, primarily reflecting the overall brightness information of the image. The label classifier is a pre-trained machine learning model capable of classifying images based on input feature parameters. By analyzing the target DC coefficient, the main object types contained in the image can be identified. For example, if the features of the target DC coefficient best match the features of the "person" category, the classifier will output "person" as the target object type.

[0071] Next, the target communication coefficients are classified and analyzed using a label classifier to determine the target scene type. The target communication coefficients are the communication coefficients extracted from the target feature parameters; they primarily reflect the texture and detail information of the image. Similarly, by analyzing the target communication coefficients using a label classifier, the scene type of the image can be identified. For example, if the features of the target communication coefficients best match the features of the "landscape" category, the classifier will output "landscape" as the target scene type.

[0072] Finally, the target object type and target scene type are combined into an arbitrary category label. An arbitrary category label is a comprehensive label that includes not only information about the main objects in the image but also information about the scene in which the image is located. This combination method can more comprehensively describe the content of the image, providing richer semantic information for subsequent association analysis. For example, if the target object type is "people" and the target scene type is "landscape," then the arbitrary category label can be represented as "people-landscape."

[0073] By designing a dual-channel classification path in this step, we can move beyond the traditional single-dimensional image classification method and construct a multi-level, multi-label semantic recognition framework based on frequency domain features. This effectively improves the semantic richness and discrimination accuracy of image nodes in the image-text index structure, enabling the system to have stronger adaptability and intelligent analysis capabilities when processing diverse document structures (such as research reports, patent specifications, and teaching materials).

[0074] P30: Based on the aforementioned image and text index structure, perform a search and match on the target search request for the PDF document to obtain the target search information.

[0075] Furthermore, step P30 in this embodiment of the application also includes:

[0076] P31: Determine whether the target retrieval request contains an object; P32: If it does, coordinate with the image and text index structure to search and match the object in the PDF document to obtain the target retrieval information; P33: If it does not exist, determine whether the target retrieval request contains a scene; P34: If it does, coordinate with the image and text index structure to search and match the scene in the PDF document to obtain the target retrieval information.

[0077] It should be understood that during the retrieval phase, the system first receives the target retrieval request input by the user, and then performs multimodal element decomposition on the request statement based on the natural language parser to extract possible object nouns, scene description words, and contextual limiting information. Subsequently, the intent recognition engine is called to perform semantic annotation on the parsing results, generating a retrieval vector containing a set of object candidates, a set of scene candidates, and a query confidence threshold.

[0078] When the retrieval process enters the determination stage, the object existence determination module is invoked to quickly scan the aforementioned retrieval vector. If the retrieval vector contains at least one high-confidence entity mapped to the "target object type" tag in the image-text index structure, it is determined that "the target retrieval request contains an object." The object-image-text coupled retrieval channel is then immediately activated to extract object-related feature information from the retrieval request, such as the object's name, descriptive keywords, or image feature parameters. Using the image category tags (including target object type and target scene type) stored in the image-text index structure, image groups matching the object features in the retrieval request are searched. For example, if the object specified in the retrieval request is "person," then image groups tagged "person" are searched in the image-text index structure. Further retrieval matching is performed on the matched image groups, combining image feature parameters (such as target feature parameters) and semantic vectors of text content to determine the images and text content most relevant to the retrieval request. The final target retrieval information includes the matched images and their related text descriptions.

[0079] If the object existence determination result is negative, the scene existence determination module is triggered again. If a high-confidence descriptive word matching the "target scene type" label appears in the search vector, it is considered that "the target search request exists in a scene". Then, the scene-image-text coupled search channel is used. In the image-text index structure, the image-text joint traversal is performed with the scene type label as the anchor point. The image nodes and corresponding text nodes that meet the scene semantic constraints are selected by joint sorting through page proximity, semantic vector angle and index confidence. The structured target search information is also returned.

[0080] If neither of the two-level judgments is triggered, the remaining keywords in the search vector will be used as the text priority search conditions. The full-text inverted index and semantic vector recall module will be called to perform supplementary matching on the content processing results, so as to ensure that the most relevant search feedback can still be given even in the absence of clear object or scene identifiers.

[0081] The aforementioned branching retrieval strategy is supported by a unified image and text index structure. It performs dynamic matching based on object priority, scenario priority, or text priority for different query intentions, which not only ensures the accuracy and interpretability of the retrieval results, but also avoids the semantic ambiguity and insufficient recall problems of traditional single-channel retrieval under multimodal requests.

[0082] Furthermore, after determining whether the target retrieval request exists, step P30 in this embodiment of the application further includes:

[0083] P35: If it does not exist, perform semantic analysis on the target retrieval request according to the text strategy to obtain the retrieval semantic vector; P36: In conjunction with the image and text index structure, perform retrieval matching on the retrieval semantic vector in the PDF document to obtain the target retrieval information.

[0084] Specifically, to achieve a comprehensive response to multi-type target retrieval requests, a supplementary retrieval path based on semantic vectors can be further introduced to handle natural language queries that do not contain a clear object type or a scene description in the retrieval request.

[0085] Specifically, after determining whether the target retrieval request contains a scene, if it is determined that the retrieval request does not contain either the target object type label that can be mapped to the image and text index structure, or the target scene type label, then a text-based semantic analysis mechanism is triggered to perform semantic feature extraction and embedding modeling operations on the original retrieval request.

[0086] In this process, natural language processing techniques are used to perform semantic analysis on the text content of the retrieval request, extracting key semantic information and converting it into a semantic vector. This process can be achieved using pre-trained language models (such as BERT, Word2Vec, etc.), mapping the text in the retrieval request to a high-dimensional semantic space to obtain the retrieval semantic vector. This semantic vector retains the core semantic intent and logical structure features of the query statement, making it suitable for alignment with the vector space in the text and image index structure.

[0087] Subsequently, the retrieved semantic vector is used as the main query vector, and semantic similarity matching is performed across the entire PDF document, in conjunction with the semantic space information in the image-text index structure. For example, all text and image nodes in the image-text index structure are traversed, and their pre-stored or online-generated semantic vectors are compared with the retrieved semantic vector using similarity calculations (such as cosine similarity, Euclidean distance, or Mahalanobis distance). The results are then sorted according to the matching scores, and the image-text content fragments most semantically similar to the query are selected. If some image nodes in the image-text index structure are not bound to text nodes, their semantic content can be completed through image-text alignment paths, achieving unsupervised semantic completion matching.

[0088] Finally, results with similarity thresholds are packaged into target retrieval information output, including the image ID and page position of the matching image, the paragraph number and original content of the matching text segment, the similarity score, and its contextual path in the document structure. Multiple high-similarity candidate paths can be retained simultaneously for user selection, supporting multi-round interactive iterative query optimization. This process fully utilizes the image category labels, feature parameters, and semantic vectors of the text content stored in the image-text index structure. Through precise matching operations, it can quickly locate the images and text content most relevant to the user's search request. The resulting target retrieval information not only meets the user's search needs but also improves the accuracy and efficiency of the search, providing efficient support for intelligent image and content association retrieval in PDF documents.

[0089] In summary, the embodiments of this application have at least the following technical effects:

[0090] This application constructs an image-text index structure by performing multi-dimensional feature extraction and semantic association analysis on the image and text content in PDF documents. This achieves deep semantic linkage between images and text, significantly improving the retrieval accuracy and matching capability of image-text information. Target retrieval request matching based on this index structure has efficient response capabilities, avoiding the resource waste caused by full-text traversal. Simultaneously, this method supports the identification and branch retrieval path selection of multiple types of target requests, including images, text, and scenes, enhancing the system's adaptability in complex query contexts. Discrete cosine transform, multi-dimensional difference analysis, and variation weighting strategies are introduced during image processing to further improve image feature discrimination and index classification accuracy. The overall process achieves standardized and automated processing of image-text information, enabling the construction of high-quality image-text correspondences without manual annotation, demonstrating good adaptability, scalability, and engineering practical value.

[0091] The technology achieves the goal of intelligently associating and retrieving images and text content by constructing an image-text index structure, thereby improving retrieval efficiency and accuracy.

[0092] Example 2, based on the same inventive concept as the intelligent retrieval method for associating images and content in PDF documents as described in the previous examples, such as... Figure 2 As shown, this application provides an intelligent retrieval system for associating images and content in PDF documents. The system and method embodiments in this application are based on the same inventive concept. The system includes:

[0093] The document preprocessing module 11 is used to retrieve a document preprocessing strategy to preprocess the PDF document and obtain the document processing result, wherein the document processing result includes image processing result and content processing result.

[0094] Image-text association analysis module 12 is used to perform association analysis on the image processing results and the content processing results according to the association indexing mechanism to obtain the image-text index structure.

[0095] The retrieval matching module 13 is used to perform retrieval matching on the target retrieval request of the PDF document based on the image and text index structure, so as to obtain the target retrieval information.

[0096] Furthermore, the document preprocessing module 11 is also used to perform the following steps:

[0097] The PDF document is analyzed and extracted using object structure as a discrimination constraint to obtain extraction results; the image objects in the extraction results are processed according to the image strategy in the document preprocessing strategy to obtain the image processing results; wherein, the process includes: extracting a first object and a second object from the image objects; performing image standardization processing on the first object and the second object sequentially according to the image standardization scheme in the image strategy to obtain a first standard image and a second standard image respectively; sequentially calculating the first hash value of the first standard image and the second hash value of the second standard image; comparing the first hash value and the second hash value to obtain the Hamming distance; if the Hamming distance does not reach a predetermined threshold, the first object and the second object are separated to obtain the image processing results.

[0098] Furthermore, the document preprocessing module 11 is also used to perform the following steps:

[0099] The text objects in the extracted results are processed according to the text strategy in the document preprocessing strategy to obtain the content processing result; wherein, the process includes: extracting a third object from the text object; performing semantic analysis on the third object according to the text strategy to obtain semantic features; constructing a semantic vector based on the semantic features and combining it with the third object to obtain a mapping relationship; and forming the content processing result based on the mapping relationship.

[0100] Furthermore, the image-text association analysis module 12 is also used to perform the following steps:

[0101] Obtain any image group from the image processing results; perform multi-dimensional feature collection on the first arbitrary image in the arbitrary image group to obtain a first arbitrary feature parameter; analyze the first arbitrary feature parameter to determine the target feature parameter; activate the label classifier in the association indexing mechanism to perform classification analysis on the target feature parameter to obtain an arbitrary category label for the arbitrary image group; traverse the content processing results with the arbitrary category label as a traversal constraint to obtain the traversal result; establish the image and text index structure based on the traversal result.

[0102] Furthermore, the image-text association analysis module 12 is also used to perform the following steps:

[0103] Multi-dimensional feature collection yields a second arbitrary feature parameter for the second arbitrary image in the arbitrary image group; analysis of the first arbitrary feature parameter and the second arbitrary feature parameter determines the target feature parameter; wherein, before analyzing the first arbitrary feature parameter and the second arbitrary feature parameter to determine the target feature parameter, the process includes: sequentially performing discrete cosine transform on the first arbitrary image and the second arbitrary image to obtain a first transform coefficient and a second transform coefficient, respectively; constructing the first arbitrary feature parameter based on the first DC coefficient and the first AC coefficient in the first transform coefficient; and constructing the second arbitrary feature parameter based on the second DC coefficient and the second AC coefficient in the second transform coefficient.

[0104] Furthermore, the image-text association analysis module 12 is also used to perform the following steps:

[0105] Obtain the first difference between the first DC coefficient and the second DC coefficient; obtain the second difference between the first AC coefficient and the second AC coefficient; perform a variation weighted calculation on the first difference and the second difference to obtain the first deviation index; optimize the target pair with the first deviation index as the target; take the mean of the first target image group and the second target image group in the target pair to form the target feature parameters.

[0106] Furthermore, the image-text association analysis module 12 is also used to perform the following steps:

[0107] The target DC coefficient is classified and analyzed by the label classifier to obtain the target object type; the target AC coefficient is classified and analyzed by the label classifier to obtain the target scene type; the target object type and the target scene type constitute the arbitrary category label.

[0108] Furthermore, the retrieval and matching module 13 is also used to perform the following steps:

[0109] Determine whether the target retrieval request contains an object; if it does, coordinate with the image and text index structure to search and match the object in the PDF document to obtain the target retrieval information; if it does not, determine whether the target retrieval request contains a scene; if it does, coordinate with the image and text index structure to search and match the scene in the PDF document to obtain the target retrieval information.

[0110] Furthermore, the retrieval and matching module 13 is also used to perform the following steps:

[0111] If it does not exist, semantic analysis is performed on the target retrieval request according to the text strategy to obtain a retrieval semantic vector; the retrieval semantic vector is then matched with the image and text index structure in the PDF document to obtain the target retrieval information.

[0112] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.

[0113] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

[0114] This specification and accompanying drawings are merely illustrative examples of this application and are intended to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from its scope. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application intends to include such modifications and variations.

Claims

1. An intelligent retrieval method for associating images with content in PDF documents, characterized in that, include: The document preprocessing strategy is invoked to preprocess the PDF document to obtain the document processing result, wherein the document processing result includes image processing result and content processing result; Based on the association indexing mechanism, the image processing results and the content processing results are analyzed to obtain the image and text index structure; Based on the aforementioned image and text index structure, the target retrieval request of the PDF document is searched and matched to obtain the target retrieval information; Specifically, the image processing results and content processing results are correlated using an association indexing mechanism to obtain an image-text index structure, including: Obtain any group of images from the image processing results; Multi-dimensional feature collection is performed on the first arbitrary image in the arbitrary image group to obtain the first arbitrary feature parameters; Analyze the first arbitrary feature parameters to determine the target feature parameters; Activate the label classifier in the association index mechanism to classify and analyze the target feature parameters to obtain arbitrary category labels for the arbitrary image group; The content processing result is traversed using the arbitrary category label as a traversal constraint to obtain the traversal result. The image and text index structure is established based on the traversal results; The step of analyzing the first arbitrary feature parameter to determine the target feature parameter includes: Multi-dimensional feature collection yields the second arbitrary feature parameters of the second arbitrary image in the arbitrary image group; Analyze the first arbitrary feature parameter and the second arbitrary feature parameter to determine the target feature parameter; Before analyzing the first arbitrary feature parameter and the second arbitrary feature parameter to determine the target feature parameter, the process includes: The first arbitrary image and the second arbitrary image are sequentially subjected to discrete cosine transform to obtain the first transform coefficient and the second transform coefficient, respectively. Based on the first DC coefficient and the first AC coefficient in the first transformation coefficient, the first arbitrary characteristic parameter is constructed; Based on the second DC coefficient and the second AC coefficient in the second transformation coefficient, the second arbitrary characteristic parameter is constructed; The process of analyzing the first arbitrary feature parameter and the second arbitrary feature parameter to determine the target feature parameter includes: Obtain the first difference between the first DC coefficient and the second DC coefficient; Obtain the second difference between the first AC coefficient and the second AC coefficient; A first deviation index is obtained by performing a variation-weighted calculation on the first difference and the second difference; The target pair is obtained by optimizing the first deviation index to maximize it. The average values ​​of the first target image group and the second target image group in the target pair are used to form the target feature parameters; Specifically, activating the label classifier in the association indexing mechanism to classify and analyze the target feature parameters to obtain arbitrary category labels for the arbitrary image group includes: The target DC coefficient is classified and analyzed by the label classifier to obtain the target object type; The target scene type is obtained by classifying and analyzing the target communication coefficients using the label classifier. The target object type and the target scene type together form the arbitrary category label.

2. The intelligent retrieval method for associating images and content in a PDF document as described in claim 1, characterized in that, The document preprocessing strategy is invoked to preprocess the PDF document, and the document processing results are obtained, including: The PDF document is analyzed and extracted using the object structure as a discrimination constraint to obtain the extraction result; The image objects in the extracted results are processed according to the image strategy in the document preprocessing strategy to obtain the image processing result; This includes: Extract the first object and the second object from the image object; According to the image standardization scheme in the image strategy, the first object and the second object are standardized sequentially to obtain the first standard image and the second standard image, respectively. The first hash value of the first standard image and the second hash value of the second standard image are calculated sequentially. The Hamming distance is obtained by comparing the first hash value with the second hash value; If the Hamming distance does not reach the predetermined threshold, the first object and the second object are separated to obtain the image processing result.

3. The intelligent retrieval method for associating images and content in a PDF document as described in claim 2, characterized in that, The document preprocessing strategy is invoked to preprocess the PDF document, and the document processing results are obtained, including: The text objects in the extracted results are processed according to the text strategy in the document preprocessing strategy to obtain the content processing result; This includes: Extract the third object from the text object; Semantic analysis is performed on the third object according to the text strategy to obtain semantic features; A semantic vector is constructed based on the semantic features, and a mapping relationship is obtained by combining it with the third object; The content processing result is formed based on the mapping relationship.

4. The intelligent retrieval method for associating images and content in a PDF document as described in claim 3, characterized in that, Based on the aforementioned image and text index structure, the target retrieval request for the PDF document is searched and matched to obtain target retrieval information, including: Determine whether the target retrieval request contains an object; If it exists, the object is searched and matched in the PDF document using the image and text index structure to obtain the target search information; If it does not exist, determine whether the target retrieval request exists in the scenario; If it exists, the scene is searched and matched in the PDF document using the image and text index structure to obtain the target search information.

5. The intelligent retrieval method for associating images and content in a PDF document as described in claim 4, characterized in that, After determining whether the target retrieval request exists, the process also includes: If it does not exist, perform semantic analysis on the target retrieval request according to the text strategy to obtain a retrieval semantic vector; The image and text index structure is used in conjunction with the search semantic vector in the PDF document to obtain the target search information.

6. An intelligent retrieval system that associates images with content in PDF documents, characterized in that, The system is used to perform the intelligent retrieval method for associating images and content in a PDF document as described in any one of claims 1 to 5, the system comprising: A document preprocessing module is used to retrieve a document preprocessing strategy to preprocess the PDF document and obtain the document processing result, wherein the document processing result includes image processing result and content processing result; The image-text association analysis module is used to perform association analysis on the image processing results and the content processing results according to the association indexing mechanism to obtain the image-text index structure. The retrieval and matching module is used to perform retrieval and matching on the target retrieval request of the PDF document based on the image and text index structure, so as to obtain the target retrieval information.

Citation Information

Patent Citations

  • PDF image screening method driven by intelligent analysis

    CN119441531A