Biding document multi-mode duplicate checking method and system based on large model
Through the multimodal plagiarism check method based on the large language model, dynamic context perception and hierarchical position coding, combined with cross-modal attention calculation and distributed computing framework, the problems of insufficient semantic understanding of bid plagiarism check, poor field adaptability and low computing efficiency are solved, and efficient and real-time plagiarism check for professional term intensive bids are achieved.
Patent Information
- Application Number
- CN202510787107.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-13
AI Technical Summary
The existing bid plagiarism checking technology has insufficient semantic understanding, poor field adaptability and low computational efficiency, and is particularly difficult to deal with the content of bid intensive professional terms and complex structures, and lacks the fusion analysis of multimodal data.
A multimodal plagiarism check method based on large language models is adopted, and deep semantic plagiarism check is achieved through dynamic context perception and hierarchical position coding. Combining cross-modal attention calculation and distributed computing framework, structured feature reconstruction and correlation analysis are carried out for non-text content such as tables, charts, and formulas.
It improves the accuracy of dilution checking, significantly reduces the cost of cross-field migration, meets the real-time response needs under 100 million data volumes, solves the problem of missed detection of non-text elements by traditional methods, and improves the accuracy of dilution checking of intensive long texts in professional terms.
Smart Images

Figure CN120337898A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing and information retrieval, and particularly relates to a multi-modal bid document duplicate checking method and system based on a large model. Background Art
[0002] The statements in this part only provide background art related to the present invention and do not necessarily constitute prior art.
[0003] Current bid document duplicate checking technologies mainly rely on two types of methods: rule matching and shallow semantic analysis. Rule matching technology checks for duplicates through preset text patterns or fragment overlap. Although it is simple to implement, it is difficult to handle semantic repetition problems with different expressions, such as being unable to recognize content with synonym replacement, sentence pattern adjustment, or logical restructuring, and the construction and maintenance costs of the rule library are relatively high. Methods based on shallow semantic analysis use statistical machine learning models to calculate text similarity through word frequency or shallow semantic features. Although it alleviates the problem of semantic generalization to a certain extent, due to the limited model representation ability, the accuracy of capturing complex logical relationships in long texts is insufficient. Especially when facing bid document content with dense professional terms and complex structures, the duplicate checking accuracy drops significantly.
[0004] The limitations of the prior art are mainly reflected in three aspects: First, the depth of semantic understanding is insufficient. Traditional models lack the ability to dynamically learn context associations and domain knowledge, resulting in a high missed detection rate for content that is semantically equivalent but differently expressed. Second, the domain adaptability is poor. The professional terms and specifications involved in bid documents in different industries vary greatly, and existing methods need to be optimized separately for each domain, with high implementation costs. Third, the computational efficiency is limited. Especially when processing large-scale bid document libraries, traditional algorithms are difficult to meet real-time requirements due to the complexity of high-dimensional sparse vector calculations and long text segmentation processing.
[0005] In recent years, with the popularization of pre-trained language models, some improved solutions have tried to introduce general semantic models to improve the duplicate checking accuracy, but such solutions still have significant deficiencies. For example, they do not design differentiated processing strategies for the unique structured content (such as technical solutions, quotation lists) of bid documents, resulting in poor duplicate checking effects for non-text elements such as tables and clauses; at the same time, directly using unoptimized general models is prone to response delay problems and lacks support for dynamic updates of industry terms; in addition, existing technologies generally ignore the fusion analysis of multi-modal data, such as the joint duplicate checking of charts, formulas, and text, further restricting the expansion of application scenarios. Summary of the Invention
[0006] To address the deficiencies of the prior art, the present invention provides a multi-modal bid duplicate checking method and system based on a large model. By leveraging the context awareness ability of the large language model, it realizes in-depth semantic duplicate checking, accurately identifies semantic equivalence between texts, reconstructs the structured features of non-text content such as tables, charts, and formulas, achieves cross-modal semantic correlation analysis, and improves the accuracy of duplicate checking.
[0007] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a multi-modal bid duplicate checking method based on a large model.
[0008] A multi-modal bid duplicate checking method based on a large model includes the following processes: Preprocess the original bid to obtain text chunks with hierarchical labels and multi-modal features of non-text; Based on a large language model with dynamic context awareness, perform semantic encoding on the text chunks to obtain semantic vectors corresponding to the text chunks; Construct an industry-segmented hierarchical index based on the semantic vectors, dynamically adjust the hierarchical structure of the index, and determine the semantic similarity according to the hierarchical structure of the index; According to the semantic vectors of the original bid, the multi-modal features of non-text, and the multi-modal features of the candidate document, perform cross-modal attention calculation to determine the weighting coefficients of each modality, and dynamically allocate the weights of each modality feature in the candidate document according to the proportion of each modality feature in the original bid; According to the weighting coefficients, the weights, and the special matching degrees of various modalities, determine the non-text similarity, and determine the global similarity between the original bid and the candidate document according to the weighted sum of the semantic similarity and the non-text similarity.
[0009] In an implementation manner of the first aspect of the present invention, after preprocessing the original bid, multi-modal associated metadata is obtained, and the semantic vectors are associated with the multi-modal associated metadata; Extract the keyword set of the original bid, and determine multiple candidate documents through correlation calculation. This process is distributed to multiple computing nodes for parallel execution; Calculate the semantic similarity between the semantic vectors of each candidate document and the query vector of the original bid. This process performs parallel vector operations through the SIMD instruction set.
[0010] In an implementation manner of the first aspect of the present invention, preprocessing the original bid to obtain text chunks with hierarchical labels includes: Match the title based on regular expressions, and calculate the weight of the title by comprehensively considering the font size, bold attribute, and position information. The weight is dynamically determined by the linear combination of the normalized font size value, bold flag, and position weight. When the weight exceeds the preset threshold, it is marked as a title and a hierarchical index tree is generated, and the original tender document is split into independent text blocks with hierarchical labels.
[0011] In one implementation of the first aspect of the present invention, the original tender document is preprocessed to obtain multi-modal features, including: converting the table content into a logical tree structure; Locate the axes, legends, and data points in the chart through a target detection model, extract the numerical range vector, and construct a topological relationship matrix of the data points based on a graph neural network; Convert the formula obtained by the formula editor or the image formula into a symbol tree structure, generate a serialized expression through breadth-first traversal, and generate a hash fingerprint using locality-sensitive hashing; The logical tree structure, the topological relationship matrix, and the hash fingerprint constitute the multi-modal features.
[0012] In one implementation of the first aspect of the present invention, for any text block, it is jointly represented by word embedding, paragraph type embedding, and hierarchical position encoding. The hierarchical position encoding dynamically adjusts the encoding frequency according to the title level. Before the joint representation result is input into the large language model, a domain adaptation unit is used to dynamically splice the industry-specific prompt template.
[0013] In one implementation of the first aspect of the present invention, the hierarchical structure of the index is dynamically adjusted, including: for the sub-library of the industry, the number of index levels is adaptively calculated according to the data scale : , where is the hierarchical sparsity factor.
[0014] In one implementation of the first aspect of the present invention, cross-modal attention calculation is performed, including: , where is the semantic vector corresponding to the th text block of the original tender document , is the multi-modal fusion function of the candidate document, represents the th text block of the candidate document corresponding semantic vector, represents the th text block of the candidate document corresponding non-text multi-modal feature, are all trainable projection matrices, is the vector dimension.
[0015] In a second aspect, the present invention provides a multi-modal bid duplicate checking system based on a large model.
[0016] A multi-modal bid duplicate checking system based on a large model includes: A preprocessing unit configured to: preprocess the original bid to obtain text chunks with hierarchical labels and multi-modal features of non-text; A semantic processing unit configured to: perform semantic encoding on the text chunks based on a large language model with dynamic context awareness to obtain semantic vectors corresponding to the text chunks; A retrieval calculation unit configured to: construct an industry-segmented hierarchical index based on the semantic vectors, dynamically adjust the hierarchical structure of the index, and determine the semantic similarity according to the hierarchical structure of the index; A multi-modal calculation unit configured to: perform cross-modal attention calculation according to the semantic vectors of the original bid, the multi-modal features of non-text, and the multi-modal features of the candidate document to determine the weighting coefficients of each modality, and dynamically allocate the weights of each modality feature in the candidate document according to the proportion of each modality feature in the original bid; A joint duplicate checking unit configured to: determine the non-text similarity according to the weighting coefficients, the weights, and the special matching degrees of various modalities, and determine the global similarity between the original bid and the candidate document according to the weighted sum of the semantic similarity and the non-text similarity.
[0017] In a third aspect, the present invention provides a computer device, including: a processor and a computer-readable storage medium; The processor is adapted to execute a computer program; The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the multi-modal bid duplicate checking method based on a large model as described in the first aspect of the present invention.
[0018] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the multi-modal bid duplicate checking method based on a large model as described in the first aspect of the present invention.
[0019] Compared with the prior art, the beneficial effects of the present invention are: 1. The present invention innovatively proposes a bid duplicate checking method based on large models, which utilizes the context awareness ability of large language models to achieve in-depth semantic duplicate checking, accurately identify the semantic equivalence between texts, reconstruct the structured features of non-text content such as tables and charts, realize cross-modal semantic correlation analysis, and construct an intelligent duplicate checking strategy covering text and non-text content by integrating dynamic context awareness encoding, domain adaptation optimization and multi-modal joint analysis techniques.
[0020] 2. The present invention designs a multi-level collaborative optimization mechanism for the limitations of traditional technologies. At the semantic understanding level, the semantic break problem of long texts is solved by sliding window and hierarchical position encoding techniques. For example, multi-level headings in bids are mapped to hierarchical vectors to capture the context dependencies across paragraphs and chapters, so as to accurately identify semantic repetitions caused by sentence restructuring or logical adjustment. At the domain adaptation level, pluggable parameter adapters and incremental fine-tuning frameworks are introduced to support the dynamic loading of term libraries and bid specifications in different industries. Only a small number of industry samples are required to update the adapter parameters, significantly reducing the cross-domain migration cost. At the computing architecture level, a hybrid pipeline combining inverted index and dense vector retrieval is proposed. In the coarse screening stage, the candidate set range is quickly narrowed based on keywords and metadata. In the fine ranking stage, accurate ranking is achieved through semantic vector similarity calculation. At the same time, distributed task scheduling and SIMD instruction set parallelization are used to accelerate the operation, ensuring real-time response under the data volume of hundreds of millions.
[0021] 3. The present invention breaks through the processing bottleneck of non-text elements in traditional methods and proposes a multi-modal joint duplicate checking scheme. For tables, charts and formulas in bids, structured feature extraction methods are designed respectively: table content is parsed into a logical tree and row-column relationship features are extracted, charts match topological similarity through graph neural networks, and formulas are converted into symbol trees to generate hash fingerprints. The text and non-text features are fused through cross-modal attention mechanisms to realize the joint analysis of graphic and text semantics. For example, the text description in the technical solution is associated and matched with the table data in the quotation sheet to avoid missed detection problems caused by scattered content.
[0022] 4. The present invention has achieved remarkable breakthroughs in semantic understanding, domain adaptation, and multimodal processing. Through the dynamic context encoding mechanism, the system can accurately identify semantic duplicate content in bids caused by sentence restructuring, synonym replacement, or logical adjustment. In particular, the accuracy of duplicate checking for long texts with a high density of professional terms has been significantly improved. The introduction of the parameter adapter and the incremental fine-tuning framework has greatly shortened the migration and deployment cycle of bid specifications across industries, while avoiding the resource consumption of full-model fine-tuning. The combination of the hybrid retrieval architecture and the distributed optimization strategy ensures that the retrieval response time for a bid library of hundreds of millions of entries is reduced to the millisecond level, meeting the real-time requirements of the bidding scenario. For non-text content, the multimodal joint duplicate checking scheme effectively solves the problem of missed detection of elements such as tables and charts by traditional methods through structured feature extraction and cross-modal correlation analysis. For example, semantic integrity verification of merged cells, quantitative matching of chart topology similarity, and fuzzy hashing comparison of formula symbol trees are carried out, thus forming a duplicate checking ability covering all elements.
[0023] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0025] Figure 1 A schematic flowchart of a bid duplicate checking method based on a large model provided for an exemplary embodiment of the present invention; Figure 2 A schematic diagram of a bid duplicate checking system based on a large model provided for an exemplary embodiment of the present invention; Figure 3 A schematic diagram of a computer device provided for an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0026] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0027] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0028] As described in the background art, in the prior art, bid duplicate checking mainly relies on traditional text matching algorithms (such as cosine similarity, TF-IDF) or shallow semantic analysis models, and there are core problems such as insufficient semantic understanding, poor domain adaptability, and low efficiency in processing massive data. For example, traditional methods are difficult to identify content with the same semantics but different expressions (such as synonym replacement, sentence pattern restructuring), and have weak context correlation analysis capabilities for professional terms and industry specifications, resulting in a high missed detection rate and limited cross-scenario applicability. In addition, when processing large-scale bid data, traditional algorithms are difficult to achieve real-time response due to high computational complexity. In view of this, this implementation method proposes a bid duplicate checking method based on a large model, which performs structured parsing and multi-modal feature extraction on the input bid to generate feature representations of semantic chunks and non-text elements; uses a large language model with dynamic context awareness to perform deep semantic encoding on the text chunks, and combines hierarchical position encoding to retain the logical relevance of long texts; realizes efficient matching of massive semantic vectors through a hybrid retrieval architecture, and combines a distributed computing framework and hardware acceleration instructions to optimize the computing efficiency; finally, reconstructs the structured features of non-text content such as tables and charts to achieve cross-modal semantic correlation analysis and output a comprehensive duplicate checking result.
[0029] More specifically, as Figure 1 shown, it includes the following processes: S101: Preprocess the original bid to obtain text chunks with hierarchical labels and multi-modal features of non-text; S102: Based on a large language model with dynamic context awareness, perform semantic encoding on the text chunks to obtain semantic vectors corresponding to the text chunks; S103: Build an industry-level hierarchical index based on the semantic vectors, dynamically adjust the hierarchical structure of the index, and determine the semantic similarity according to the hierarchical structure of the index; S104: According to the semantic vectors of the original bid, the multi-modal features of non-text, and the multi-modal features of the candidate document, perform cross-modal attention calculation to determine the weighting coefficients of each modality, and dynamically allocate the weights of the multi-modal features in the candidate document according to the proportion of each modality feature in the original bid; S105: Determine the non-text similarity according to the weighting coefficients, the weights, and the special matching degrees of various modalities, and determine the global similarity between the original bid and the candidate document according to the weighted sum of the semantic similarity and the non-text similarity.
[0030] In step S101 of this implementation method, more specifically, it includes: Through structured parsing and multi-modal feature extraction techniques, the original tender document is transformed into standardized semantic units and non-text feature representations, providing a basis for subsequent in-depth semantic analysis. The core tasks of preprocessing include tender-level splitting, non-text element separation, and multi-modal semantic reconstruction.
[0031] Structured parsing adopts a combination of a rule engine and an adaptive algorithm to identify multi-level headings and divide semantic blocks in the tender. Based on regular expressions to match headings (such as "1.1.1 Technical Parameters"), and comprehensively calculate the heading weight by combining the font size, bold attribute, and position information. The weight is dynamically determined by the linear combination of the normalized font size value, bold flag, and position weight. When the weight exceeds the preset threshold, it is marked as a heading and a hierarchical index tree is generated, thus splitting the tender into independent semantic blocks such as technical solutions and quotation sheets. For the problem of broken cross-page tables, an association algorithm based on merged cell features and header consistency is designed: by detecting the repeatability of header keywords and the continuity of numerical distributions, automatically splice cross-page tables to ensure logical integrity.
[0032] The separation of non-text content is achieved through a hybrid OCR engine, which combines a convolutional neural network (CNN) and a hidden Markov model (HMM) to improve the recognition accuracy of scanned documents, and uses a context semantic error correction mechanism to correct easily confused characters. For example, for the blurred characters output by OCR, dynamically optimize the candidate character selection based on the probability distribution of adjacent words to ensure the accurate recognition of characters such as "0" and "O".
[0033] More specifically, for tables, charts, and formulas, a differentiated feature extraction scheme is designed. Table parsing converts the table content into a logical tree structure, extracts row-column relationships, numerical distribution features, and merged cell ranges, generates a global feature vector, assigns higher weights to the header rows, and constructs a dense vector reflecting the table content by weighted aggregation of the semantic roles (such as "data item", "unit") and statistical features (mean, variance) of each cell; Chart processing locates the coordinate axes, legends, and data points through an object detection model, extracts the numerical range vector, and constructs the topological relationship of data points based on a graph neural network (GNN). The weights of the adjacency matrix are jointly determined by the spatial distance and numerical difference of data points to capture the structural similarity of the chart; Formula encoding converts LaTeX or image formulas into a symbol tree structure, generates a serialized expression through breadth-first traversal, and uses locality-sensitive hashing (LSH) to generate fuzzy fingerprints to support fuzzy matching of variable substitution and typesetting fine-tuning.
[0034] After preprocessing, text chunks with hierarchical labels, multi-modal feature sets (logical trees, topological relationship matrices, hash fingerprints), and associated metadata (such as "a certain figure corresponds to a certain section") are output to ensure that downstream strategies can integrate the combined semantic information of text and non-text.
[0035] In step S102 of this implementation method, more specifically, it includes: Based on a large language model with dynamic context awareness, perform deep semantic encoding on the preprocessed structured text chunks, and optimize cross-industry semantic representation through domain adaptation technology. Collaborate hierarchical position encoding with pluggable parameter adapters to solve the problems of long text semantic break and industry term generalization.
[0036] For the th text chunk (associated hierarchical label such as "2.3.1 Technical Solution"), it can be jointly represented by word embedding, paragraph type embedding, and hierarchical position encoding. Hierarchical position encoding dynamically adjusts the encoding frequency according to the title hierarchy (chapter, section, paragraph), and captures long-range dependencies through multi-level sine / cosine functions. Specifically, the th-level position encoding has the following calculation formula: (1); Wherein, is the vector dimension, is the maximum number of title levels, is the frequency adjustment factor. By superimposing the encoding results of different levels, the model can distinguish the global logic of chapters and the local details of paragraphs in the technical solution. For example, map the hierarchical relationship between "1.1 Overall Design" and "1.1.2 Core Parameters" to differential position features.
[0037] The domain adaptation unit uses lightweight parameter adapters to inject industry knowledge. The adapter is superimposed on the base LLM model in the form of residuals, and its core is a bottleneck structure of dimension reduction - activation - dimension increase, which significantly reduces the computational overhead. During training, only update the adapter parameters, and keep the base model parameters frozen to ensure the efficiency of industry migration. Dynamically splice industry-specific prompt templates before the input text to guide the model to focus on industry key terms and specification constraints.
[0038] The semantic vector set obtained by the present invention integrates multi-granularity semantic information: high-level encoding focuses on the framework consistency of the technical solution, and low-level encoding captures the differences in implementation details. At the same time, each semantic vector is associated with the multi-modal associated metadata provided by the preprocessing process (such as "a certain figure corresponds to a certain paragraph"), providing semantic alignment anchors for subsequent cross-modal joint duplicate checking. For example, the "equipment installation process" described in the technical solution will be cross-modally matched with the "equipment list table" in the quotation through association marks, avoiding missed detection caused by scattered content.
[0039] In step S103 of this implementation method, more specifically, it includes: Through a hybrid retrieval architecture and a distributed optimization strategy, efficient semantic matching and multimodal joint duplicate checking of a billion-level tender library are achieved. Specifically: First, a hierarchical index by industry is constructed based on the semantic vector set output during the semantic processing process. The improved HNSW (Hierarchical Navigable Small World) algorithm is used to dynamically adjust the hierarchical structure of the index. For the sub-library of an industry the number of index levels is adaptively calculated according to the data scale as: (2); where is the hierarchical sparsity factor (default value 500), and the number of levels dynamically expands as the number of tenders grows. The high-level index stores coarse-grained feature vectors to quickly narrow the search range, and the low-level index retains fine-grained vectors to improve the matching accuracy. Navigation between levels is performed through a greedy search algorithm, significantly reducing the retrieval latency under billion-level data. During index construction, each industry sub-library is stored independently, ensuring that only the target industry index needs to be loaded during cross-industry queries, reducing ineffective calculations.
[0040] The retrieval process is divided into a two-stage pipeline: The rough screening stage (i.e., the first screening stage) quickly screens the candidate document set based on the inverted index and metadata (such as industry classification, tender type), extracts the keyword set of the original tender (such as the core terms in the technical solution), calculates the document relevance score through TF-IDF weights, and screens out the Top-M candidate document set , and the rough screening score is calculated as: (3); where is the indicator function, screening documents with a score higher than the threshold , is the keyword, is the term frequency, is the inverse document frequency.
[0041] The fine ranking stage (i.e., the second screening stage) calculates the similarity between the semantic vector of the documents in the candidate set (derived from the document ) and the query vector (derived from the original tender), and performs joint ranking by fusing multimodal features. The semantic similarity (i.e., text similarity) adopts a weighted combination of cosine similarity and Manhattan distance: (4); where Adjust the contribution degree of semantic matching, is the distance scaling factor, optimized through cross-validation. This design takes into account both the direction similarity and absolute distance difference of vectors, and is especially suitable for scenarios where there are differences in expression but similar semantics in technical solutions.
[0042] To meet the real-time requirements under massive data, the present invention uses the MapReduce distributed framework to optimize the calculation process. The tender library is stored in different computing nodes by industry sharding, and the rough screening tasks are distributed to each node for parallel execution; in the refined ranking stage, vector operations are parallelized through the SIMD instruction set to accelerate the calculation of cosine similarity and distance. The dynamic resource scheduler allocates GPU / CPU resources in real time according to the query load, preferentially processes high-priority tasks, and automatically expands computing nodes in high-concurrency scenarios.
[0043] In steps S104 and S105 of this implementation method, specifically, it includes: Through the cross-modal attention mechanism and dynamic feature fusion strategy, deep semantic association analysis of text and non-text elements is realized, solving the problem of missed detection of text-image separated content by traditional methods. Based on the multi-modal feature set (table logic tree, chart topology matrix, hash fingerprint of formula) extracted by preprocessing and the semantic vectors generated by semantic processing, a unified cross-modal semantic space is constructed, and the global similarity is quantified through adaptive weight allocation.
[0044] In the cross-modal attention mechanism, multi-modal feature interaction is realized through the query-key-value (QKV) model. For the text chunks of the original tender and their associated non-text features (such as tables ), calculate their cross-modal attention scores with the candidate tender : (5); Among them, is the semantic vector corresponding to the th text chunk of the original tender, is the multi-modal fusion function of the candidate document, represents the th text chunk of the candidate document corresponding semantic vector, represents the th non-text multi-modal feature corresponding to the text chunk of the candidate document, is the vector dimension, is the multi-modal fusion function of the candidate document. The fusion function Map the text vector and non-text features to the same space, and dynamically weight the contributions of each modality through an independent encoding network to automatically align the text description with the associated non-text elements (such as "equipment parameters in the technical solution" and "model list in the quotation form").
[0045] The dynamic feature fusion generates a comprehensive similarity based on the attention scores. For the candidate documents , its global similarity is calculated by weighting the text similarity and the non-text similarity : (6); (7); Among them, is the dynamic weight coefficient, which is adaptively adjusted according to the modality distribution of the query content. The non-text similarity is further decomposed into the contributions of tables, charts, and formulas, and the weights are dynamically allocated according to the proportion of non-text elements in the original tender. For example, if the original tender contains multiple tables, the weight of the table similarity is significantly increased to ensure the adaptability of the business scenario. is the special matching similarity of the th non-text modality. Table duplicate checking is based on the logical tree feature vector and the intersection over union (IoU) of row-column relationships, and combines the numerical distribution difference (such as KL divergence) to quantify the similarity; chart duplicate checking calculates the structural similarity of the adjacency matrix and the similarity of node feature matching through the topological features extracted by the graph neural network; formula duplicate checking uses local sensitive hashing fingerprints and symbolic tree edit distances to quantify the similarity speed, supporting fuzzy matching with variable substitution and layout fine-tuning. is the weighted coefficient of the cross-modal attention for the th non-text modality, which is obtained by normalizing the weight distribution output by the attention mechanism. For example, when the attention score between the query text and the candidate table is high, will be close to 1, increasing the proportion of the table similarity in the non-text fusion, represents the weighted coefficient corresponding to the table. After all the non-text similarity results are weighted and fused, they jointly generate the final duplicate checking score with the text similarity.
[0046] Training and optimization adopt a contrastive learning framework, and the model is driven to learn the consistent expression of text and image semantics through the triplet loss function. For example, through a large number of sample trainings, the model can automatically associate the text description of the "construction flow chart" in the technical solution with the topological features of the corresponding chart, avoiding the missed detection caused by the separation of text and image in traditional methods.
[0047] Figure 2 Figure shows a multi-modal tender duplicate checking system based on a large model, including: The preprocessing unit 201 is configured to: preprocess the original tender document to obtain text chunks with hierarchical tags and non-text multi-modal features; The semantic processing unit 202 is configured to: perform semantic encoding on the text chunks based on a large language model with dynamic context awareness to obtain semantic vectors corresponding to the text chunks; The retrieval and calculation unit 203 is configured to: construct an industry-segmented hierarchical index based on the semantic vectors, dynamically adjust the hierarchical structure of the index, and determine the semantic similarity according to the hierarchical structure of the index; The multi-modal calculation unit 204 is configured to: perform cross-modal attention calculation based on the semantic vectors of the original tender document, non-text multi-modal features, and multi-modal features of the candidate document to determine the weighting coefficients of each modality, and dynamically allocate the weights of each modality feature in the candidate document according to the proportion of each modality feature in the original tender document; The combined duplicate-checking unit 205 is configured to: determine the non-text similarity according to the weighting coefficients, the weights, and the special matching degrees of various modalities, and determine the global similarity between the original tender document and the candidate document according to the weighted sum of the semantic similarity and the non-text similarity.
[0048] It can be understood that the above-mentioned units can be respectively or all combined into one or several other units to form, or some of them can be further split into multiple smaller units with more specific functions to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present invention. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present invention, the system may also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.
[0049] According to another embodiment of the present invention, the system described in this embodiment can be constructed by running a computer program (including program code) capable of executing the respective steps involved in the corresponding method described in Embodiment 1 on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a random access memory (RAM), and a read-only memory (ROM). The computer program can be recorded on a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0050] Figure 3A computer device is shown. The electronic device includes a processor 301, a communication interface 302, and a computer-readable storage medium 303. Among them, the processor 301, the communication interface 302, and the computer-readable storage medium 303 can be connected through a bus or other means.
[0051] Among them, the communication interface 302 is used to receive and send data. The computer-readable storage medium 303 can be stored in the memory of the electronic device. The computer-readable storage medium 303 is used to store computer programs. The computer programs include program instructions. The processor 301 is used to execute the program instructions stored in the computer-readable storage medium 303.
[0052] The processor 301 is the computing core and control core of the electronic device. It is suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.
[0053] The processor 301 is configured to execute the following process: Preprocess the original tender document to obtain text chunks with hierarchical tags and non-text multi-modal features; Based on a large language model with dynamic context awareness, semantically encode the text chunks to obtain semantic vectors corresponding to the text chunks; Build an industry-specific hierarchical index based on the semantic vectors, dynamically adjust the hierarchical structure of the index, and determine the semantic similarity according to the hierarchical structure of the index; According to the semantic vectors of the original tender document, non-text multi-modal features, and multi-modal features of the candidate documents, perform cross-modal attention calculation to determine the weighting coefficients of each modality, and dynamically allocate the weights of each modality feature in the candidate documents according to the proportion of each modality feature in the original tender document; Determine the non-text similarity according to the weighting coefficients, the weights, and the special matching degrees of various modalities, and determine the global similarity between the original tender document and the candidate document according to the weighted sum of the semantic similarity and the non-text similarity.
[0054] The present invention also provides a computer-readable storage medium. The computer-readable storage medium is a memory device in the electronic device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the electronic device.
[0055] Moreover, one or more instructions suitable for being loaded and executed by a processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0056] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the processor loads and executes one or more instructions stored in the computer-readable storage medium to implement the following process: Preprocess the original tender document to obtain text chunks with hierarchical tags and multi-modal features of non-text; Based on a large language model with dynamic context awareness, semantically encode the text chunks to obtain semantic vectors corresponding to the text chunks; Build an industry-segmented hierarchical index based on the semantic vectors, dynamically adjust the hierarchical structure of the index, and determine semantic similarity according to the hierarchical structure of the index; Perform cross-modal attention calculation according to the semantic vectors of the original tender document, multi-modal features of non-text, and multi-modal features of candidate documents to determine the weighting coefficients of each modality, and dynamically allocate the weights of each modality feature in the candidate document according to the proportion of each modality feature in the original tender document; Determine non-text similarity according to the weighting coefficients, the weights, and the special matching degrees of various modalities, and determine the global similarity between the original tender document and the candidate document according to the weighted sum of the semantic similarity and the non-text similarity.
[0057] The present invention also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the electronic device to perform the following process: Preprocess the original tender document to obtain text chunks with hierarchical tags and multi-modal features of non-text; Based on a large language model with dynamic context awareness, semantically encode the text chunks to obtain semantic vectors corresponding to the text chunks; Build an industry-segmented hierarchical index based on the semantic vectors, dynamically adjust the hierarchical structure of the index, and determine semantic similarity according to the hierarchical structure of the index; Based on the semantic vectors of the original tender, the multi-modal features of non-text, and the multi-modal features of the candidate document, cross-modal attention calculation is performed to determine the weighting coefficients of each modality, and the weights of each modality feature in the candidate document are dynamically allocated according to the proportion of each modality feature in the original tender; According to the weighting coefficients, the weights, and the special matching degrees of various modalities, the non-text similarity is determined, and according to the weighted sum of the semantic similarity and the non-text similarity, the global similarity between the original tender and the candidate document is determined.
[0058] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0059] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data processing device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0060] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A multi-modal bid duplicate checking method based on large models, characterized in that, It includes the following processes: Preprocess the original tender document to obtain text chunks with hierarchical tags and non-text multimodal features; Based on a large language model with dynamic context awareness, semantically encode the text chunks to obtain semantic vectors corresponding to the text chunks; Construct an industry-segmented hierarchical index based on the semantic vectors, dynamically adjust the hierarchical structure of the index, and determine semantic similarity according to the hierarchical structure of the index; Perform cross-modal attention calculation based on the semantic vectors of the original tender document, non-text multimodal features, and multimodal features of candidate documents to determine the weighting coefficients of each modality, and dynamically allocate the weights of each modality feature in the candidate document according to the proportion of each modality feature in the original tender document; Determine the non-text similarity according to the weighting coefficients, the weights, and the special matching degrees of various modalities, and determine the global similarity between the original tender document and the candidate document according to the weighted sum of the semantic similarity and the non-text similarity.
2. The multi-modal duplicate checking method for tender documents based on a large model according to claim 1, wherein After preprocessing the original tender document, obtain multi-modal associated metadata, and associate the semantic vectors with the multi-modal associated metadata; Extract the keyword set of the original tender document, and determine multiple candidate documents through correlation calculation. This process is distributed to multiple computing nodes for parallel execution; Calculate the semantic similarity between the semantic vectors of each candidate document and the query vector of the original tender document. This process performs parallel vector operations through the SIMD instruction set.
3. The multi-modal duplicate checking method for tender documents based on a large model according to claim 1, wherein Preprocess the original tender document to obtain text chunks with hierarchical tags, including: Match the title based on regular expressions, comprehensively calculate the weight of the title according to the font size, bold attribute, and position information. The weight is dynamically determined by the linear combination of the normalized value of the font size, bold flag, and position weight. When the weight exceeds the preset threshold, mark it as the title and generate a hierarchical index tree, and split the original tender document into independent text chunks with hierarchical tags.
4. The multi-modal duplicate checking method for tender documents based on a large model according to any one of claims 1-3, wherein Preprocess the original tender document to obtain multi-modal features, including: converting the table content into a logical tree structure; locating the coordinate axes, legends, and data points in the chart through an object detection model, extracting the numerical range vector, and constructing a topological relationship matrix of the data points based on a graph neural network; converting the formula obtained by the formula editor or the image formula into a symbol tree structure, generating a serialized expression through breadth-first traversal, and generating a hash fingerprint using locality-sensitive hashing; the logical tree structure, the topological relationship matrix, and the hash fingerprint constitute the multi-modal features.
5. The multi-modal duplicate checking method for tender documents based on a large model according to any one of claims 1-3, wherein For any text chunk, it is jointly represented by word embedding, paragraph type embedding, and hierarchical position encoding. The hierarchical position encoding dynamically adjusts the encoding frequency according to the title level; Before the combined representation result is input into the large language model, a domain adaptation unit is used to dynamically splice industry-specific prompt templates.
6. The large model-based multi-modal bid duplicate checking method according to any one of claims 1-3, characterized in that Dynamically adjust the hierarchical structure of the index, including: for the sub-library of the industry the number of index levels is adaptively calculated according to the data scale where , and is the hierarchical sparsity factor.
7. The large model-based multi-modal bid duplicate checking method according to any one of claims 1-3, characterized in that Perform cross-modal attention calculation, including: , where is the semantic vector corresponding to the th text chunk of the original tender is the candidate document multi-modal fusion function, represents the th text chunk of the candidate document represents the th non-text multi-modal feature of the text chunk of the candidate document, are all trainable projection matrices, is the vector dimension.
8. A multi-modal bid duplicate checking system based on a large model, characterized in that, It includes: A preprocessing unit configured to preprocess the original bid to obtain text chunks with hierarchical labels and multi-modal features of non-text; A semantic processing unit configured to semantically encode the text chunks based on a large language model with dynamic context awareness to obtain semantic vectors corresponding to the text chunks; A retrieval calculation unit configured to construct an industry-segmented hierarchical index based on the semantic vectors, dynamically adjust the hierarchical structure of the index, and determine the semantic similarity according to the hierarchical structure of the index; A multi-modal calculation unit configured to perform cross-modal attention calculation to determine the weighting coefficients of each modality according to the semantic vectors of the original bid, the multi-modal features of non-text, and the multi-modal features of the candidate document, and dynamically allocate the weights of each modality feature in the candidate document according to the proportion of each modality feature in the original bid; A combined duplicate checking unit configured to determine the non-text similarity according to the weighting coefficients, the weights, and the special matching degrees of various modalities, and determine the global similarity between the original bid and the candidate document according to the weighted sum of the semantic similarity and the non-text similarity.
9. A computer device, characterized in that, It includes: A processor and a computer-readable storage medium; The processor is adapted to execute a computer program; The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the large model-based multi-modal bid duplicate checking method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor for the large model-based multi-modal bid duplicate checking method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Test question duplicate checking method and system
CN118227850A
Method and device for constructing feature comparison library supporting semantic duplicate checking and novelty checking
CN118245564A
Memory retrieval method for enhancing multi-modal long-context dialogue ability of large language model
CN119293139A
Domain speech recognition method and system based on RAG
CN119296516A
Calculation method and system for unstructured text data
CN119474383A
Cited By
Short message template classification and automatic matching method and system based on AI technology
CN120597863A
Multi-modal bidding document intention anchoring and rule pluggable auditing method and system
CN120766304A
Intention anchoring and rule pluggable auditing method and system for multi-modal proposals
CN120766304B
RAG-oriented document analysis method and system and computer equipment
CN120849350A
Multimodal large model-based multimedia material label management method and system
CN120873214A