Intelligent Abstract Extraction Method and System for Bidding Documents Based on Image Recognition
Through image recognition and natural language processing technology, an abstract extraction model for bid documents was established, which solved the problem that existing technology was difficult to deal with bid documents in different industries, realized the intelligent and automated abstract extraction of bid documents, and improved the accuracy of extraction.
Patent Information
- Application Number
- CN202510421859.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The prior art is difficult to cover all scenarios when processing bid documents in different industries, and a single model is difficult to understand complex legal and regulatory texts, resulting in poor performance in the face of new types or formats of documents.
Through image recognition and natural language processing technology, bid file information is entered and abstract extraction model is established, keywords of bid files are extracted, and the conciseness and semantic consistency of the abstract are checked through the tuning correction unit, synonymous replacement and root association are performed to optimize the abstract.
The intelligent and automated abstract extraction of bid documents is realized, the ability to understand file images is enhanced, the accuracy of extraction is improved, and the bid documents in different fields can be better adapted to bid documents.
Smart Images

Figure CN119938908B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent document extraction, and discloses an intelligent abstract extraction method and system for tender documents based on image recognition. Background Art
[0002] Image recognition technology has developed rapidly in recent years. Especially driven by deep learning algorithms, the accuracy of image recognition has been significantly improved. Intelligent abstract extraction of tender documents is an application field of image recognition technology, which mainly involves automatically extracting key information from documents in the form of scanned copies or pictures to generate document summaries or for subsequent data processing. However, there are still many deficiencies. For example, tender documents may have various formats and layouts, including complex elements such as tables and graphics, which pose challenges to image recognition. High-quality training data is crucial for building an accurate model. However, in practical applications, it is not easy to obtain a large amount of well-annotated training data. Tender documents in different industries have differences in terms, formats, etc. It is difficult for a single model to cover all scenarios and needs to be adjusted or retrained for different fields.
[0003] For example, the Chinese patent application with the publication number CN118504559A discloses an intelligent extraction method and system for legal and regulatory annotation documents, including collecting texts with legal and regulatory annotation bases as the input original legal and regulatory annotation basis texts, and performing data preprocessing on the original legal and regulatory annotation basis texts; implementing feature construction based on feature engineering, at least constituting the title feature, the table text feature, the non-table text feature and the symbolic feature; using the constructed features to extract key information from the original legal and regulatory annotation basis texts, and based on the extracted key information, automatically identifying key entity information in the legal and regulatory texts through text scanning, splitting, feature comparison, regular matching, etc. Compared with the prior art, the invention can improve the accuracy of 2D fixation point estimation.
[0004] The dataset of the training model of the above patent may not be diverse enough, resulting in poor performance of the model when facing new types or formats of documents. For example, different documents may have different term usage habits, format layouts, etc. Although natural language processing technology has made great progress, it still has certain limitations in understanding and extracting complex semantics. Legal and regulatory texts often contain a large number of professional terms and complex logical relationships, which pose high requirements for the machine's understanding ability. Relying on fixed rules for information extraction, such methods may fail when facing flexible text formats. For example, regular expressions are very useful in some scenarios, but may become inflexible when facing complex text structures. Summary of the Invention
[0005] The purpose of this section is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract, and the title, and such simplifications or omissions shall not be used to limit the scope of the present invention.
[0006] To solve the above technical problems, the main object of the present invention is to provide an intelligent abstract extraction method for tender documents based on image recognition, including:
[0007] S1. Select the application field to which the tender document belongs according to the tender document (detailed description of the application field, with examples);
[0008] S2. Enter the tender document information through image scanning (specifically, the technical means of image scanning, which can be achieved by conventional technical means), and establish an abstract extraction model to extract the keywords of the tender document;
[0009] S3. Establish an abstract of the tender document by sorting the keywords, and have the tuning and correction unit check whether the abstract is smooth and whether it conforms to the semantics of the original tender document;
[0010] S4. If the abstract is smooth, continue to check whether it conforms to the semantics of the original tender document. If the abstract is not smooth, perform synonym replacement and root word association on the keywords, and re-sort the keywords;
[0011] S5. If it conforms to the semantics of the original tender document, complete the intelligent extraction of the abstract of the tender document. If it does not conform to the semantics of the original tender document, re-check whether the keywords are correct.
[0012] As a preferred solution of the intelligent abstract extraction method for tender documents based on image recognition of the present invention, wherein:
[0013] The image scanning for entering the tender document information converts the document into an editable electronic text format through image recognition and natural language processing;
[0014] The image recognition is used to extract the text content from the editable electronic text format;
[0015] The natural language processing is used to initially understand the extracted text.
[0016] As a preferred solution of the intelligent abstract extraction method for tender documents based on image recognition of the present invention, wherein:
[0017] The abstract extraction model includes a word density analysis unit, a word priority screening unit, a word priority sorting unit, and an abstract generation and verification unit;
[0018] The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function;
[0019] The word priority screening unit calculates the priority weight of words by setting the word priority sorting logic;
[0020] The word priority sorting unit is used to receive the word priority weight and sort the word priorities;
[0021] The abstract generation and verification unit is used to generate a tender document abstract according to the word priority sorting and perform integrity verification on the tender document abstract.
[0022] As a preferred solution of the intelligent abstract extraction method for tender documents based on image recognition of the present invention, wherein:
[0023] The word density analysis unit divides the text content extracted from the editable electronic text format into regions by establishing a two-dimensional coordinate system with a size of a×a, numbers the grids from left to right and from top to bottom, and calculates the density of words appearing in each part of the grid. The calculation expression of the word density analysis unit is as follows:
[0024]
[0025] Wherein, is the word density of the grid with the abscissa i and the ordinate j in the electronic text format, is the weight of the kth word, is the length of the kth word, is the distance factor of the kth word in the grid (i, j), is the total number of words;
[0026] The calculation expression of the distance factor is:
[0027]
[0028] Wherein, is the distance factor, e is the exponential constant, is the abscissa position of the word k, is the abscissa of the grid center position, is the ordinate position of the word k, is the ordinate of the grid center position, α is the distance coefficient, related to the word spacing;
[0029] The distance factor is a value between 0 and 1, indicating that the importance of the word changes with the position.
[0030] As a preferred solution of the intelligent abstract extraction method for tender documents based on image recognition of the present invention, wherein:
[0031] The word priority screening unit collects the density and position of grid words in the region segmentation, extracts features from the collected density and position of grid words, calculates weights based on the extracted features, and finally sorts and outputs according to the priority weights of all grid words;
[0032] The feature extraction is used to extract the word frequency and inverse document frequency index of the grid word, the length of the grid word, and the position weight of the grid word.
[0033] As a preferred scheme of the intelligent abstract extraction method for tender documents based on image recognition of the present invention, wherein:
[0034] The calculation expression of the weight calculation is as follows:
[0035]
[0036] Wherein, is the priority weight of the grid word, is the word frequency and inverse document frequency index of the grid word, is the length of word k, is the position weight of word k;
[0037] The calculation expression of the position weight is as follows:
[0038] P k = { β , if [ P i = " Title "] δ , if [ P i = " First sentence "] λ , othrewise
[0039] Wherein, is the position weight coefficient of the grid word as the title, is the position weight coefficient of the grid word as the first sentence, is the position weight coefficient of the grid word in the sentence, is the position of the k-th word in the tender document, is other situations;
[0040] The priority weight of the grid word, the word frequency and inverse document frequency index of the grid word, the word length, and the word position weight are normalized by maximum normalization processing;
[0041] The calculation expression of the maximum normalization is as follows:
[0042] Y ' = F ( S ) − min F ( S )] max F ( S )] − min F ( S )]
[0043] Wherein, is the data after maximum normalization processing, is the data to be normalized input, min[] is the minimum value function, and max[] is the maximum value function;
[0044] The word priority sorting unit is used to receive the word priority weights after maximum normalization processing and sort the word priorities.
[0045] As a preferred solution of the intelligent abstract extraction method for tender documents based on image recognition of the present invention, wherein:
[0046] The keyword sorting sorts the grid words in descending order of weight according to the extracted grid words and their weights;
[0047] The establishment of the tender document abstract includes selecting the top N keywords from the sorted keyword list and constructing a logically coherent sentence or paragraph as the abstract;
[0048] A screening threshold is set through the word priority sorting to screen out high-priority words as the keywords;
[0049] The tuning and correction unit constructs a correction function and a pruning function to perform grammar checking and semantic consistency checking on the abstract. The correction function is used to correct the symmetry, compactness, vanishing distance, and orthogonality problems during image feature extraction. The pruning function compensates the word density analysis unit through the database and text image information.
[0050] As a preferred solution of the intelligent abstract extraction method for tender documents based on image recognition of the present invention, wherein:
[0051] If the abstract is smooth, the semantic consistency of the abstract is checked again through the semantic similarity algorithm. If the abstract is not smooth, it is made smooth through synonym replacement, root word association, and reordering;
[0052] If it does not conform to the semantics of the initial tender document, the keyword density analysis unit and the word priority sorting unit are used to screen the keywords of the tender document again, and it is checked again whether there are missing and incorrect keywords;
[0053] The semantic similarity algorithm measures the similarity between two vectors through cosine similarity, and the calculation expression is as follows:
[0054]
[0055] Wherein, is the similarity between the abstract and the text, is the cosine similarity, is the vector representation of the abstract, is the vector representation of the tender document;
[0056] For the extracted keyword k, the synonym replacement matches a group of synonyms from the preset synonym library and selects the synonym with the highest semantic matching degree to replace the original keyword;
[0057] For a specific word t, the root word association is based on the root word root(t), and its synonyms and derivatives are retrieved from a preset word library for replacement.
[0058] An intelligent abstract extraction system for tender documents based on image recognition, comprising:
[0059] A document recognition module, including a scanning unit for scanning tender document information and a classification unit for defining the technical field to which the tender document belongs;
[0060] An input module, including an input unit for image scanning input, an image recognition unit for extracting text content from an editable electronic text format, and a natural language processing unit for preliminarily understanding the extracted text;
[0061] An abstract extraction model, including a word density analysis unit, a word priority screening unit, a word priority sorting unit, and an abstract generation and verification unit;
[0062] A feedback module includes a tuning and correction unit, a semantic similarity algorithm, synonym replacement, and root word association.
[0063] As a preferred solution of the intelligent abstract extraction system for tender documents based on image recognition of the present invention, wherein:
[0064] The word density analysis unit is used to extract and verify the density of word occurrences, and extract word features through a feature extraction function;
[0065] The word priority screening unit calculates the priority weight of words by setting a word priority sorting logic;
[0066] The word priority sorting unit is used to receive the word priority weight and sort the word priorities;
[0067] The abstract generation and verification unit is used to generate a tender document abstract according to the word priority sorting and perform integrity verification on the tender document abstract.
[0068] Advantages of the present invention:
[0069] The problem that different documents may have different term usage habits is solved by the abstract extraction model. The key information of the tender document is extracted through semantic intelligent segmentation and intelligent feature extraction of image recognition, the understanding ability of the document image is enhanced, the intelligence and automation of extraction are realized, and the accuracy of extraction is improved. Description of the Drawings
[0070] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the attached drawings required for the description of the embodiments. Obviously, the attached drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other attached drawings can also be obtained based on these attached drawings. Among them:
[0071] Figure 1 It is a flowchart of the intelligent abstract extraction method for tender documents based on image recognition of the present invention;
[0072] Figure 2 It is a flowchart of the working process of the word density analysis unit of the intelligent abstract extraction method for tender documents based on image recognition of the present invention;
[0073] Figure 3 It is a system composition diagram of the intelligent abstract extraction system for tender documents based on image recognition of the present invention;
[0074] Figure 4 It is a general structure diagram of the intelligent abstract extraction system for tender documents based on image recognition of the present invention;
[0075] Figure 5 It is an internal structure diagram of the intelligent abstract extraction system for tender documents based on image recognition of the present invention;
[0076] Figure 6 It is a physical diagram of the structure of the intelligent abstract extraction system for tender documents based on image recognition of the present invention;
[0077] Figure 7 It is a module diagram of the intelligent abstract extraction system for tender documents based on image recognition of the present invention.
[0078] Reference numerals: 1, cabinet body; 2, total layer hatch; 3, lower layer hatch; 4, roller-equipped floor; 5, wireless communication module; 6, keyboard; 7, recognition camera; 8, laser printer; 9, printer paper output tray; 10, host; 11, document recognition module; 12, input module; 13, lightning protection and power distribution unit; 14, data processing unit; 15, storage drawer. Detailed implementation manners
[0079] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific implementation manners of the present invention in conjunction with the attached drawings of the specification.
[0080] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0081] Second, the "one embodiment" or "embodiment" mentioned herein refers to specific features, structures or characteristics that may be included in at least one implementation manner of the present invention. The appearances of "in one embodiment" in different places in this specification do not all refer to the same embodiment, nor are they separate or alternative embodiments that exclude each other with other embodiments.
[0082] Embodiment 1
[0083] As Figure 1 shown, the intelligent abstract extraction method for tender documents based on image recognition includes:
[0084] S1. Select the application field to which the tender document belongs according to the tender document;
[0085] In this embodiment, the tender document may belong to multiple fields such as construction projects, information technology, medical equipment, energy development, environmental protection projects, etc. By analyzing key information such as project descriptions, technical requirements, and contract terms in the tender document, its belonging field can be initially judged; for example, if a tender document contains a large amount of content about "building structure design", "construction material selection", and "construction schedule plan", it can be judged that it belongs to the construction project field; by determining the application field to which the tender document belongs, it helps the subsequent information processing and abstract extraction to be more accurate for a specific field.
[0086] S2. Enter the tender document information through image scanning, and establish an abstract extraction model to extract the keywords of the tender document;
[0087] In this embodiment, the tender document information is entered through image scanning. The recognition accuracy of modern OCR (Optical Character Recognition) technology has reached a relatively high level. The image scanning technology of the present invention selects OCR technology to convert paper tender documents or electronic tender documents in formats such as PDF into editable text formats. OCR technology converts the characters in the image into computer-readable text data by recognizing the character shapes and arrangements in the image; by entering the tender document information through image scanning, the time and cost of manual entry are greatly reduced; establish an abstract extraction model to extract the keywords of the tender document, and extract the keywords that can summarize the main information of the document by analyzing the text content of the tender document; by condensing a large amount of text information into a small number of keywords, it is convenient to quickly understand the main content of the tender document. The extraction of keywords helps non-professionals quickly understand the key points of the tender document; for example, extract keywords such as "project name", "tender amount", "technical requirements", and "construction period" from a tender document.
[0088] S3. Establish a tender document abstract by sorting the keywords, and have the tuning and correction unit check whether the abstract is smooth and whether it conforms to the semantics of the original tender document;
[0089] In this embodiment, the keywords are sorted according to their importance and relevance, and then combined into an abstract according to a certain logical structure (such as chronological order, logical hierarchy, etc.). The basis for sorting can be the frequency of occurrence of the keywords in the document, their positions, the degree of association with other keywords, etc. Through reasonable sorting and combination, the abstract can be made more understandable. For example, the extracted keywords are combined into an abstract in the order of project name, bid amount, technical requirements, and construction period.
[0090] Specifically, the optimization and correction unit is used to check the grammar, semantics, and logic of the generated abstract. By comparing the abstract with the text content of the initial tender document, it is judged whether the abstract is smooth and whether it accurately conveys the semantics of the initial tender document. By correcting the abstract, it is ensured that the abstract accurately conveys the semantics of the initial tender document and avoids misunderstandings or ambiguities caused by inaccurate abstracts. For example, if there are deviations in the "technical requirements" part of the abstract from the description in the initial tender document, the optimization and correction unit will make corrections to ensure the accuracy of the abstract.
[0091] S4. If the abstract is smooth, continue to check whether it conforms to the semantics of the initial tender document. If the abstract is not smooth, perform synonym replacement and root word association on the keywords, and re-sort the keywords.
[0092] In this embodiment, when the abstract is not smooth or does not conform to the semantics of the initial tender document, synonym replacement and root word association techniques can be used to optimize the keywords. Synonym replacement is to replace inappropriate keywords with their synonyms or near-synonyms; root word association is to associate other relevant words based on the root or affix of the keyword to enrich the content of the abstract. Through synonym replacement and root word association, the content of the abstract can be flexibly adjusted to make the abstract more fluent and understandable. For example, if the term "construction period" in the abstract causes the abstract to be not smooth, it can be replaced with synonyms such as "duration" or "construction cycle".
[0093] S5. If it conforms to the semantics of the initial tender document, complete the intelligent extraction of the abstract of the tender document. If it does not conform to the semantics of the initial tender document, re-check whether the keywords are correct.
[0094] After performing synonym replacement and root word association, it is necessary to re-check whether the keywords are correct and whether they still conform to the semantics of the initial tender document. Through re-checking, it is ensured that the keywords in the abstract are correct, the quality of the abstract meets the requirements, it can accurately convey the semantics of the initial tender document, making the abstract more credible and reliable, and helping to improve the overall quality of the tender document.
[0095] Furthermore, the information of the tender documents input by image scanning is converted into an editable electronic text format through image recognition and natural language processing; image recognition is used to extract text content from the editable electronic text format; natural language processing is used to initially understand the extracted text.
[0096] In this embodiment, the paper tender documents are converted into digital images by using a scanner or a camera. These images are usually of high resolution to ensure that the text is clearly distinguishable. The images are preprocessed, including denoising, binarization (converting the image into black and white for clearer text recognition), and correction (such as rotation correction to ensure that the text is neatly arranged). Through high-quality scanning, every detail in the document can be accurately captured, and the document is converted into an editable electronic text format, which is easy to store and process; the text content is extracted from the preprocessed image through the OCR technology in image recognition technology, and the text in the image is converted into editable text for subsequent processing and analysis; the initially extracted text content is understood through natural language processing. The natural language processing technology can identify the keywords, phrases, and sentence structures in the document, thereby initially understanding the theme and key points of the document.
[0097] Furthermore, the abstract extraction model includes a word density analysis unit, a word priority screening unit, a word priority sorting unit, and an abstract generation and verification unit;
[0098] The word density analysis unit is used to extract and verify the density of word occurrences, and extract word features through a feature extraction function;
[0099] The word priority screening unit calculates the priority weights of words by setting a word priority sorting logic;
[0100] The word priority sorting unit is used to receive the word priority weights and sort the word priorities;
[0101] The abstract generation and verification unit is used to generate an abstract of the tender document according to the word priority sorting and perform an integrity verification on the abstract of the tender document.
[0102] Furthermore, as Figure 2 , the word density analysis unit divides the area of the text content extracted in the editable electronic text format, establishes a two-dimensional coordinate system of size a×a, numbers the grids from left to right and from top to bottom, and calculates the density of word occurrences in each part of the grid. The calculation expression of the word density analysis unit is as follows:
[0103]
[0104] Where is the word density of the grid with abscissa \(i\) and ordinate \(j\) in the electronic text format, is the weight of the \(k\)-th word, is the length of the \(k\)-th word, is the distance factor of the \(k\)-th word in the grid \((i, j)\), is the total number of words;
[0105] Furthermore, is used to represent all the words in the grid;
[0106] The calculation expression of the distance factor is:
[0107]
[0108] where, is the distance factor, \(e\) is the exponential constant, is the abscissa position of the word \(k\), is the abscissa of the grid center position, is the ordinate position of the word \(k\), is the ordinate of the grid center position, and \(\alpha\) is the distance coefficient, related to the word spacing;
[0109] The distance factor is a value between 0 and 1, indicating that the importance of the word changes with the position.
[0110] Furthermore, the word priority screening unit extracts features from the density and position of the grid words collected in the region segmentation, calculates weights according to the features extracted, and finally sorts and outputs according to the priority weights of all grid words;
[0111] The feature extraction is used to extract the word frequency and inverse document frequency index of the grid words, the length of the grid words, and the position weight of the grid words.
[0112] In this embodiment, the word priority screening unit first performs regional segmentation on the text or image content, divides the overall content into multiple grids, identifies and collects all the grid words within each grid. Through regional segmentation, the text or image content can be analyzed more meticulously, capturing more details. Multiple features are extracted from the collected grid words, including word frequency, inverse document frequency index, grid word length, and grid word position weight. By extracting multiple features, the importance and priority of grid words can be evaluated more comprehensively. Combining multiple features for screening can more accurately identify key information. A priority weight is calculated for each grid word based on the extracted features. Through weight calculation, the importance of grid words can be quantified into specific values, facilitating comparison and sorting. All grid words are sorted according to the calculated priority weights and output according to the sorting result. The sorting usually adopts the descending order, that is, the grid word with the highest weight is ranked first. The sorted result can intuitively show which grid words are more important, facilitating users to quickly obtain key information.
[0113] Furthermore, the calculation expression for the weight calculation is as follows:
[0114]
[0115] Wherein, is the grid word priority weight, is the word frequency and inverse document frequency index of the grid word, is the length of word k, is the position weight of word k;
[0116] The calculation expression for the position weight is as follows:
[0117] P k = { β , if [ P i = " Title "] δ , if [ P i = " First sentence "] λ , othrewise
[0118] Wherein, is the position weight coefficient when the grid word is a title, is the position weight coefficient when the grid word is the first sentence, is the position weight coefficient when the grid word is in the sentence, is the position of the k-th word in the tender document, is other situations;
[0119] The grid word priority weight, the word frequency and inverse document frequency index of the grid word, the word length, and the word position weight are normalized through maximum normalization processing;
[0120] The calculation expression for the maximum normalization is as follows:
[0121] Y ' = F ( S ) − min F ( S )] max F ( S )] − min F ( S )]
[0122] Among them, is the data after maximum normalization, is the data that needs to be normalized and input, min[] is the minimum value function, and max[] is the maximum value function;
[0123] The word priority sorting unit is used to receive the word priority weights after maximum normalization and sort the word priorities.
[0124] Furthermore, the keyword sorting sorts the grid words according to the extracted grid words and their weights in descending order of weights;
[0125] The establishment of the tender document abstract includes selecting the top N keywords from the sorted keyword list and constructing a logically coherent sentence or paragraph as the abstract;
[0126] The screening threshold is set through the word priority sorting to screen out high-priority words as the keywords;
[0127] The tuning and correction unit constructs a correction function and a pruning function to perform grammar checking and semantic consistency checking on the abstract. The correction function is used to correct the symmetry, compactness, vanishing distance, and orthogonality problems during image feature extraction. The pruning function compensates the word density analysis unit through the database and text image information.
[0128] In this embodiment, the keywords are sorted from high to low according to the weight values. Through sorting, the keywords with higher weights can be processed preferentially, thereby improving the efficiency of information processing and text analysis; the top N (N is a preset value) keywords are selected from the sorted keyword list as the core content of the abstract, and a logically coherent and information-complete sentence or paragraph is constructed as the abstract according to these keywords. When constructing the abstract, the relevance and context logic between the keywords need to be considered to ensure the accuracy and readability of the abstract; by setting a screening threshold in the word priority sorting, the high-priority keywords are screened out. Through the screening threshold, the keywords with lower weights can be removed, thereby optimizing the quality of the keyword list; the tuning and correction unit performs grammar checking and semantic consistency checking on the abstract by constructing a correction function and a pruning function. The correction function is used to solve the symmetry, compactness, vanishing distance, and orthogonality problems during image feature extraction to ensure the accuracy of the image information in the abstract. The pruning function compensates the word density analysis unit through the database and text image information to improve the accuracy and integrity of the abstract.
[0129] Further, if the abstract is smooth, the semantic consistency of the abstract is checked again through the semantic similarity algorithm. If the abstract is not smooth, it is made smooth through synonym replacement, root word association, and reordering;
[0130] If it does not conform to the semantics of the initial tender document, the keyword screening of the tender document is re-performed through the word density analysis unit and the word priority sorting unit, and whether there are missing and incorrect keywords is re-checked;
[0131] The semantic similarity algorithm measures the similarity between two vectors through cosine similarity, and the calculation expression is as follows:
[0132]
[0133] Among them, is the similarity between the abstract and the text, is the cosine similarity, is the vector representation of the abstract, is the vector representation of the tender document;
[0134] For the extracted keyword k, the synonym replacement matches a group of synonyms from the preset synonym library and selects the synonym with the highest semantic matching degree to replace the original keyword;
[0135] For the specific word t, the root word association is based on the root word root(t), and its synonyms and derivative words are retrieved from the preset word library for replacement.
[0136] Embodiment 2
[0137] As Figure 3 shown, the intelligent abstract extraction system for tender documents based on image recognition includes:
[0138] The document recognition module includes a scanning unit for scanning tender document information and a classification unit for defining the technical field to which the tender document belongs;
[0139] The input module includes an input unit for image scanning input, an image recognition unit for extracting text content from an editable electronic text format, and a natural language processing unit for initially understanding the extracted text;
[0140] The abstract extraction model includes a word density analysis unit, a word priority screening unit, a word priority sorting unit, and an abstract generation and verification unit;
[0141] The feedback module includes a tuning and correction unit, a semantic similarity algorithm, synonym replacement, and root word association;
[0142] Among them, a correction function and a pruning function are constructed by the tuning and correction unit to perform a syntax check and a semantic consistency check on the abstract. The correction function is used to correct the symmetry, compactness, vanishing distance, and orthogonality problems during image feature extraction. The pruning function compensates the word density analysis unit through a database and text image information;
[0143] The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function;
[0144] The word priority screening unit calculates the priority weights of words by setting a word priority sorting logic;
[0145] The word priority sorting unit is used to receive the word priority weights and sort the word priorities;
[0146] The abstract generation and verification unit is used to generate a tender document abstract according to the word priority sorting and perform an integrity check on the tender document abstract.
[0147] Embodiment III
[0148] As Figure 4 shown, a schematic structural diagram of an intelligent abstract extraction system for tender documents based on image recognition includes: a cabinet 1, a main layer hatch 2, a lower layer hatch 3, and a roller floor 4. The cabinet 1 is used to support a display, and the lower layer hatch 3 is used to store a lightning protection and power distribution unit 13, a data processing unit 14, and a storage drawer 15 as Figure 5 shown;
[0149] Furthermore, the main layer hatch 2 is used to store a host 10, a document recognition module 11, and an input module 12;
[0150] As Figure 5 shown, it further includes a wireless communication module 5, a keyboard 6, an identification camera 7, a laser printer 8, and a printer paper output tray 9.
[0151] Among them, the keyboard 6 is used to control the input of tender documents, the identification camera 7 is used to identify the content of paper tender documents, and the laser printer 8 is used to print the abstract generation result.
[0152] As Figure 6 is a physical diagram of the structure of an intelligent abstract extraction system for tender documents based on image recognition.
[0153] As Figure 7As shown, for the intelligent abstract extraction system of tender documents, after inputting the tender documents for office supplies and entering them into the system, the file size is parsed and the input time is recorded, and the original text is previewed. Through marking, the file is scanned and keywords are extracted to conduct preliminary identification and processing on the tender documents for office supplies. The generation process of the tender document abstract is monitored through the processing progress, and finally, the final abstract generation result of the tender documents for office supplies is generated from the abstract result.
[0154] Importantly, it should be noted that the construction and arrangement of the present application shown in multiple different exemplary embodiments are merely illustrative. Although only two embodiments are described in detail in this disclosure, those who refer to this disclosure should easily understand that many modifications are possible without substantially departing from the novel teachings and advantages of the subject matter described in this application. For example, changes in the dimensions, scales, structures, shapes and proportions of various elements, as well as parameter values (such as temperature, pressure, etc.), installation arrangements, use of materials, colors, orientations, etc. For example, an element shown as integrally formed may be composed of multiple parts or elements, the position of the element may be inverted or otherwise changed, and the nature, number or position of discrete elements may be changed or altered. Therefore, all such modifications are intended to be included within the scope of the present invention. The order or sequence of any process or method steps may be changed or reordered according to alternative embodiments. Any "means-plus-function" clause is intended to cover the structures that perform the functions described herein, and not only structural equivalents but also equivalent structures. Other substitutions, modifications, changes and omissions may be made in the design, operating conditions and arrangement of the exemplary embodiments without departing from the scope of the present invention. Therefore, the present invention is not limited to specific embodiments, but extends to various modifications that still fall within the scope of the appended claims.
[0155] In addition, in order to provide a concise description of the exemplary embodiments, not all features of the actual embodiments may be described (i.e., those features that are not relevant to the currently considered best mode of implementing the present invention, or those features that are not relevant to the implementation of the present invention).
[0156] It should be understood that in the development process of any actual implementation, as in any engineering or design project, a large number of specific implementation decisions may be made. Such development efforts may be complex and time-consuming, but for those of ordinary skill in the art who benefit from this disclosure, without excessive experimentation, the development efforts will be a routine task of design, manufacturing and production.
[0157] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. The intelligent summary extraction method of bidding documents based on image recognition is characterized by: include: S1. Select the application field to which the bidding document belongs according to the bidding document; S2, input bidding document information through image scanning, and establish a summary extraction model to extract bidding document keywords; The abstract extraction model includes a word density analysis unit, a word priority screening unit, a word priority sorting unit and an abstract generation verification unit; The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function; The word priority screening unit calculates the priority weights of the words by setting the word priority sorting logic; The word priority sorting unit is used to receive word priority weights and sort the word priorities; The summary generation and verification unit is used to generate a summary of the bidding document according to the word priority ranking, and perform integrity verification on the summary of the bidding document; S3. Establish a summary of the bidding document by sorting keywords, and use the optimization and correction unit to check whether the summary is fluent and conforms to the semantics of the initial bidding document; S4. If the abstract is fluent, continue to check whether it conforms to the semantics of the initial bidding document. If the abstract is not fluent, perform synonym replacement and root association on the keywords and re-order the keywords; S5. If the bidding document semantics are consistent, the intelligent extraction summary of the bidding document is completed. If the bidding document semantics are not consistent, the keywords are rechecked to see if they are correct.
2. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 1 is characterized in that: The image scanning and input bidding document information converts the document into an editable electronic text format through image recognition and natural language processing; The image recognition is used to extract text content from an editable electronic text format; The natural language processing is used to initially understand the extracted text.
3. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 2 is characterized in that: The word density analysis unit extracts text content from the editable electronic text format and performs regional segmentation, establishes a two-dimensional coordinate system of size a×a, numbers the grids from left to right and from top to bottom, and calculates the density of word occurrences in each part of the grid. The word density analysis unit calculates the expression as follows: ; in, is the word density of the grid with horizontal coordinate i and vertical coordinate j in the electronic text format, is the k-th word weight, is the length of the kth word, is the distance factor of the kth word in the grid (i, j), is the total number of words; The distance factor calculation expression is: ; in, is the distance factor, e is the exponential constant, is the horizontal coordinate position of word k, is the horizontal coordinate of the grid center position, is the ordinate position of word k, is the vertical coordinate position of the grid center, α is the distance coefficient; The distance factor is a value between 0 and 1, indicating that the importance of a word changes with its position.
4. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 3 is characterized by: The word priority screening unit collects the density and position of the grid words in the area segmentation, extracts features from the collected density and position of the grid words, calculates weights based on the features extracted from the features, and finally sorts and outputs the grid words according to the priority weights of all the grid words; The feature extraction is used to extract the word frequency and inverse text frequency index of the grid word, the grid word length and the grid word position weight.
5. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 4 is characterized in that: The calculation expression of the weight calculation is as follows: ; in, is the grid word priority weight, is the frequency and inverse text frequency index of the grid words, is the length of word k, is the position weight of word k; The position weight calculation expression is as follows: ; in, is the position weight coefficient of the grid word as the title, is the position weight coefficient of the grid word as the first sentence, is the position weight coefficient of the grid word in the sentence, is the position of the kth word in the bidding document, For other cases; Normalizing the grid word priority weight, the grid word frequency and inverse text frequency index, the word length and the word position weight by maximum normalization; The maximum normalization calculation expression is as follows: ; in, is the data after maximum normalization processing, is the input data that needs to be normalized, min[] is the minimum value function, and max[] is the maximum value function; The word priority sorting unit is used to receive the word priority weights after maximum normalization and sort the word priorities.
6. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 5 is characterized in that: The keyword sorting is based on the extracted grid words and their weights, and the grid words are sorted in descending order of weight; The step of establishing the bid document summary includes selecting the first N keywords from the sorted keyword list and constructing a logically coherent sentence or paragraph as the summary; Setting a screening threshold by ranking the word priorities to screen out high-priority words as the keywords; The correction function and the pruning function are constructed by the tuning correction unit to perform grammatical check and semantic consistency check on the summary. The correction function is used to correct the symmetry, compactness, vanishing distance and orthogonality problems when extracting image features. The pruning function compensates the word density analysis unit through the database and text image information.
7. The method for extracting intelligent abstracts from bidding documents based on image recognition according to claim 6 is characterized in that: If the abstract is fluent, the semantic consistency of the abstract is checked again through the semantic similarity algorithm. If the abstract is not fluent, synonym replacement, root association and re-ordering are performed to make the abstract fluent. If it does not conform to the semantics of the initial bidding document, the keywords of the bidding document will be screened again through the word density analysis unit and the word priority sorting unit to recheck whether there are any missing or wrong keywords; The semantic similarity algorithm measures the similarity of two vectors by cosine similarity. The calculation expression is as follows: ; in, is the similarity between the abstract and the text, is the cosine similarity, is the vector representation of the summary, is the vector representation of the bidding document; For the extracted keyword k, the synonym replacement matches a group of synonyms from a preset synonym library, and selects the synonym with the highest semantic matching degree to replace the original keyword; For a specific word t, the root association is based on the root root(t), and its synonyms and derivatives are retrieved from the preset word library for replacement.
8. A system for intelligent abstract extraction of bidding documents based on image recognition, used to implement the method for intelligent abstract extraction of bidding documents based on image recognition as claimed in any one of claims 1 to 7, characterized in that: include: The document identification module includes a scanning unit for scanning the bidding document information and a classification unit for defining the technical field to which the bidding document belongs; An input module, including an input unit for image scanning and input, an image recognition unit for extracting text content from an editable electronic text format, and a natural language processing unit for preliminary understanding of the extracted text; A summary extraction model, including a word density analysis unit, a word priority screening unit, a word priority sorting unit and a summary generation verification unit; The feedback module includes a tuning correction unit, a semantic similarity algorithm, synonym replacement, and root association.
9. The bidding document intelligent summary extraction system based on image recognition according to claim 8 is characterized by: The word density analysis unit is used to extract and check the density of word occurrences, and extract word features through a feature extraction function; The word priority screening unit calculates the priority weights of the words by setting the word priority sorting logic; The word priority sorting unit is used to receive word priority weights and sort the word priorities; The summary generation and verification unit is used to generate a summary of the bidding document according to the word priority ranking, and to perform integrity verification on the summary of the bidding document.
Citation Information
Patent Citations
Intelligent extraction method and system for law and regulation annotation files
CN118504559A
Course target achievement condition evaluation rationality evaluation method based on semantic analysis
CN116362927A