Target traceability method, electronic equipment, storage medium and program product

By filtering bidding pairs with the same keywords as the project name, and combining the similarity between text semantics, numerical differences and time differences, the problems of accuracy and efficiency in bidding data traceability are solved, and efficient and accurate bidding traceability are achieved.

CN120217006APending Publication Date: 2025-06-27BEIJING QIANLIMA NETWORK INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510263626.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the field of bidding, with the surge in data volume, it has become difficult to effectively locate and screen related information, the processing speed and efficiency seem to be ineffective, and the diversification of data formats and structures reduces accuracy. The existing technology is not accurate enough in the traceability of bid documents, resulting in unreliable results.

Method used

By screening the pairs of the same item name keywords from the target library, combining text semantic similarity, numerical differences and time differences, the comprehensive similarity of the pair is determined. If the preset similarity threshold is exceeded, it is judged that the pair belongs to the same item.

Benefits of technology

It improves the accuracy and efficiency of the tracing of the tracing of the tracing of the tracing of the tracing of the tracing of the tracing of the tracing of the tracing, and accurately determines whether the tracing belongs to the same project by comprehensively considering various information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217006A_ABST
    Figure CN120217006A_ABST
Patent Text Reader

Abstract

According to the bid text tracing method, the electronic equipment, the storage medium and the program product, the bid text pairs with the same project name keywords are screened from the bid library, the comparison range is effectively reduced, the efficiency is improved, and the direct relevance of the bid text pairs on the project names is ensured. And then, the similarity of the bid text pairs is evaluated by combining the text semantic similarity, the numerical difference and the relationship between the time difference and the target time window, so that efficient and accurate source tracing of the bid text is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data mining and analysis. Specifically, it relates to a method for tracing the source of a bid document, an electronic device, a storage medium, and a program product. Background Art

[0002] With the rapid progress of modern technology, the speed of data generation and dissemination has reached an unprecedented level, and the amount of data shows an explosive growth trend. In the field of bidding and tendering, the introduction of new policies requires the full publicity of all information and process stages related to bidding and tendering projects, which undoubtedly further exacerbates the sharp expansion of bidding and tendering data volume. These data are widely penetrated into all walks of life, covering every subtle link from project initiation to completion.

[0003] However, this rapid increase in data volume is accompanied by a series of severe challenges. On the one hand, it has become increasingly difficult to accurately track the progress of a single project among numerous data, because effectively locating and screening relevant information has become an arduous task. On the other hand, the huge scale of data makes it feel powerless in terms of processing speed and efficiency. In addition, the diversification of data formats and structures further increases the complexity and uncertainty in extracting key feature fields, reducing the accuracy.

[0004] In related technologies, although there have been some methods attempting to solve these problems, there are still some deficiencies. For example, the feature extraction is not accurate enough, resulting in unreliable results of tracing the source of the bid document. Therefore, there is an urgent need for a more efficient and accurate method to address the problem of tracing the source of bidding and tendering data. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide a method for tracing the source of a bid document, an electronic device, a storage medium, and a program product, so as to achieve the technical effect of improving the efficiency of tracing the source of the bid document.

[0006] The first aspect of the embodiments of the present application provides a method for tracing the source of a bid document, and the method includes:

[0007] Screen out bid document pairs with the same project name keywords from the bid document library; the bid document pair includes at least two bid documents, and each bid document in the bid document pair includes a text type field, a numerical type field, and a time type field;

[0008] Determine a first similarity based on the semantic similarity of the text in the text type field of the bid document pair, determine a second similarity based on the numerical difference of the numerical fields in the bid document pair, and determine a third similarity based on the relationship between the time difference of the time fields in the bid document pair and the size of the target time window; the target time window is determined based on the project type to which the bid document pair belongs;

[0009] Based on the first similarity, the second similarity, and the third similarity, determine the comprehensive similarity of the caption pair, and determine that the caption pair belongs to the same project when the comprehensive similarity exceeds a preset similarity threshold.

[0010] In the above implementation process, by comprehensively considering the similarity of the caption pair in text semantics, numerical differences, and time windows, accurately judge whether the caption pair belongs to the same project, effectively improving the accuracy and efficiency of caption traceability.

[0011] Further, before screening out the caption pairs with the same project name keywords from the caption library, it further includes:

[0012] Perform word segmentation processing and part-of-speech tagging processing on the project name of each caption in the caption library to obtain the word segmentation included in each caption and the part-of-speech tagging corresponding to the word segmentation.

[0013] For each caption, determine the project name keywords from all the word segmentations included in each caption according to the weight corresponding to each part-of-speech tagging result.

[0014] In the above implementation process, before screening out the caption pairs with the same project name keywords, first perform word segmentation and part-of-speech tagging on the project name of each caption, and determine the keywords based on the part-of-speech weights, effectively improving the accuracy of keyword extraction and laying a solid foundation for subsequent caption traceability.

[0015] Further, before screening out the caption pairs with the same project name keywords from the caption library, it further includes:

[0016] For the structured content in each caption, use a preset regular expression to extract fields to obtain a first field extraction result; the first field extraction result includes any one or more of a text type field result, a numerical type field result, and a time type field result;

[0017] For the unstructured content in each caption, extract the semantic information of the unstructured content and determine a second field extraction result according to the semantic information; the second field extraction result includes any one or more of a text type field result, a numerical type field result, and a time type field result;

[0018] According to the first field extraction result and the second field extraction result, determine the text type field, the numerical type field, and the time type field included in each caption.

[0019] In the above implementation process, for the different characteristics of the structured content segment and the unstructured content segment in the caption, different field acquisition methods are selected.

[0020] Further, determining the text type field, the numerical type field, and the time type field included in each caption according to the extraction results of the first field and the extraction results of the second field includes:

[0021] Correcting outliers in the extraction results of the first field and the extraction results of the second field, and / or filling in missing values in the extraction results of the first field and the extraction results of the second field to obtain processed extraction results;

[0022] Determining the text type field, the numerical type field, and the time type field according to the processed extraction results.

[0023] In the above implementation process, by correcting outliers and filling in missing values, the accuracy and integrity of the field extraction results are ensured.

[0024] Further, determining the first similarity based on the semantic similarity of the text on the text type field in the caption pair includes:

[0025] Determining the semantic similarity of the text according to the word vector similarity of the caption pair on the text type field, determining the weight assignment value corresponding to the text type field according to the importance degree of the text type field, and determining the first similarity based on the semantic similarity of the text and the weight assignment value;

[0026] Determining the second similarity based on the numerical difference of the numerical fields in the caption pair includes:

[0027] Determining the relative error rate of the numerical values of the caption pair on the numerical field as the second similarity.

[0028] In the above implementation process, by combining the word vector similarity and the field weight to determine the text semantic similarity, and at the same time using the relative error rate of the numerical field to measure the numerical similarity, the accuracy of the caption pair similarity evaluation is improved.

[0029] Further, before determining the third similarity based on the relationship between the time difference of the time fields in the caption pair and the size of the target time window, it further includes:

[0030] Based on the target field of each caption in the caption pair, determining the project type of each caption;

[0031] According to the project type of each caption, determining the project type to which the caption pair belongs.

[0032] In the above implementation process, first clarify the project type based on the target field of the caption, and then determine the project type of the caption pair.

[0033] Further, the method further includes:

[0034] If the item types of each of the said caption texts are the same, then determine the item type of any one of the caption texts in the caption text pair as the item type to which the caption text pair belongs;

[0035] If the item types of each of the said caption texts are different, then determine the preset initial time window as the target time window.

[0036] In the above implementation process, if the item types of multiple caption texts in the caption text pair are the same, then determine the item type of the caption text as the item type of the caption text pair, and then set the size of the time window according to the item type of the caption text pair. If the item types of multiple caption texts in the caption text pair are different, then the item type of the caption text pair cannot be determined. In this case, it is also impossible to determine the target time window according to the item type of the caption text pair. Therefore, set the initial time window as the target time window.

[0037] The second aspect of the embodiments of the present application provides an electronic device, and the electronic device includes:

[0038] A processor;

[0039] A memory for storing instructions executable by the processor;

[0040] Wherein, when the processor calls the executable instructions, the method described in any one of the first aspect is implemented.

[0041] The third aspect of the embodiments of the present application provides a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the steps of the method described in any one of the first aspect are implemented.

[0042] The fourth aspect of the embodiments of the present application provides a computer program product, and the computer program product includes a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspect is implemented. Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained according to these drawings without creative efforts.

[0044] Figure 1 It is a schematic flowchart of a caption text traceability method provided by an embodiment of the present application;

[0045] Figure 2The overall process schematic diagram of a document tracing method provided by an embodiment of the present application;

[0046] Figure 3 The structural block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.

[0048] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0049] In the related art, document tracing mainly extracts key information points in the document item name, retrieves data containing these key information in the entire database to obtain a basic database. Then, manual data verification is performed through key features such as the bidding unit, project number, and other fields. However, due to reasons such as huge data volume, inaccurate extraction of key feature fields, and high labor costs, the efficiency of document tracing is greatly affected.

[0050] In view of any of the above-mentioned problems, an embodiment of the present application provides a document tracing method, referring to Figure 1 , Figure 1 The process schematic diagram of a document tracing method provided by an embodiment of the present application.

[0051] In this embodiment, the method includes:

[0052] Step S10: Screen out document pairs with the same project name keywords from the document library; the document pair includes at least two documents, and each document in the document pair includes a text type field, a numerical type field, and a time type field;

[0053] It should be noted that the bid library is a database containing multiple bid texts, among which there are bid texts belonging to the same project and those belonging to different projects. Each bid text consists of multiple fields, which, as the basic components of the bid text, describe various attributes of the bid text in detail. These fields include, but are not limited to, the project name field (used to identify the project name), the tendering unit field (indicating the organization or unit initiating the tender), the tender budget field (showing the budget amount of the tender project), the winning bid unit field (indicating the winning organization or unit), the winning bid amount field (showing the amount of the winning bid project), and the tender number field (used to uniquely identify the number of a tender project), etc. And these fields can be roughly divided into text type fields, numerical type fields, and time type fields according to the data type. Before identifying multiple bid texts belonging to the same project, it is necessary to first screen out one or more pairs of bid texts with the same project name keywords from the bid library. This step aims to narrow the comparison scope and focus on those bid texts that are most likely to belong to the same project. Exemplarily, a keyword matching algorithm, such as regular expression matching or fuzzy matching, can be used to find pairs of bid texts containing the same project name keywords.

[0054] Step S20: Determine the first similarity based on the semantic similarity of the words in the text type field of the pair of bid texts, determine the second similarity based on the numerical difference of the numerical fields in the pair of bid texts, and determine the third similarity based on the relationship between the time difference of the time fields in the pair of bid texts and the size of the target time window; the target time window is determined based on the project type to which the pair of bid texts belongs;

[0055] It should be noted that different types of fields contain different information, which plays an important role in determining whether bid texts belong to the same project. For example, text type fields usually contain information such as project descriptions, requirements, and specifications, which are very important for understanding the essence and characteristics of the project; numerical type fields usually contain specific data such as budgets, quotations, dimensions, and quantities, which can reflect the scale, cost, and requirements of the project; time type fields provide time information such as the release date and deadline of the project, and these time information helps to determine whether the bid texts are within the same project timeline, that is, whether they are bid texts submitted for different stages or related activities of the same project. Therefore, specific processing for different types of fields can more comprehensively and accurately evaluate the similarity degree between bid texts.

[0056] Exemplarily, to evaluate the similarity of two captions in terms of text content and determine whether they are likely to belong to the same project, natural language processing techniques such as word vectors and semantic models (such as BERT, GPT, etc.) are used to calculate the semantic similarity of caption pairs in the text type field. If the text similarity of two captions is high, it indicates that they have a high degree of consistency in describing the key information of the project, and thus are more likely to belong to the same project.

[0057] To evaluate the similarity of two captions in the numerical type field and determine whether they are likely to belong to the same project, for each numerical type field, calculate the numerical differences between caption pairs, such as absolute difference, relative difference, etc., and determine the second similarity based on the magnitude of the numerical differences. Different stages or related activities of the same project should have a certain degree of coherence and consistency in numerical information. Therefore, the smaller the difference, the higher the similarity. If the information of two captions in the numerical type field is closer, it indicates that they have a high degree of consistency in the scale, cost, and requirements of the project, and are more likely to belong to the same project.

[0058] To evaluate the degree of association of two captions in terms of time information and determine whether they are likely to belong to different stages or related activities of the same project, extract all time fields from the caption pair, such as the release time field, and determine a reasonable target time window according to the project type to which the caption pair belongs. This time window should be able to reflect the reasonable time interval between different stages or related activities of the same project. For example, for a construction project, the target time window may be determined based on the average construction period of the project. Then calculate the time difference of the time fields between the caption pair, and determine the third similarity according to the size relationship between the time difference and the target time window. If the time difference is within the target time window, the similarity is high; otherwise, it is low. If the target time window is set to 90 days (i.e., the average construction period of the project), then two captions with a time difference within 90 days are very likely to be captions of different stages of the same construction project.

[0059] In addition, the similarity of fields can be calculated through the following replacement method: Use siamese neural network end-to-end learning to calculate the field similarity, and its replacement logic is to directly input caption pairs (such as tender announcements and winning bid announcements), and output the similarity score through the siamese network, replacing the manually designed multi-level algorithm. Specifically, the network structure: BERT two-tower model → fully connected layer → similarity scoring. The advantage of using siamese neural network end-to-end learning to calculate the field similarity is that it does not rely on weight assignment and sub-item calculation, but instead uses end-to-end deep learning, which can automatically learn complex semantic relationships and reduce manual feature engineering.

[0060] In addition, fuzzy logic matching can be used to calculate the similarity of fields. Its alternative logic is to design a fuzzy rule base (such as a membership function) to determine the similarity for fuzzy scenarios (such as "the abbreviated name of the tendering unit is inconsistent with the full name"). Exemplarily, a membership function for "similarity of unit names" is defined: character coincidence rate ≥ 80% → membership degree 1.0; 50% - 80% → membership degree 0.5. The advantage of using fuzzy logic matching to calculate field similarity is that it does not use numerical / text separation calculation, but instead uses fuzzy rule integration, which can flexibly handle problems with unclear boundaries.

[0061] Step S30: Based on the first similarity, the second similarity, and the third similarity, determine the comprehensive similarity of the bid document pair, and determine that the bid document pair belongs to the same project when the comprehensive similarity exceeds a preset similarity threshold.

[0062] It should be understood that in order to comprehensively consider various aspects of information such as text semantic similarity, numerical field differences, and time field differences to determine the overall similarity degree of the bid document pair and judge whether they belong to the same project, the first similarity, the second similarity, and the third similarity can be weighted and summed to obtain the comprehensive similarity of the bid document pair. Compare the comprehensive similarity with a preset similarity threshold. If the comprehensive similarity exceeds the preset similarity threshold, it is determined that the bid document pair belongs to the same project; otherwise, it is considered that they do not belong to the same project.

[0063] In specific implementation, it is necessary to calculate the comprehensive similarity between multiple bid documents according to the text type field, numerical type field, and time type field. However, if a certain type of field is missing, such as the time type field is missing, the time window matching can be skipped, and the comprehensive similarity is calculated only relying on the similarity of other fields. If a certain field has an incorrect format and cannot be parsed, it can be marked as "pending manual review".

[0064] The calculation formula for the comprehensive similarity is shown in Formula 1 below:

[0065]

[0066] In the formula, Sim total is the comprehensive similarity after weighting the similarities of all fields, n is the total number of fields participating in the calculation. Optionally, the comprehensive similarity is calculated according to the similarities of the tendering unit field, winning bidder unit field, project name field, winning bid amount field, budget amount field, and bid document release time field. i is the index of the current field, indicating the i-th field being calculated. w i is the weight of the i-th field, and this weight represents the importance of this field. Sim i is the similarity value of the i-th field.

[0067] Optionally, set the similarity threshold to 0.85. If Sim total ≥0.85, it is determined as the same project; if 0.6 ≤ Sim total <0.85, it is a situation where it is uncertain whether it is the same project, triggering manual review; if Sim total <0.6, it is determined as a different project.

[0068] In this embodiment, by screening for bid text pairs with the same project name keywords from the bid library, the comparison scope is effectively narrowed, the efficiency is improved, and the direct relevance of the bid text pairs in terms of the project name is ensured. Then, by combining the text semantic similarity, numerical differences, and the relationship between the time difference and the target time window to evaluate the similarity degree of the bid text pairs, efficient and accurate traceability of the bid text is achieved.

[0069] Based on any of the above embodiments, before step S10, it further includes:

[0070] Perform word segmentation processing and part-of-speech tagging processing on the project name of each bid text in the bid library to obtain the word segmentation included in each bid text and the part-of-speech tagging corresponding to the word segmentation;

[0071] It should be noted that a word segmentation algorithm (such as rule-based word segmentation, statistics-based word segmentation, or deep learning models, etc.) is used to segment the project name of each bid text into individual words or phrases.

[0072] Perform part-of-speech tagging on each word segmentation, that is, determine the grammatical role of each word segmentation in the sentence, such as noun, verb, adjective, etc. Part-of-speech tagging helps to subsequently determine which word segmentations are more likely to be project name keywords based on the part of speech. This is because, in the project name, keywords often have specific part-of-speech characteristics. For example, nouns (especially proper nouns) are more likely to be part of the project name because they usually represent entities, concepts, or events. Through part-of-speech tagging, word segmentations with specific parts of speech can be quickly screened out, and these word segmentations are more likely to be project name keywords. Exemplarily, by combining part-of-speech information and context, it is possible to more accurately determine which word segmentations are project name keywords. For example, in the "XX Highway Construction Project", "Highway" and "Construction", as a noun and a verb respectively, together constitute the core part of the project name, so "Highway Construction" is the project name keyword.

[0073] For each of the bid texts, determine the project name keywords from all the word segmentations included in each bid text according to the weight corresponding to each part-of-speech tagging result.

[0074] It is understandable that a weight can be assigned to each part of speech according to historical experience or statistical results. For example, nouns (especially proper nouns) play an important role in project names and are thus given a higher weight, while the weights of parts of speech such as adjectives are relatively low.

[0075] According to the part-of-speech tagging results and the corresponding weights, calculate the importance scores of each segmented word. The segmented words with higher scores are more likely to be the keywords of the project name. Sort the scores and select one or more segmented words with the highest scores as the keywords of the project name.

[0076] In a specific implementation, after receiving the project name retrieval term input by the user or an automated program, use a multi-dimensional feature extraction algorithm to extract keywords from the retrieval term. This step involves natural language processing (NLP) techniques such as word segmentation and part-of-speech tagging. Specifically, perform word segmentation on the retrieval term and split it into individual words or phrases. For example, for the project name "Bid-winning Announcement for the General Integration Service of the Provincial Unified Medical Security Information Platform (Phase I)", the system can segment it into ["province-wide", "unified", "medical security", "information platform", "(Phase I)", "project", "general integration service", "bid-winning", "announcement"]. Based on the word segmentation, perform part-of-speech tagging on each segmented word to determine its grammatical role, which helps to determine which segmented words are more likely to be the keywords of the project name subsequently. Next, assign weights to each word segmentation result. The weight calculation rules include: the weight of a core noun (such as "information platform") is 1.0; the weight of a modifier (such as "province-wide") is 0.5; the weight of a stop word (such as "of", "and") is 0. Then, based on the weight assignment results, filter out the segmented words with higher weights as the core keywords and combine them into the final keywords. For example, from the above word segmentation results, extract "medical security information platform" as the keyword. To improve the retrieval efficiency, all the bid texts in the bid text library have been preprocessed, and all the preprocessed bid texts carry their corresponding keywords and indexes. After determining the keywords in the retrieval term, directly search for the same keywords in the index without re-extracting keywords for each bid text, and thus the bid text pairs with the same project name keywords as the retrieval term and matching the retrieval term can be determined from the bid text library.

[0077] In this embodiment, the method of performing word segmentation processing and part-of-speech tagging processing on the project name of each bid text in the bid text library and extracting keywords based on part-of-speech weights significantly improves the accuracy and efficiency of keyword extraction and provides strong support for subsequent retrieval and matching.

[0078] Based on any of the above embodiments, before step S10, it further includes:

[0079] For the structured content in each of the said captions, use a preset regular expression to extract fields, obtaining a first field extraction result; the first field extraction result includes any one or more of a text type field result, a numerical type field result, and a time type field result;

[0080] Specifically, define corresponding regular expressions for each type of field to be extracted. For example, to extract a date, define a regular expression that matches common date formats (such as YYYY - MM - DD). Then, apply these regular expressions to the structured content part of the caption to identify and extract the field values that match the regular expression pattern. Finally, the extracted field values are classified according to their types (text, numerical, time) to form the first field extraction result.

[0081] In a specific implementation, a field positioning method based on a structured template is adopted. First, for common structured formats in captions, such as "Tendering Unit: XXX" and "Winning Bid Amount: XXX", a template library is predefined. These templates use regular expressions to quickly locate the positions of fields in the caption. For example, to match the "Tendering Unit" field, the system uses the regular expression re.search(r'Tendering Unit[::]\s*([^\n]+)', text); and to match the "Winning Bid Amount" field, the regular expression re.search(r'Winning Bid Amount[::]\s*([\d,,.]+)\s*yuan', text) is used. By applying these regular expressions, the field values that match the template can be identified and extracted. These extracted field values are classified and sorted according to their types (text, numerical, time, etc.), thereby obtaining the first field extraction result, which includes any one or more of a text type field result (such as "Tendering Unit"), a numerical type field result (such as "Winning Bid Amount"), and a time type field result (such as "Release Time"). In addition, to enhance the accuracy and robustness of field extraction, context association is used to correct the results. When the field name is relatively ambiguous or there is ambiguity, such as "Party A: XXX", keywords in the context, such as "Purchasing Party" or "Winning Bid Party", are combined to correct the attribution of the field. For example, if "Purchasing Party: A certain Municipal Education Bureau" appears in the text, the system will automatically map it to "Tendering Unit: A certain Municipal Education Bureau" to ensure the accuracy of field extraction.

[0082] For the unstructured content in each of the said captions, extract the semantic information of the unstructured content, and determine a second field extraction result according to the semantic information; the second field extraction result includes any one or more of a text type field result, a numerical type field result, and a time type field result;

[0083] It should be noted that structured content refers to information with a fixed format and position in the bid text, such as titles, dates, amounts, etc. Such information can usually be directly parsed from the web page source code through web crawler technology. Unstructured content, on the other hand, refers to descriptive information in free text form in the bid text. Such information usually contains richer details and background, such as project background, technical requirements, contract terms, etc. Compared with structured content, unstructured content is more flexible and diverse, but also more difficult to process and parse. In actual application scenarios, unstructured content often exists in attachments, such as Word documents, PDF files, etc. These attachments need to be parsed and processed through specific technologies, such as OCR (Optical Character Recognition) and NLP (Natural Language Processing), to extract useful information from them. In fact, as an official tender document, the bid text needs to contain complete information for bidders to refer to. Structured content provides the basic information framework, while unstructured content supplements more detailed descriptions and background information. The combination of the two together constitutes the complete information system of the bid text. In related technologies, over-reliance on fixed templates to extract fields results in the ability to only recognize fixed formats such as "Tendering Unit: XXX", but unstructured expressions (such as "Party A is the Education Bureau of a certain city") often appear in actual bid texts, so the failure rate of field extraction is relatively high.

[0084] In specific implementation, for the natural language processing (NLP) of unstructured text, the following strategies are adopted to extract semantic information and determine the second field extraction result. This result includes any one or more of the text type field result, numerical type field result, and time type field result. Entity recognition and relationship extraction strategy: Use a pre-trained NLP model to identify key entities and their semantic relationships in the text. Entity types include but are not limited to organization names (such as tendering unit, winning bid unit), amounts (such as budget, winning bid price), time (such as release time), and project names. Through relationship extraction, information such as "XXX Company won the bid with XXX yuan" can be extracted from the text and transformed into fields such as winning bid unit and winning bid amount; similarly, information such as "The project budget is XXX yuan" can be transformed into the tender budget field. Semantic role annotation strategy: To understand the text information more deeply, semantic role annotation technology is adopted to analyze the role relationship between predicates and entities in the sentence. Through this technology, field information can be extracted from the sentence. For example, when processing the sentence "Hunan Provincial Medical Security Bureau (tendering unit) released a smart medical project with a budget of 5.3 million yuan on December 1, 2023", key fields such as the actor (tendering unit), time (release time), amount (tender budget), and project name can be identified. By combining entity recognition, relationship extraction, and semantic role annotation and other strategies, useful semantic information can be effectively extracted from unstructured text, and the second field extraction result can be determined accordingly.

[0085] Determine the text type field, numeric type field, and time type field included in each caption based on the first field extraction result and the second field extraction result.

[0086] It should be noted that the reason for determining the text type field, numeric type field, and time type field of each caption by combining the first field extraction result and the second field extraction result is to utilize the complementarity of the two extraction results to solve the problems of possible incomplete fields or inconsistent extraction results. Specifically, when both extraction results contain the same field results, the accuracy and integrity of these fields will be compared. If they are exactly the same, the field results will be directly adopted. If the two extraction results are inconsistent but the differences are small (such as minor differences in format), the optimal result will be selected based on the integrity of the information and the standardization of the format. In addition, when a certain field only appears in one extraction result, that is, a field extraction result contains a field that does not exist in the other extraction result, the uniqueness principle needs to be followed, and the only existing field will be directly adopted. For example, if the text type field "tendering unit" is unique in the first field extraction result, this field will be directly adopted. For the differences or inconsistencies in a certain field between the two extraction results, some technical means will be used for correction, such as context analysis, regular expression matching optimization, fine-tuning of the NLP model, etc. Taking the time type field as an example, if the dates given by the two extraction results are different, other time clues in the caption (such as project cycle, bid closing date) will be combined for comprehensive analysis to ensure that the value of the finally determined time type field is accurate. After integrating the first field extraction result and the second field extraction result and handling the differences between the two extraction results, the final field values of each caption will be obtained, including the text type field (such as tendering unit, winning bid unit, etc.), numeric type field (such as winning bid amount, budget amount, etc.), and time type field (such as release time, bid closing time, etc.).

[0087] In addition, fields can be extracted based on the structured parsing of the graph neural network (GNN). The specific steps include: regarding the caption as graph-structured data (nodes = fields, edges = semantic associations), and using GNN to automatically learn the relationships between fields to replace traditional template matching. Exemplarily, input caption → construct a text graph (such as the association between "tendering unit" and "budget amount") → the GNN model outputs the field mapping result. The advantage of extracting fields based on the structured parsing of the graph neural network (GNN) is that it is applicable to unstructured text content (such as the paragraph description part), does not rely on NLP entity recognition, uses graph structure modeling instead, and does not require predefined templates.

[0088] In addition, knowledge graph-driven entity association can be used to correct the first field extraction result and the second field extraction result to determine the text type field, numerical type field, and time type field included in each caption. The alternative logic is to complement the field information through graph query based on an industry knowledge graph (such as an enterprise name library, a project classification tree). Exemplarily, detecting "winning bidder: a certain construction company" → querying the knowledge graph → associating with "a certain provincial construction engineering co., ltd.". The advantage of using knowledge graph-driven entity association to correct the first field extraction result and the second field extraction result is that it does not rely on dynamic rule generation, but instead uses an external knowledge base to drive, solves the problem of inconsistent name abbreviations and full names, and enhances the robustness of field mapping.

[0089] In this embodiment, the field information of the structured content in the caption is extracted through a preset regular expression, and the field information of the unstructured content is extracted by using semantic analysis technology, realizing the accurate extraction of the field values of text, numerical, and time fields.

[0090] Based on any of the above embodiments, determining the text type field, the numerical type field, and the time type field included in each caption according to the first field extraction result and the second field extraction result includes:

[0091] Correcting the outliers in the first field extraction result and the second field extraction result, and / or filling the missing values in the first field extraction result and the second field extraction result to obtain a processed extraction result;

[0092] Determining the text type field, the numerical type field, and the time type field according to the processed extraction result.

[0093] It should be noted that outliers refer to data with incorrect formats, unreasonable logics, or significantly deviating from the normal range. The purpose of correcting outliers is to ensure that the extracted fields meet the expected format and logical requirements.

[0094] Missing values may be caused by incomplete information, mismatched extraction rules, or non-existence of the data itself. For the identified missing values, they need to be filled according to the context information and default values to ensure the completeness of the fields.

[0095] In specific implementation, the anomaly detection model is first trained using historical data to enable it to dynamically identify common error patterns in field extraction, including but not limited to the absence of currency units, improper name abbreviations, etc. Once the model detects an outlier in the first field extraction result or the second field extraction result, it will trigger an automatic generation mechanism for correction rules. For example, when the model detects data expressed as "Budget: 5 million", it will automatically complete it to "5,000,000 yuan" according to the preset rules to ensure the accuracy and consistency of the amount field. Similarly, if the model identifies an abbreviated name such as "Winning bidder: A certain construction company", it will use the enterprise database for intelligent matching and accurately correct it to the full name "A certain Provincial Construction Engineering Co., Ltd.". In addition, due to the diversity of the source of tender documents, a multi-source data adaptation engine is introduced. This engine can build corresponding correction rules for tender document formats from different sources such as government announcements and enterprise bidding platforms, and dynamically adjust the field extraction strategy to meet different requirements. For example, when processing government announcements, the engine will automatically correct fields such as "Purchaser" and "Budget amount" to more general expressions, such as the bidding unit and the bidding amount. When processing enterprise platform data, it will make corresponding corrections to fields such as "Party A" and "Contract amount" and convert them into standard fields such as the bidding unit and the winning bid amount.

[0096] In addition, a rule optimization method based on reinforcement learning can be used to determine the processed extraction result. Its alternative logic is to regard rule generation as a reinforcement learning task and dynamically optimize the rule library through a "trial and error - feedback" mechanism. Exemplarily, Action: Generate a correction rule (such as "Complete the currency unit"); Reward: Adjust the strategy according to the correction success rate. The advantage of using the rule optimization method based on reinforcement learning to determine the processed extraction result is that it does not rely on historical error pattern statistics, but instead uses an online learning mechanism, which can adapt to rapidly changing tender document formats.

[0097] In this embodiment, when integrating the first field extraction result and the second field extraction result, by correcting outliers and filling in missing values, the accuracy and integrity of the extraction result are ensured, thereby accurately determining the keyword fields of text, numerical values, time, etc. in the tender document.

[0098] Based on any of the above embodiments, determining the first similarity based on the literal semantic similarity of the text type field in the tender document includes:

[0099] Determining the literal semantic similarity according to the word vector similarity of the text type field in the tender document, determining the weight assignment value corresponding to the text type field according to the importance degree of the text type field, and determining the first similarity based on the literal semantic similarity and the weight assignment value;

[0100] It is understandable that for the text type fields in the bid document pair, the text needs to be converted into word vectors first. The word vector technology can map the words in the text into a high-dimensional space to capture the semantic relationships between words. On this basis, similarity calculation methods such as cosine similarity or Euclidean distance can be used to calculate the similarity between the word vectors of two text fields. Among them, cosine similarity mainly measures the similarity degree of two vectors in direction, while Euclidean distance focuses on measuring the straight-line distance between two vectors in space. Given that different text type fields (such as the tendering unit, winning bid unit, etc.) have different importance in describing the project name or bid document content, it is necessary to assign a weight value to each text type field. This weight value can reflect the importance of the field in the overall similarity calculation. The assignment of weights can comprehensively consider factors such as the occurrence frequency of the field, the uniqueness of the field content, and the key nature of the field's description of the project name or bid document content. Optionally, according to the contribution degree of different fields to the project uniqueness, combined with historical data training, the weight of the tender number field is 0.35 (the uniqueness of the tender number is the strongest, so the weight is the highest), the weight of the tendering unit field is 0.25 (the relevance of projects of the same unit is high), the weight of the project name field is 0.2 (the names are similar but the content may be different), the weight of the budget amount field is 0.15 (similar amounts may be for bid documents at different stages of the same project), and the weight of the release time field is 0.05 (similar times are used to assist in determining whether it is the same project).

[0101] In the specific implementation, for the text type fields, an improved cosine similarity calculation method is used, which combines word vectors (such as Word2Vec / BERT) and term frequency (TF-IDF) for weighted calculation. The formula for the improved cosine similarity is shown in Equation 2 below:

[0102] Simtext = α·SimTF-IDF + β·SimWord2Vec Equation 2;

[0103] Wherein, Simtext is the first similarity between two text fields, α and β are weight coefficients, and the values of α and β usually range from 0 to 1, and α + β = 1. Through the optimization of training data, the specific values of α and β can be determined. Optionally, α = 0.6 and β = 0.4. SimTF-IDF is the literal semantic similarity calculated based on the Term Frequency-Inverse Document Frequency (TF-IDF) method. The TF-IDF method evaluates the importance of a term by considering the frequency of the term in a document and the rarity of the term in the entire document collection. Based on these importance weights, the literal semantic similarity between two texts can be calculated. SimWord2Vec is the text similarity calculated based on the word vector (such as Word2Vec) method. The word vector method maps terms to a high-dimensional space through training, so that terms that are semantically similar are closer in space. Based on these word vectors, the similarity between two texts can be calculated.

[0104] Taking the text type field of the project name field as an example, assume there are two project names A: "Construction of Medical Security Information Platform" and project name B: "Informatization Project of Smart Medical Platform". First, their TF-IDF similarity (such as 0.7) and word vector similarity (such as 0.8 obtained using BERT) can be calculated. Then, according to the above formula and the determined weight values, the final text similarity of these two project names can be calculated (such as 0.6×0.7 + 0.4×0.8 = 0.74).

[0105] Determining the second similarity based on the numerical difference of the numerical fields in the bid text pair includes:

[0106] Determining the relative error rate of the numerical values of the bid text pair in the numerical field as the second similarity.

[0107] It should be noted that for the numerical fields in the bid text pair, the relative error rate between two numerical values needs to be calculated to measure the degree of difference between them in terms of numerical values. The relative error rate is obtained by calculating the ratio of the absolute value of the difference between two numerical values to the larger value among them. The smaller this ratio is, the smaller the difference between the two numerical values and the higher the similarity. After calculating the relative error rate of the numerical field, this relative error rate can be directly used as the second similarity of the bid text pair in the numerical field.

[0108] In specific implementation, for numerical type fields (such as budget amount field, winning bid amount field), the relative error rate is used to calculate the similarity, and the calculation formula of the relative error rate is as shown in formula 3 below:

[0109]

[0110] Example: The budget A for bid text A is 5,000,000 yuan; the budget B for bid text B is 5,200,000 yuan. Then the second similarity between bid text A and bid text B in the numerical field is 1 - 200,000 / 5,200,000 = 0.96.

[0111] In this embodiment, by calculating the word vector similarity of the bid text pair in the text type field and combining the weights assigned according to the field importance, the similarity degree of the text semantics can be accurately evaluated; at the same time, the second similarity is determined by calculating the relative error rate of the numerical values in the numerical field. This comprehensive evaluation method not only considers the semantic matching of the text content but also takes into account the accurate comparison of numerical data, thus improving the accuracy of the similarity evaluation of the bid text pair.

[0112] Based on any of the above embodiments, before determining the third similarity according to the relationship between the time difference of the time field in the bid text pair and the size of the target time window, it further includes:

[0113] Based on the target field of each bid text in the bid text pair, determine the project type of each bid text.

[0114] It should be noted that the target field is not limited to a single field. It can be a combination of any one or more fields such as the project name field, the project code field, or other fields that can describe the project characteristics.

[0115] The project type can include engineering construction, equipment procurement, service consulting, and so on.

[0116] As an example, use a preset project type library or rule set to match the identified target field to determine the project type to which each bid text belongs. Among them, the project type library or rule set contains the definitions and identification criteria of various project types. For example: A project type library is preset, and the library defines various project types such as engineering construction projects and equipment procurement projects. When processing a bid text, first, the target fields in the bid text need to be identified, such as the project name field "XX Road Bridge Construction Project" and the project code field "GJ - 2023 - 001". Then, these fields are matched with the definitions in the project type library, and it is determined that the project name "XX Road Bridge Construction Project" and the project code "GJ - 2023 - 001" are likely to match the definition of engineering construction projects, so this bid text is classified as an engineering construction project.

[0117] According to the project type of each bid text, determine the project type to which the bid text pair belongs.

[0118] It should be noted that after determining the project type of each caption, a set of judgment logics need to be applied to determine whether the caption pairs belong to the same project type. This judgment process can be a simple direct match, that is, the project types of the two captions are exactly the same; it can also be a more complex logical judgment. For example, although the project types of the two captions are different, they belong to the same major category or have a certain specific association. When the project types of the two captions are exactly the same, directly determine that the project type to which the caption pair belongs is this type. When the project types of the two captions are not completely the same but belong to a broader category or are related to each other, a set of association matching logics need to be used to determine the project type to which the caption pair belongs. For example, one caption is "engineering construction category" (such as a bridge construction project), and the other caption is "engineering service category" (such as an engineering supervision service project). Although they are not completely the same, they both belong to engineering-related projects. At this time, a broader category (such as "engineering-related category") can be defined to cover these two captions. However, if the project types of the two captions are neither the same nor have any association, then the project type to which the caption pair belongs is empty.

[0119] In this embodiment, by identifying the project type of each caption according to the target field of the caption and accordingly determining the project type to which the caption pair belongs, the accuracy of caption classification is effectively improved, laying a foundation for subsequent similarity evaluation based on time difference and target time window.

[0120] Based on any of the above embodiments, the method further includes:

[0121] If the project types of each of the captions are the same, then determine the project type of any caption in the caption pair as the project type to which the caption pair belongs;

[0122] If the project types of each of the captions are not the same, then determine the preset initial time window as the target time window.

[0123] It should be noted that if the item types of the bid texts are the same, then the common item type is determined as the type to which the bid text pair belongs. However, when the item types of the bid text pair cannot be determined (for example, the item types of the two bid texts are inconsistent, or the target field of any one bid text is missing, resulting in the inability to extract the item type), then the target time window cannot be set according to the item type to calculate the third similarity. To address this situation, the concept of the initial time window is introduced. According to historical data, for short-cycle projects such as equipment procurement, the bid text life cycle usually does not exceed 30 days, so the target time window can be set to 30 days. For long-cycle projects such as engineering construction, the bid text life cycle is mostly within 90 days, so the target time window can be set to 90 days. If the item type is unknown, whether it is due to inconsistent item types of the two bid texts or due to the missing target field, the target time window is default set to 60 days to ensure flexibility and adaptability in processing.

[0124] In addition, after tracing the sources of multiple bid texts for the same project, item marking and the generation of a traceability chain can also be carried out. The specific steps include: Generation of unique identifiers: For each data cluster determined by similarity comparison, a unique project number is assigned. The coding rule of this project number incorporates the regional code, year, serial number, and hash check code, ensuring its uniqueness and accuracy. For example, the project number "BJ-2023-001-5A8F" represents the first project in Beijing in 2023, and "5A8F" as the hash check code is used to prevent duplicate numbers. Construction of the associated data chain: Arrange the bid text data according to the time sequence or business logic to generate a clear associated data chain. This data chain will cover all stages of the project, such as tendering, winning the bid, and signing the contract, providing support for the full life cycle management of the project.

[0125] Exemplarily, refer to Figure 2 , Figure 2 which is the overall process schematic diagram of a bid text traceability method provided by an embodiment of the present application. In the scenario of determining whether two bid texts belong to the same project, the following bid texts are used as examples for illustration:

[0126] Bid text A includes the tender number "HLJ-2023-001", project name "Construction of the Smart Medical Platform in Heilongjiang Province", tendering unit "Heilongjiang Provincial Health Commission", budget amount of 5,000,000 yuan, and release time of 2023-01-10.

[0127] The tender number of Bid Document B is "HLJ-2023-001-A", the project name is "Heilongjiang Provincial Medical Informatization Platform Project", the tendering unit is "Heilongjiang Provincial Health Commission" (matching the abbreviation and full name of A), the winning bidder is xxxx Technology Company, the budget amount is 5,200,000 yuan, and the release time is January 25, 2023.

[0128] The tender number of Bid Document C is the same as that of A, the project name is the same as that of B, the winning bidder and the budget amount are also the same as those of B, and the release time is January 28, 2023.

[0129] In the comparison process, first calculate the similarity of each field, including the tender number (0.8, the suffix "-A" is regarded as a sub-project), the project name (0.9, high semantic similarity), the tendering unit (1.0), the budget amount (0.96), and the release time (1.0, the time difference is 15 days, less than 30 days of the target time window). Subsequently, calculate the comprehensive similarity according to the preset weights (0.4×0.8 + 0.3×0.9 + 0.2×1.0 + 0.1×0.96 + 0.05×1.0 = 0.936). Optionally, set the similarity threshold to 0.85. Since the comprehensive similarity is greater than or equal to the similarity threshold of 0.85, it is determined that these three bid documents belong to the same project, and they are merged into the same data cluster and marked with the same project number "HLJ-2023-001-8C2F". Finally, an associated data chain is constructed according to the time sequence or business logic, that is, tender announcement (January 10, 2023) → winning bid announcement (January 25, 2023) → contract announcement (January 28, 2023). Through the above steps, the preprocessing, feature extraction, comparison and classification of project data, and the generation of the project team after tracing the bid document data are realized. This process improves the efficiency of data processing and project integration, reduces the errors caused by human intervention, shortens the time from data inflow to the final project output, and improves the data governance ability.

[0130] In addition, density clustering (DBSCAN) can be used to replace hierarchical clustering to complete the clustering of bid documents of the same project. The replacement logic is to automatically divide clusters according to data density without presetting a similarity threshold. Specifically, parameter settings: neighborhood radius ε = 0.1, minimum number of samples = 5 → output dense data clusters. The advantage of using density clustering is that it does not rely on the hierarchical merging strategy, but uses density drive to adapt to datasets with non-uniform distributions.

[0131] In addition, graph clustering can be used to cluster target texts of the same project. Its alternative logic is to construct the caption data as a graph (nodes = data, edges = similarity), and use a community discovery algorithm (such as Louvain) for classification. Specifically, edge weight = comprehensive similarity → the Louvain algorithm outputs the community division result. The advantage of using graph clustering to cluster target texts of the same project is that it does not rely on numerical similarity calculation, but instead uses graph theory methods, which can identify complex association relationships (such as cross-regional projects).

[0132] In this embodiment, the time window size is determined according to the project type, and the initial time window is used as the evaluation benchmark when the project types are inconsistent, so as to ensure reasonable and accurate similarity evaluation results in different situations, effectively improving the accuracy and efficiency of the evaluation.

[0133] Based on the method described in any of the above embodiments, the present application also provides a structural schematic diagram of an electronic device as shown in Figure 3 . As shown in Figure 3 , at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the method described in any of the above embodiments.

[0134] Based on the method described in any of the above embodiments, the present application also provides a computer storage medium storing a computer program, which can be used to execute the method described in any of the above embodiments when executed by a processor.

[0135] Based on the method described in any of the above embodiments, the present application also provides a computer program product including one or more computer programs or instructions. The computer program or instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. When the computer program is executed by a processor, it implements the method described in any of the above embodiments.

[0136] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0137] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0138] If the described functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0139] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0140] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all of them should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0141] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.

Claims

1. A method for tracing the origin of a mark, characterized in that: The method comprises: Filtering out bid pairs with the same keywords of the project name from the bid library; the bid pairs include at least two bids, and each bid in the bid pair includes a text type field, a numerical type field, and a time type field; A first similarity is determined based on the semantic similarity of the words of the tag pair in the text type field, a second similarity is determined based on the numerical difference of the numerical field in the tag pair, and a third similarity is determined based on the relationship between the time difference of the time field in the tag pair and the size of a target time window; the target time window is determined based on the project type to which the tag pair belongs; Based on the first similarity, the second similarity, and the third similarity, a comprehensive similarity of the pair of labels is determined, and when the comprehensive similarity exceeds a preset similarity threshold, it is determined that the pair of labels belong to the same project.

2. The method according to claim 1, characterized in that Before selecting the bid pairs with the same project name keywords from the bid library, the method further includes: Perform word segmentation and part-of-speech tagging on the project name of each bid in the bid database to obtain the word segmentation included in each bid and the part-of-speech tag corresponding to the word segmentation; For each of the tags, according to the weight corresponding to each part-of-speech tagging result, the project name keywords are determined from all the participles included in each of the tags.

3. The method according to claim 1, characterized in that Before selecting the bid pairs with the same project name keywords from the bid library, the method further includes: For each structured content in the mark text, a field is extracted using a preset regular expression to obtain a first field extraction result; the first field extraction result includes any one or more of a text type field result, a numerical type field result, and a time type field result; For each piece of unstructured content in the mark text, extract semantic information of the unstructured content, and determine a second field extraction result according to the semantic information; the second field extraction result includes any one or more of a text type field result, a numerical type field result, and a time type field result; The text type field, the value type field, and the time type field included in each of the tags are determined according to the first field extraction result and the second field extraction result.

4. The method according to claim 3, characterized in that The determining, according to the first field extraction result and the second field extraction result, the text type field, the value type field, and the time type field included in each of the mark texts comprises: Correcting abnormal values ​​in the first field extraction result and the second field extraction result, and / or filling missing values ​​in the first field extraction result and the second field extraction result to obtain a processed extraction result; The text type field, the value type field, and the time type field are determined according to the processed extraction result.

5. The method according to claim 1, characterized in that The determining of the first similarity based on the semantic similarity of the words of the tag pair in the text type field includes: Determine the text semantic similarity according to the word vector similarity of the tag pair on the text type field, determine the weight allocation value corresponding to the text type field according to the importance of the text type field, and determine the first similarity based on the text semantic similarity and the weight allocation value; The determining the second similarity based on the numerical difference of the numerical fields in the token pair includes: The relative error rate of the numerical values ​​of the pair of tokens in the numerical field is determined as the second similarity.

6. The method according to claim 1, characterized in that Before determining the third similarity based on the time difference of the time field in the token pair and the size relationship of the target time window, the method further includes: Determining the project type of each of the bids based on the target field of each bid in the bid pair; According to the project type of each of the bids, the project type to which the bid pair belongs is determined.

7. The method according to claim 6, characterized in that The method further comprises: If the project types of each of the bids are consistent, the project type of any bid in the bid pair is determined as the project type to which the bid pair belongs; If the item types of each of the bids are inconsistent, the preset initial time window is determined as the target time window.

8. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing processor-executable instructions; Wherein, when the processor calls the executable instruction, the method described in any one of claims 1-7 is implemented.

9. A computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the steps of any method described in claims 1-7 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Scientific literature extraction task-oriented field tracing method and system

    CN122019761A