A method, apparatus, electronic device, and storage medium for processing bidding and tendering data
By identifying and processing different types of bidding documents, and adopting highly targeted data extraction methods, the problem of insufficient data extraction accuracy and reliability in the prior art is solved, and more efficient and accurate data processing is achieved.
Patent Information
- Application Number
- CN202510140288.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-02-08
AI Technical Summary
The existing bidding data processing technology based on natural language processing is difficult to cope with complex file formats and variable content structures, which affects the accuracy and reliability of data extraction.
By receiving bidding documents and identifying their file types, data extraction methods corresponding to different types of files are adopted, including identifying basic information, determining the development status of the publishing unit, and performing multiple extractions and data inspections based on the extraction weights and extraction modes to ensure the accuracy and completeness of the data.
It improves the accuracy and efficiency of the bidding data processing system, ensures the accuracy and reliability of data extraction, and provides a scientific basis for decision-making.
Smart Images

Figure CN119576874B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and particularly to a method, device, electronic device and storage medium for processing bidding data. Background Art
[0002] The processing of bidding data is an important link in modern enterprise management and government procurement. With the improvement of the informatization level, the data processing in the bidding process has become increasingly complex. The traditional manual processing method is not only inefficient but also error-prone, and cannot meet the large-scale and high-frequency bidding requirements. In recent years, with the development of big data and artificial intelligence technologies, automated bidding data processing systems have gradually become a research hotspot. These systems can efficiently complete tasks such as receiving, classifying, data extraction and analysis of bidding documents, greatly improving work efficiency and accuracy.
[0003] In the existing bidding data processing technologies, one of the common methods is the technology based on natural language processing. Specifically, through a pre-trained language model, the bidding documents can be deeply parsed to extract key information and generate structured data. However, the existing bidding data processing technologies based on natural language processing often rely on fixed rules or simple feature matching, and it is difficult to cope with complex file formats and changing content structures. As a result, in practical applications, the accuracy and reliability of data extraction in the file submission process are affected. Therefore, how to improve the accuracy of the bidding data processing system has become an urgent technical problem to be solved. Summary of the Invention
[0004] In order to improve the accuracy of the bidding data processing system, the present application provides a method, device, electronic device and storage medium for processing bidding data.
[0005] In a first aspect, the present application provides a method for processing bidding data, adopting the following technical solution:
[0006] A method for processing bidding data includes:
[0007] Receiving a current bidding document and determining the file type corresponding to the current bidding document;
[0008] Based on the file type, performing data extraction on the current bidding document to obtain the bidding data corresponding to the current bidding document;
[0009] Analyzing the bidding data to obtain an analysis result corresponding to the current bidding document;
[0010] Based on the analysis result, generating an analysis report corresponding to the current bidding document.
[0011] By adopting the above technical solution, by receiving the bidding documents and identifying the file types of the bidding documents, different extraction methods are used for different types of documents to extract data, so as to obtain the bidding data corresponding to the bidding documents, which can improve the efficiency and accuracy of data processing to a certain extent. Then, through in-depth analysis of the bidding data, key information is mined to provide a basis for decision-making, and finally an analysis report is generated, thereby improving the accuracy of the bidding data processing system.
[0012] In a possible implementation, determining the file type corresponding to the current bidding document includes:
[0013] Identifying the basic information corresponding to the current bidding document, where the basic information includes the issuing unit, the file length, and the file keywords. Among them, the file keywords include time nodes, project budgets, and project requirements;
[0014] Determining the development status of the issuing unit corresponding to the current bidding document, where the development status is mature or immature;
[0015] If the development status of the issuing unit corresponding to the current bidding document is mature, determine whether the file length exceeds the length threshold; if the file length exceeds the length threshold, determine that the file type corresponding to the bidding document is the first type; if the file length does not exceed the length threshold, determine that the file type corresponding to the bidding document is the second type;
[0016] If the development status of the issuing unit corresponding to the current bidding document is immature, based on the identified file keywords, determine whether there is inconsistency in the keyword content in the current bidding document; if there is no inconsistency in the keyword content in the current bidding document, determine that the file type corresponding to the current bidding document is the third type; if there is inconsistency in the keyword content in the current bidding document, determine that the file type corresponding to the current bidding document is the fourth type.
[0017] By adopting the above technical solution, by identifying basic information such as the issuing unit, file length, and keywords, it provides an important basis for the preliminary judgment of the file type. Considering the development status of the issuing unit and adopting different judgment logics for units with different maturity levels, it not only considers the diversity of the actual situation but also improves the rationality of the judgment; furthermore, for mature units, the file types are distinguished according to whether the file length exceeds the threshold, which is convenient for quickly identifying the file scale and complexity; while for immature units, the file types are determined by checking the consistency of the keyword content, which helps to identify potential information inconsistency problems and ensure the accuracy and efficiency of file processing.
[0018] In a possible implementation, when the file type corresponding to the current bidding document is the first type, data extraction is performed on the current bidding document based on the file type to obtain the bidding data corresponding to the current bidding document, including:
[0019] Determine the extraction weight corresponding to the file type, and determine the extraction mode corresponding to the file type based on the extraction weight. The extraction mode includes an extraction model and the number of extraction times corresponding to each extraction model;
[0020] Extract the current bidding document based on each extraction model and according to the number of extraction times corresponding to each extraction model to obtain the initial data corresponding to the current bidding document;
[0021] Perform data inspection on the initial data to obtain an inspection result, where the inspection result includes the data accuracy;
[0022] Perform iterative data extraction and data inspection on the initial data based on the inspection result until a preset condition is met. The initial data that meets the preset condition is used as the bidding data corresponding to the current bidding document. The preset condition is that the number of iterations reaches the iteration number threshold or the data accuracy is greater than the accuracy threshold.
[0023] By adopting the above technical solution, by determining the extraction weight and extraction mode corresponding to the file type, it is possible to accurately match the extraction strategy most suitable for this type of file, including selecting a suitable extraction model and setting a reasonable number of extraction times, thereby improving the pertinence and efficiency of data extraction. By adopting the method of multiple extractions accompanied by data inspection, the accuracy and integrity of the extracted data are ensured. Through the process of iterative extraction and data inspection, the data quality is continuously optimized until the preset condition is met (the number of iterations reaches the threshold or the data accuracy exceeds the threshold). The finally obtained bidding data is both comprehensive and reliable, providing a solid foundation for subsequent analysis.
[0024] In a possible implementation, the extraction models include the Go language model, the CNN language model, and the RNN language model. The extracting the current bidding document based on each extraction model and according to the number of extraction times corresponding to each extraction model to obtain the initial data corresponding to the current bidding document includes:
[0025] Preprocess the current bidding document to obtain the preprocessed current bidding document;
[0026] Determine the extraction sequence corresponding to the current bidding document based on the extraction times corresponding to each extraction model. The extraction sequence includes at least two extraction model groups arranged in sequence. Each extraction model group includes at least one extraction model, and the extraction models in each extraction model group are arranged in the extraction order;
[0027] Determine the current extraction model group, perform the extraction step on the current bidding document, and obtain the initial data corresponding to the current bidding document;
[0028] The extraction step includes:
[0029] Determine whether the extraction times of the Go language model in the current extraction model group is 0. If the extraction times of the Go language model is not 0, input the preprocessed current bidding document into the Go language model, and obtain the text data output by the Go language model;
[0030] Determine whether the extraction times of the CNN language model in the current extraction model group is 0. If the extraction times of the CNN language model is not 0, input the text data into the CNN language model, and obtain the first key data and local features output by the CNN language model;
[0031] Determine whether the extraction times of the RNN language model in the current extraction model group is 0. If the extraction times of the RNN language model is not 0, input the text data into the RNN language model, and obtain the second key data and sequence features output by the RNN language model;
[0032] When the extraction times of the RNN language model in the current extraction model group is not 0 and the extraction times of the CNN language model is not 0, splice the local features and the sequence features to obtain combined features, and extract the data corresponding to the combined features to obtain the third key data.
[0033] By adopting the above technical solution, through extracting sequences and combining the advantages of different models, in-depth and multi-angle analysis of the content of bidding documents is achieved. Among them, the Go language model may be good at processing structured or semi-structured data, the CNN model can effectively capture local features, while the deep learning model is good at processing sequence data, extracting key information and temporal features; especially when the CNN and the deep learning model are applied simultaneously, the combined features obtained by splicing local features and sequence features further enrich the extracted data dimensions, improving the comprehensiveness and depth of the data; in addition, the extraction process is flexibly adjusted according to the extraction times set for each model, ensuring the sufficiency and efficiency of data extraction. Finally, the key data output by each model is integrated to form a high-quality and multi-dimensional initial data set, providing strong support for subsequent data analysis and report generation.
[0034] In a possible implementation manner, when the file type corresponding to the current bidding document is the second type, based on the file type, data extraction is performed on the current bidding document to obtain the bidding data corresponding to the current bidding document, including:
[0035] The current bidding document is respectively input into the Embedding model and the sparse coding model, and the text vector output by the Embedding model and the feature representation output by the sparse coding model are obtained;
[0036] The feature vector constructed based on artificial features corresponding to the decision tree model is obtained, and the text vector is spliced with the feature vector to form a first comprehensive feature vector, and the feature representation is spliced with the feature vector to form a second comprehensive feature vector;
[0037] The first comprehensive feature vector and the second comprehensive feature vector are respectively input into the decision tree model, and the first extraction data and the second extraction data output by the decision tree model are obtained;
[0038] Based on the first extraction data and the second extraction data, the bidding data corresponding to the current bidding document is determined.
[0039] By adopting the above technical solutions, the bidding documents are processed by the Embedding model and the sparse coding model respectively. The former generates text vectors rich in semantic information, and the latter provides a sparse feature representation of the document content. The two capture the core information of the document from different perspectives. Combining these two automatically extracted features with the feature vectors constructed based on manual experience forms the first comprehensive feature vector and the second comprehensive feature vector, thereby enhancing the expression ability and robustness of the features. Finally, these two comprehensive feature vectors are input into the decision tree model for prediction to obtain the first extracted data and the second extracted data respectively. By integrating the information of both, the final data of the bidding documents is determined, which not only improves the accuracy and comprehensiveness of data extraction, but also realizes the in-depth understanding and efficient processing of the content of the bidding documents by combining automatic feature extraction and manual feature engineering.
[0040] In a possible implementation manner, when the file type corresponding to the current bidding document is the third type, data extraction is performed on the current bidding document based on the file type to obtain the bidding data corresponding to the current bidding document, including:
[0041] The current bidding document is sequentially segmented into multiple sequence segments, and the sequence segments are sequentially input into the deep learning model in order;
[0042] Obtain the intermediate features corresponding to each sequence segment output by the deep learning model;
[0043] Calculate the correlation score corresponding to each intermediate feature as the attention weight corresponding to each intermediate feature;
[0044] Perform weighted summation based on the attention weight corresponding to each intermediate feature to obtain the weighted feature, and identify the weighted feature to obtain the bidding data corresponding to the current bidding document.
[0045] By adopting the above technical solutions, the bidding document is segmented into multiple sequence segments and sequentially input into the deep learning model, effectively capturing the temporal dependence relationship of the document content. Calculating the correlation score of each intermediate feature and using it as the attention weight enhances the model's attention to key information and improves the pertinence of feature extraction. Finally, weighted summation and identification of the intermediate features based on the attention weight not only integrate the key information in the document, but also further purify the bidding data, ensuring the accuracy and practicality of the data.
[0046] In a possible implementation manner, analyzing the bidding data to obtain the analysis result corresponding to the current bidding document includes:
[0047] Obtain historical bidding data, where the historical bidding data includes the first historical bidding data of the current enterprise corresponding to the current bidding document and the second historical bidding data corresponding to the bidding data;
[0048] Based on the historical bidding moments corresponding to the first historical bidding data and the second historical bidding data respectively, perform time series analysis on the historical bidding data to obtain historical change patterns;
[0049] Based on the historical change patterns, determine the characteristic sub-data corresponding to the bidding data, and determine the correlation degree between every two pieces of characteristic sub-data;
[0050] Based on the characteristic sub-data and the correlation degree between every two pieces of characteristic sub-data, obtain the analysis result corresponding to the current bidding document.
[0051] By adopting the above technical solution, by integrating the first historical bidding data of the current enterprise and the second historical bidding data associated with the bidding data, it provides a comprehensive and rich historical reference for analysis; then, using time series analysis to reveal the change patterns of these historical data not only captures the trends in the time dimension but also deepens the understanding of the dynamic characteristics of the bidding activities; then, based on the historical change patterns, identify the key characteristic sub-data and the correlation degree between them, and the analysis result obtained by synthesizing the characteristic sub-data and their correlation degree not only accurately reflects the core situation of the current bidding document but also provides a scientific basis for decision-making, effectively improving the efficiency and success rate of the bidding activities.
[0052] In a second aspect, the present application provides a bidding data processing device, adopting the following technical solution:
[0053] A bidding data processing device includes:
[0054] A receiving module, configured to receive a current bidding document and determine the file type corresponding to the current bidding document;
[0055] An extraction module, configured to perform data extraction on the current bidding document based on the file type to obtain the bidding data corresponding to the current bidding document;
[0056] An analysis module, configured to analyze the bidding data to obtain the analysis result corresponding to the current bidding document;
[0057] A generation module, configured to generate an analysis report corresponding to the current bidding document based on the analysis result.
[0058] In a third aspect, the present application provides an electronic device, adopting the following technical solution:
[0059] An electronic device, the electronic device comprising:
[0060] At least one processor;
[0061] A memory;
[0062] At least one application, wherein the at least one application is stored in the memory and is configured to be executed by the at least one processor, and the at least one application is configured to: execute the bidding data processing method described in the first aspect above.
[0063] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution:
[0064] A computer-readable storage medium, comprising: a computer program stored with the ability to be loaded and executed by a processor for the bidding data processing method described in the first aspect above.
[0065] In summary, the present application includes the following beneficial technical effects:
[0066] By receiving a bidding document and identifying the file type of the bidding document, different extraction methods are used for different types of files for data extraction to obtain the bidding data corresponding to the bidding document, which can improve the efficiency and accuracy of data processing to a certain extent. Then, by deeply analyzing the bidding data, key information is mined to provide a basis for decision-making, and finally an analysis report is generated, thereby improving the accuracy of the bidding data processing system. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a schematic flowchart of a bidding data processing method provided by an embodiment of the present application;
[0068] Figure 2 is a schematic block diagram of a bidding data processing device provided by an embodiment of the present application;
[0069] Figure 3 is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] The following is a further detailed description of the present application in conjunction with the attached Figure 1 - attached Figure 3 to further illustrate the present application in detail.
[0071] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0072] An embodiment of this application provides a method for processing bidding data. As Figure 1 shown, in the method provided in the embodiment of this application, it is executed by an electronic device, which can be a server or a terminal device. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here. The method includes steps S101 to S104, where:
[0073] Step S101: Receive the current bidding document and determine the file type corresponding to the current bidding document.
[0074] Among them, the file type can include conventional bidding in traditional industries, innovative bidding in emerging industries, bidding for small and simple projects, and emergency repair bidding.
[0075] When receiving a bidding document to be processed, use the bidding document to be processed as the current bidding document, and extract the bidding data corresponding to the current bidding document. To improve the accuracy of the extracted bidding data, the file type of the current bidding document can be determined, and data extraction can be performed based on the file type corresponding to the current bidding document. Specifically, a document parsing tool, such as an open-source library like Apache Tika, can be used to deeply analyze the format structure of the file to identify whether the file has a standardized chapter title (such as hierarchical divisions like 1., 1.1, 1.1.1, etc.), whether the table of contents structure is complete, and whether the font size layout is unified, and preliminarily judge the standardization of the file. At the same time, combined with natural language processing technology, a professional industry keyword library is constructed, covering many fields such as construction, electronics, medicine, and finance, and the frequency of occurrence and distribution of keywords in the file are statistically analyzed. For example, if words such as "concrete", "construction drawings", and "construction qualifications" frequently appear and the file format is standardized, it is very likely to be a conventional bidding document in the construction industry; if a large number of emerging technology words such as "blockchain", "artificial intelligence algorithm", and "gene editing" emerge and the text structure is relatively flexible, it may be an innovation tender for emerging industries. The text logic analysis algorithm can also be used to detect whether the file content is described in the conventional logical order such as project background, bidding requirements, technical specifications, commercial terms, and evaluation criteria, and determine the file type based on comprehensive information from multiple aspects.
[0076] Step S102: Based on the file type, perform data extraction on the current bidding document to obtain the bidding data corresponding to the current bidding document.
[0077] For conventional bidding documents in traditional industries: A rule-based extraction method can be adopted, and a precise matching pattern can be constructed using regular expressions. Taking the extraction of the project budget as an example, set a regular expression such as "Budget: ([\d,.]+) yuan", and the amount information such as "Budget: 500,000 yuan" can be accurately captured; for the qualification requirements, according to the qualification standard library of each industry, write rules to match the qualification level, certificate validity period, etc. of the bidding unit, so as to obtain the bidding data corresponding to the current bidding document.
[0078] For innovation bidding documents in emerging industries: First, a pre-trained model based on the Transformer architecture, such as a fine-tuned version of the GPT series, can be used to convert the file content into a semantically rich vector representation to enable the model to understand the text meaning. Then, combined with a domain knowledge graph, which contains knowledge such as professional term explanations and technical association relationships in emerging industries, to assist the model in identifying key information. For example, in the bidding for the field of quantum computing, the model understands terms such as "quantum bit" and "quantum gate" based on the knowledge graph and extracts unique information such as special quantum algorithm performance requirements and experimental environment parameters, so as to obtain the bidding data corresponding to the current bidding document.
[0079] For the bidding documents of small and simple projects: Such documents usually have concise content and concentrated key information. Therefore, lightweight text processing tools can be used to first remove redundant information in the documents, such as excessive modifiers, repeated paragraphs, etc., and streamline the text. Then, a simple keyword matching algorithm can be adopted. For common core information (such as the name of the purchased item, quantity, quotation deadline), a keyword list is set, such as "purchased item", "quantity:", "quotation deadline", etc., to quickly locate and extract key information, supplemented by a small amount of manual spot checks to ensure the reliability of the extracted information, so as to meet the requirements of rapid processing of small projects, and thus obtain the bidding data corresponding to the current bidding document.
[0080] For the bidding documents of emergency repair projects: Since such documents often have chaotic formats and urgent information expressions. Therefore, text preprocessing techniques can be used first, such as using an automatic error correction tool to correct typos and unify punctuation marks, and using a text sorting algorithm to reorder the chaotic paragraphs. Then, combined with multiple basic models for step-by-step extraction, first use a rule model to capture obvious time (such as the start time of emergency repair, the bidding deadline), amount (such as the emergency repair budget), and other key numerical information, and then use a shallow machine learning model, such as a Naive Bayes classifier, to identify key text fragments (such as the description of the technical points of emergency repair), so as to obtain the bidding data corresponding to the current bidding document.
[0081] Step S103: Analyze the bidding data to obtain the analysis result corresponding to the current bidding document.
[0082] In this embodiment, historical bidding data is obtained. The historical bidding data includes the first historical bidding data of the current enterprise corresponding to the current bidding document and the second historical bidding data corresponding to the bidding data; based on the historical bidding moments corresponding to the first historical bidding data and the second historical bidding data respectively, perform time series analysis on the historical bidding data to obtain historical change rules; based on the historical change rules, determine the characteristic sub-data corresponding to the bidding data, and determine the correlation degree between every two characteristic sub-data; based on the characteristic sub-data and the correlation degree between every two characteristic sub-data, obtain the analysis result corresponding to the current bidding document.
[0083] Among them, the second historical bidding data is other historical bidding data that is relevant to the current bidding data in terms of project type and industry.
[0084] Specifically, the first historical bidding and tendering data corresponding to the current bidding and tendering documents and the second historical bidding data corresponding to the bidding and tendering data can be directly obtained from the database corresponding to the current bidding and tendering documents. Further, the obtained first historical bidding and tendering data and second historical bidding data are preprocessed so that the time information of each data is uniformly formatted into a standard timestamp or date format for subsequent analysis.
[0085] Specifically, for numerical bidding and tendering data, time series analysis tools can be used. By plotting time series graphs and combining statistical methods such as moving average method and exponential smoothing method, the time series is smoothed to remove short-term random fluctuations, highlight the long-term trend, and by plotting a time series graph with time as the horizontal axis and the bidding amount as the vertical axis, observe the rising, falling or fluctuating trend of the bidding amount to obtain the first historical change law. For non-numerical data, such as winning bid situations (winning or not winning the bid), changes in project types, etc., count the frequencies of various categories at different time points, observe their changes over time, and calculate the proportion of the number of projects winning the bid in different years to the total number of tendered projects, and analyze the change law of this proportion over time to obtain the second historical change law. Further, the first historical change law and the second historical change law are combined to obtain the historical change law.
[0086] Further, after obtaining the historical change law, based on the historical change law and combined with the characteristics of the current bidding and tendering data, it can be determined which data are key characteristic sub-data. Specifically, review the historical change trend obtained from time series analysis. For such trend laws, the current bidding and tendering data closely related to the key driving factors are used as characteristic sub-data. Exemplarily, if historical data shows that the bidding amount of a certain type of project shows an increasing trend year by year, and at the same time it is found that the intensity of industry policy support is synchronized with the growth of the bidding amount, then the policy support factor is the key factor driving the change of the bidding amount.
[0087] For historical data with periodic changes, such as the bidding and tendering activities of some seasonal projects showing specific periodic laws. Suppose the bidding volume of some outdoor projects in the construction industry is relatively large in spring and autumn, and this periodicity is closely related to climatic conditions. Then in the current bidding and tendering data, data related to climatic conditions such as the expected start time of the project and the construction period can be used as characteristic sub-data.
[0088] Furthermore, for every two feature sub - data, if the feature sub - data is numerical, such as project budget and expected profit, the relationship between the two feature sub - data can be calculated by computing the Pearson correlation coefficient between them. The value range of this coefficient is between - 1 and 1. The closer the absolute value is to 1, the stronger the linear correlation between the two variables. If the feature sub - data is non - numerical, such as project type and the region where the winning bidder enterprise is located, the electronic device can use an association rule mining algorithm (such as the Apriori algorithm) to calculate the confidence of the association rules between them to determine the degree of association between the two.
[0089] Furthermore, integrate and visually present the feature sub - data and the degree of association between them. Specifically, the degree of association between different feature sub - data can be shown by drawing a heat map, where the darker the color, the higher the degree of association. At the same time, list the feature sub - data in tabular form and mark their association with other feature sub - data beside them, so as to obtain the analysis result corresponding to the current bidding document.
[0090] Step S104: Generate an analysis report corresponding to the current bidding document based on the analysis result.
[0091] After obtaining the analysis result corresponding to the current bidding document, select a suitable visualization tool to generate an analysis report corresponding to the current bidding document, such as Tableau, PowerBI, etc., and display the data analysis result in an intuitive chart form. For time - series trend analysis, a line chart clearly shows the changing trend; the degree of association between every two feature sub - data is displayed in the form of a network diagram.
[0092] The embodiment of the present application provides a method for processing bidding data. By receiving a bidding document and identifying the file type of the bidding document, different extraction methods are used for different types of files to extract data, so as to obtain the bidding data corresponding to the bidding document, which can improve the efficiency and accuracy of data processing to a certain extent. Then, through in - depth analysis of the bidding data, key information is mined to provide a basis for decision - making, and finally an analysis report is generated, thereby improving the accuracy of the bidding data processing system.
[0093] A possible implementation manner of the embodiment of the present application is that in the above - mentioned step S101, determining the file type corresponding to the current bidding document includes:
[0094] Identifying the basic information corresponding to the current bidding document, where the basic information includes the issuing unit, the length of the document, and the document keywords. Among them, the document keywords include time nodes, project budget, and project requirements;
[0095] Determining the development status of the issuing unit corresponding to the current bidding document, where the development status is either developed or under - developed;
[0096] If the development status of the issuing unit corresponding to the current bidding document is mature development, determine whether the document length exceeds the length threshold; if the document length exceeds the length threshold, determine that the document type corresponding to the bidding document is the first type; if the document length does not exceed the length threshold, determine that the document type corresponding to the bidding document is the second type;
[0097] If the development status of the issuing unit corresponding to the current bidding document is immature, then based on the identified file keywords, it is determined whether there is any keyword content inconsistency in the current bidding document; if there is no keyword content inconsistency in the current bidding document, then the file type corresponding to the current bidding document is determined to be the third type; if there is any keyword content inconsistency in the current bidding document, then the file type corresponding to the current bidding document is determined to be the fourth type.
[0098] Since the capabilities, requirements, and industry information of different publishers are different, the types of bidding documents issued by different publishers may be different. Therefore, text extraction tools can be used to find text with unit identification at the beginning, end, or specific signature area of the file for "publishing unit". For example, regular expressions can be used to match common formats such as company names and full names of government departments, and combined with industrial and commercial registration information databases and government agency directories for verification to ensure accurate identification of the publishing unit. For "file length", if the file is in text format, directly count the number of characters or lines; if it is in PDF format, use open source libraries such as Apache PDFBox to convert the file into text and then count the length. For "file keywords", obtain a pre-built dictionary containing various time keywords, budget keywords ("budget amount", "capital limit", etc.), and project requirement keywords ("qualification requirements", "technical standards", etc.) from the database corresponding to the current bidding document, use keyword matching algorithms to scan the file text, mark and extract the keywords and their contextual information, so as to facilitate subsequent in-depth analysis.
[0099] After obtaining the issuing unit of the current bidding document, the establishment age of the issuing unit can be queried. Entities that have been established for more than a certain number of years (e.g., 10 years), have a relatively high credit rating, and have a rich record of winning bids in past projects are initially determined to be mature in development; conversely, entities that have been established for a relatively short time (e.g., within 3 years), have less credit information, and have no experience in large successful projects are regarded as immature in development. Further, after obtaining the issuing unit, the industry to which the issuing unit belongs can be obtained, and the corresponding length threshold for this industry can be obtained from the database corresponding to the current bidding document. For example, the length threshold for the bidding text of large-scale projects in the construction industry may be set at 5000 characters, and for small-scale decoration projects, it may be 1000 characters. Compare the statistically obtained length of the current document with the corresponding length threshold. If the document length is greater than the set threshold, determine that the document type of the current bidding document is the first type; if it is less than or equal to the corresponding threshold, determine that the document type of the current bidding document is the second type.
[0100] Further, after extracting the document keywords and their contexts, a logical verification model is established for the key information. For the "time node", check whether the bid deadline is later than the release time of the tender announcement, and whether the opening time is within a reasonable range and connected to the bid deadline; for the "project budget", analyze whether the budget amount matches the project scale description, and whether there are unreasonable phenomena such as being too low or too high; in terms of "project requirements", use the semantic understanding technology of natural language processing to determine whether there are self-contradictory expressions in requirements such as qualifications and technologies. For example, using rule-based semantic analysis, set the corresponding rules between the qualification level and the project complexity. When it is determined that there is no inconsistency in the keyword content, determine that the document type corresponding to the current bidding document is the third type. When it is determined that there is an inconsistency in the keyword content, determine that the document type corresponding to the current bidding document is the fourth type.
[0101] In a possible implementation manner of the embodiments of the present application, in the above embodiments, when the document type corresponding to the current bidding document is the first type, based on the document type, data extraction is performed on the current bidding document to obtain the bidding data corresponding to the current bidding document, including:
[0102] Determine the extraction weight corresponding to the document type, and based on the extraction weight, determine the extraction mode corresponding to the document type. The extraction mode includes an extraction model and the number of extractions corresponding to each extraction model;
[0103] Based on each extraction model and according to the number of extractions corresponding to each extraction model, perform extraction on the current bidding document to obtain the initial data corresponding to the current bidding document;
[0104] Perform data inspection on the initial data to obtain an inspection result, and the inspection result includes the accuracy of the data;
[0105] Iteratively extract and check the initial data based on the inspection results until the preset conditions are met. The initial data that meets the preset conditions is used as the tendering data corresponding to the current tendering document. The preset conditions are that the number of iterations reaches the iteration threshold or the data accuracy is greater than the accuracy threshold.
[0106] The documents of the first type come from a mature and standardized field. The file format has a high degree of standardization, the content expression is rigorous and follows a fixed paradigm. For example, in the tendering for infrastructure maintenance projects regularly carried out by government departments, a unified announcement and file template have been formed over a long time. The key information involved (such as project budget, construction period requirements, qualification review standards, etc.) is clear and there are rarely ambiguous or unclear expressions. Therefore, for the documents of the first type, high accuracy and high recall rate need to be ensured. Specifically, obtain the extraction model corresponding to the first type, the extraction weight corresponding to each extraction model, and the total number of extractions from the database corresponding to the current tendering document. Then, based on the number of extractions of the extraction model = extraction weight * total number of extractions, obtain the number of extractions corresponding to each extraction model, and obtain the extraction mode corresponding to the current tendering document. Further, after obtaining the number of extractions corresponding to each extraction model, the current tendering document can be extracted to obtain the initial data corresponding to the current tendering document.
[0107] Specifically, when the extraction models include the Go language model, the CNN language model, and the RNN language model, extract the current tendering document based on each extraction model and according to the number of extractions corresponding to each extraction model. The initial data corresponding to the current tendering document obtained may include: preprocessing the current tendering document to obtain the preprocessed current tendering document; determining the extraction sequence corresponding to the current tendering document based on the number of extractions corresponding to each extraction model. The extraction sequence includes at least two extraction model groups arranged in order. Each extraction model group includes at least one extraction model, and the extraction models in each extraction model group are arranged in the extraction order; determining the current extraction model group and performing the extraction step on the current tendering document to obtain the initial data corresponding to the current tendering document.
[0108] The extraction steps include: determining whether the extraction times of the Go language model in the current extraction model group are 0. If the extraction times of the Go language model are not 0, then input the preprocessed current bidding document into the Go language model and obtain the text data output by the Go language model; determining whether the extraction times of the CNN language model in the current extraction model group are 0. If the extraction times of the CNN language model are not 0, then input the text data into the CNN language model and obtain the first key data and local features output by the CNN language model; determining whether the extraction times of the RNN language model in the current extraction model group are 0. If the extraction times of the RNN language model are not 0, then input the text data into the RNN language model and obtain the second key data and sequence features output by the RNN language model; when the extraction times of the RNN language model and the CNN language model in the current extraction model group are both not 0, splice the local features and sequence features to obtain combined features, and extract the data corresponding to the combined features to obtain the third key data.
[0109] More specifically, a text processing tool can be used to convert the format of the current bidding document into UTF-8 format through an encoding conversion library, and remove the redundant whitespace characters such as spaces, line breaks, and tab characters in the file to make the text compact. For some special symbols, such as garbled characters and non-text elements (which may be the remaining marks during the file conversion process), they are filtered and cleared according to a predefined list of illegal characters to achieve the preprocessing of the current bidding document.
[0110] Furthermore, after obtaining the extraction times corresponding to each extraction model, an extraction sequence corresponding to the current bidding document can be generated, and the preprocessed current bidding document can be extracted according to this extraction sequence to obtain the initial data corresponding to the current bidding document. Specifically, after obtaining the extraction times corresponding to each extraction model, obtain the extraction priority corresponding to each extraction model from the database corresponding to the current bidding document, and sort the extraction times of the extraction models according to this extraction priority to obtain the extraction sequence corresponding to the current bidding document. Exemplarily, the extraction models are model a, model b, and model c respectively, the priority of model a is greater than that of model b which is greater than that of model c, and the extraction times corresponding to model a, model b, and model c are 1, 2, and 3 respectively. Then the extraction sequence is that the first extraction model group is model a, model b, and model c, the second extraction model group is model b and model c, and the third extraction model group is model c.
[0111] Further, in the process of extracting the current bidding document using the extraction sequence, the current bidding document is extracted in the order of the extraction model groups in the extraction sequence. Therefore, the current extraction model group can be obtained, that is, the current extraction model group is determined according to the current extraction times. When the current extraction times is 1, the current extraction model group is the first extraction model group. After obtaining the current extraction model group, when entering the Go language model processing link of the current extraction model group, first read the extraction times configuration of the Go language model in the current extraction model group. If the extraction times of the Go language model is not 0, then input the preprocessed current bidding document into the trained Go language, and process the current bidding document through the built-in string in the Go language model, and obtain the text data output by the Go language model.
[0112] When it is judged that the CNN language model in the current extraction model group needs to perform the extraction operation, first preprocess the previously obtained text data (if it is the first time to use the CNN language model, it is the text of the preprocessed bidding document), and convert it into a format suitable for the input of the CNN language model. Specifically, a pre-trained word embedding model (such as Word2Vec or GloVe) can be used to convert each word into a vector of a fixed dimension, and then multiple word vectors are combined into a matrix. Input this matrix into the constructed CNN language model. Multiple convolutional kernels in the convolutional layer of the CNN language model scan the text matrix in a sliding window manner to extract local pattern features of key terms such as "chip main frequency" and "concrete strength grade" and their contexts. The pooling layer performs dimensionality reduction, and the fully connected layer maps and outputs. Obtain the first key data output by the CNN language model, that is, the identified key terms and related text fragments, such as "Server configuration requirements: The processor main frequency is not less than 3.0GHz, and the memory is 16GB"; at the same time, obtain the local feature vector output by the CNN language model, and this local feature vector records the local semantic features of these key information.
[0113] After confirming that the RNN language model in the current extraction model group needs to perform the extraction task, input the corresponding text data (if there is a Go language model processed before, it is its output text; if the deep learning model is directly started, it is the preprocessed bidding document) into the RNN language model word by word or sentence by sentence in order. The memory unit inside the RNN language model continuously updates the hidden state according to the input order and records the previous information, so as to be able to understand the semantics in combination with the previous text when processing the subsequent text. For example, when the previous text of the document mentions "This project is the construction of an intelligent factory and needs to meet the requirements of automated production", and the subsequent text describes "The production line control system should have functions of real-time monitoring and fault warning, and the response time does not exceed 5 seconds", the RNN language model can accurately grasp that the requirements of the control system here are key information based on the previous context, output the second key data, and at the same time output the sequence feature vector, and the electronic device obtains this second key data and the sequence feature.
[0114] Further, in the current extraction model group, when it is confirmed that both the CNN language model and the RNN language model have completed extraction and the extraction times are not zero, the local feature vectors and sequence feature vectors output by the two are obtained, and they are concatenated into a combined feature vector by dimension using a vector concatenation function. Then, based on the combined feature vector, a simple classifier (such as logistic regression) or a rule-based discrimination method is used to extract the key information text fragment according to the matching degree between the eigenvalue and the known key information pattern, so as to obtain the initial data corresponding to the current bidding document.
[0115] Furthermore, when it is confirmed that the extraction times of the CNN language model or the RNN language model are zero, the first key data or the second key data is used as the initial data corresponding to the current bidding document; when it is confirmed that the extraction times of both the CNN language model and the RNN language model are zero, the text data is used as the initial data corresponding to the current bidding document.
[0116] Furthermore, after obtaining the initial data corresponding to the current extraction model group, the initial data can be checked to determine whether the accuracy of the initial data meets the requirements. Specifically, the initial data is input into a trained rule-based finite state model, and the inspection result including the data accuracy level output by the rule-based finite state model is obtained. Among them, the rule-based finite state model is trained through a large number of data samples and a large number of inspection results of samples including data preparation levels.
[0117] After obtaining the data accuracy level of the initial data corresponding to the current extraction model group, the data accuracy level is compared with the accuracy threshold. If the data accuracy level is greater than the accuracy threshold, the initial data corresponding to the current extraction model group is used as the bidding data corresponding to the current bidding document; if the data accuracy level is not greater than the accuracy threshold, it is determined whether the extraction times corresponding to the current extraction model group reach the iteration times threshold. If the extraction times corresponding to the current extraction model group reach the iteration times threshold, the initial data corresponding to the current extraction model group is used as the bidding data corresponding to the current bidding document; if the extraction times corresponding to the current extraction model group do not reach the iteration times threshold, the extraction times corresponding to the previous extraction model group are incremented by one, and the extraction model group corresponding to the incremented extraction times is obtained as the current extraction model group, and iterative data extraction and data inspection are performed until the initial data that meets the preset conditions is obtained.
[0118] In a possible implementation manner of the embodiment of the present application, when the file type corresponding to the current bidding document is the second type, data extraction is performed on the current bidding document based on the file type to obtain the bidding data corresponding to the current bidding document, including:
[0119] Input the current bidding documents into the Embedding model and the sparse coding model respectively, and obtain the text vectors output by the Embedding model and the feature representations output by the sparse coding model;
[0120] Obtain the feature vectors constructed based on artificial features corresponding to the decision tree model, and concatenate the text vectors with the feature vectors to form the first comprehensive feature vector, and concatenate the feature representations with the feature vectors to form the second comprehensive feature vector;
[0121] Input the first comprehensive feature vector and the second comprehensive feature vector into the decision tree model respectively, and obtain the first extraction data and the second extraction data output by the decision tree model;
[0122] Based on the first extraction data and the second extraction data, determine the bidding data corresponding to the current bidding document.
[0123] The documents corresponding to the second type are those with relatively concise and clear content, concentrated and prominent key information, but with certain information limitations. It is common in simple procurement bids of small enterprises or units, such as the procurement bid for food ingredients in a small restaurant, the procurement bid for community activity supplies, etc. Such documents focus on a few core information, such as the procurement item list, the deadline for quotation, the delivery requirements, etc. The text length is short, and there is less background information and additional terms. Thus, it can be seen that the documents corresponding to the second type focus on ensuring the absolute correctness of the extracted information. Even if some less critical auxiliary information may be omitted, the key information must be error-free. Therefore, the documents corresponding to the second type need to ensure high accuracy. Specifically, load a pre-trained Embedding model (such as a model trained based on the Word2Vec or GloVe algorithm), input the current bidding document line by line or paragraph by paragraph into the pre-trained Embedding model. The pre-trained Embedding model, through the forward propagation process, converts each word in the text into a corresponding fixed-dimensional vector, such as a 300-dimensional vector. Finally, the entire document text is converted into a sequence of text vectors, and the average value is taken or other aggregation methods are used to obtain a text vector representing the overall semantics of the document as the output. The electronic device obtains the text vector output by the pre-trained Embedding model.
[0124] For the sparse coding model (such as using an open-source sparse coding library based on the orthogonal matching pursuit algorithm), the electronic device inputs the current bidding document line by line or paragraph by paragraph into the sparse coding model. The sparse coding model learns the key feature patterns of each part of the text through sparse decomposition of the text and outputs the corresponding sparse coding feature representation in matrix form. The electronic device obtains the feature representation output by the sparse coding model. Among them, the rows in the feature representation represent the features of a text segment, and the columns represent the learned feature dimensions.
[0125] Further, obtain the feature vector constructed based on artificial features corresponding to the decision tree model from the database corresponding to the current bidding document, and use a vector concatenation function (such as the numpy.concatenate function in Python, or the corresponding array concatenation method in Go language) to concatenate the text vector output by the Embedding model and the artificial feature vector dimension by dimension to obtain the first comprehensive feature vector; Similarly, concatenate the feature representation output by the sparse coding model and the artificial feature vector to form the second comprehensive feature vector, providing rich and diverse input information for the subsequent decision tree model.
[0126] Even further, load the trained decision tree model. The electronic device inputs the first comprehensive feature vector into the decision tree model. The decision tree model starts from the root node and, based on the feature segmentation conditions stored in the internal nodes (such as based on the text vector semantic similarity threshold and the artificial feature value range), makes judgments by branching down layer by layer. Finally, the first extraction data is output at the leaf node, and the electronic device obtains the first extraction data output by the decision tree model. Similarly, the electronic device inputs the second comprehensive feature vector into the decision tree model and repeats the above decision process, and the electronic device obtains the second extraction data. Among them, the decision tree model was pre-trained on a large number of second-type bidding document samples and their corresponding annotated key information, and mastered the mapping rule from the comprehensive feature vector to the key information.
[0127] After obtaining the first extraction data and the second extraction data, the first extraction data and the second extraction data can be merged to remove possible duplicate information, thereby obtaining the bidding data corresponding to the current bidding document.
[0128] In a possible implementation manner of the embodiment of the present application, when the file type corresponding to the current bidding document is the third type, data extraction is performed on the current bidding document based on the file type to obtain the bidding data corresponding to the current bidding document, including:
[0129] Sequentially divide the current bidding document into multiple sequence segments, and sequentially input the sequence segments into the deep learning model in order;
[0130] Obtain the intermediate features corresponding to each sequence segment output by the deep learning model;
[0131] Calculate the correlation score corresponding to each intermediate feature as the attention weight corresponding to each intermediate feature;
[0132] Perform weighted summation based on the attention weight corresponding to each intermediate feature to obtain the weighted feature, and identify the weighted feature to obtain the bidding data corresponding to the current bidding document.
[0133] The third type of documents generally come from emerging industries or innovative project biddings. The industry terminology has not yet been unified, the business model is changeable, and the document content is rich, complex and full of uncertainty. Therefore, the third type of documents pay more attention to capturing information as comprehensively as possible. Therefore, the deep learning model (such as the model based on the Transformer architecture) can be used to mine potential key information with its powerful semantic understanding ability and adaptability to complex texts. Even if there are some misjudgments or inaccurate extractions, it is necessary to ensure that important clues are not missed.
[0134] Specifically, the sentence segmenter in the natural language processing tool (such as an open source segmenter based on punctuation and grammatical rules) can be used to process the current bidding document and split it into complete sentence sequence fragments, that is, multiple sequence fragments. After the splitting is completed, these sequence fragments are input into the pre-loaded and initialized deep learning model one by one in a predetermined order. Taking the LSTM model as an example, before inputting the sequence fragment, each word in the sequence fragment is converted into a fixed-dimensional vector (assuming 300 dimensions) with the help of a pre-trained word embedding model (such as Word2Vec or GloVe), and then the word vector sequence is sent to the LSTM model to start the model's deep mining process for text information. The LSTM model gradually processes each sequence fragment based on its own memory unit and gating mechanism, and integrates the context information.
[0135] Furthermore, after the deep learning model completes the forward propagation calculation for each sequence segment, the electronic device obtains the intermediate feature vector output by the model as the intermediate feature corresponding to each sequence segment.
[0136] The query vector corresponding to the current bidding document is obtained from the database corresponding to the current bidding document, and the correlation between each intermediate feature and the query vector is calculated through the similarity measurement algorithm. Taking the cosine similarity calculation as an example, for each intermediate feature vector, it and the query vector are substituted into the cosine similarity calculation formula, and the result is the correlation score. Then, the softmax function is used to normalize all correlation scores to ensure that the sum of all weights is 1, thereby obtaining the attention weight corresponding to each intermediate feature.
[0137] Furthermore, according to the calculated attention weights corresponding to each intermediate feature, a weighted summation operation is performed on the stored intermediate feature sequence. Specifically, using the mathematical formula for vector weighted summation, each intermediate feature vector is multiplied element-wise with the corresponding attention weight, and then all the product results are added according to the vector addition rule to obtain a weighted feature vector that comprehensively reflects the key information of the document. Further, the weighted features can be recognized through a classifier (such as a logistic regression model) to obtain the key data corresponding to the current bidding document, and these extracted key data are regularized according to a pre-designed data classification framework to serve as the bidding data corresponding to the current bidding document.
[0138] Furthermore, when the document type corresponding to the current bidding document is the fourth type, since the documents corresponding to the fourth type are usually some bidding documents generated under temporary, emergency or informal processes, the document quality is uneven, the information expression is chaotic, and there is a lack of order. For example, in the bidding initiated by some small private enterprises in case of sudden needs, there may be no professional legal or bidding team to review, and problems such as missing key information, format errors, and mixed use of terms frequently occur in the documents. To reduce the occurrence of overfitting, basic text preprocessing means can be used, such as using string processing functions to remove garbled characters, extra spaces, and irregular punctuation in the document, and then using an encoding conversion library to unify the document into a standard encoding format to ensure the basic text quality and improve readability. Then, the preprocessed current bidding document is input into a shallow machine learning model trained based on labeled data, such as a decision tree or a naive Bayes model, and the data output by the shallow machine learning model trained based on labeled data is obtained to serve as the bidding data corresponding to the current bidding document.
[0139] The above embodiments introduce a method for processing bidding data from the perspective of the method flow. The following embodiments introduce a device for processing bidding data from the perspective of virtual modules or virtual units. For details, see the following embodiments.
[0140] See Figure 2 , the device 20 for processing bidding data may specifically include: a receiving module 201, an extraction module 202, an analysis module 203, and a generation module 204. Specifically:
[0141] A device 20 for processing bidding data, comprising:
[0142] The receiving module 201 is configured to receive the current bidding document and determine the document type corresponding to the current bidding document;
[0143] The extraction module 202 is configured to extract data from the current bidding document based on the document type to obtain the bidding data corresponding to the current bidding document;
[0144] An analysis module 203, configured to analyze the tendering and bidding data to obtain an analysis result corresponding to the current tendering and bidding document;
[0145] A generation module 204, configured to generate an analysis report corresponding to the current tendering and bidding document based on the analysis result.
[0146] In a possible implementation manner of the embodiment of the present application, when determining the file type corresponding to the current tendering and bidding document, the receiving module 201 is specifically configured to:
[0147] Identify the basic information corresponding to the current tendering and bidding document, where the basic information includes the issuing unit, the file length, and the file keywords. The file keywords include time nodes, project budgets, and project requirements;
[0148] Determine the development status of the issuing unit corresponding to the current tendering and bidding document, where the development status is either developed or underdeveloped;
[0149] If the development status of the issuing unit corresponding to the current tendering and bidding document is developed, determine whether the file length exceeds the length threshold; if the file length exceeds the length threshold, determine that the file type corresponding to the tendering and bidding document is the first type; if the file length does not exceed the length threshold, determine that the file type corresponding to the tendering and bidding document is the second type;
[0150] If the development status of the issuing unit corresponding to the current tendering and bidding document is underdeveloped, determine whether there is inconsistency in the keyword content based on the identified file keywords; if there is no inconsistency in the keyword content in the current tendering and bidding document, determine that the file type corresponding to the current tendering and bidding document is the third type; if there is inconsistency in the keyword content in the current tendering and bidding document, determine that the file type corresponding to the current tendering and bidding document is the fourth type.
[0151] In a possible implementation manner of the embodiment of the present application, when the file type corresponding to the current tendering and bidding document is the first type, when the extraction module 202 extracts data from the current tendering and bidding document based on the file type to obtain the tendering and bidding data corresponding to the current tendering and bidding document, it is specifically configured to:
[0152] Determine the extraction weight corresponding to the file type, and determine the extraction mode corresponding to the file type based on the extraction weight. The extraction mode includes an extraction model and the number of extractions corresponding to each extraction model;
[0153] Extract the current tendering and bidding document based on each extraction model and according to the number of extractions corresponding to each extraction model to obtain the initial data corresponding to the current tendering and bidding document;
[0154] Conduct a data check on the initial data to obtain a check result, where the check result includes the data accuracy;
[0155] Iteratively extract and check the initial data based on the inspection results until a preset condition is met. The initial data that meets the preset condition is used as the tendering and bidding data corresponding to the current tendering and bidding document. The preset condition is that the number of iterations reaches the iteration threshold or the data accuracy is greater than the accuracy threshold.
[0156] In a possible implementation manner of the embodiment of the present application, the extraction model includes a Go language model, a CNN language model, and an RNN language model. When the extraction module 202 extracts the current tendering and bidding document based on each extraction model and according to the extraction times corresponding to each extraction model to obtain the initial data corresponding to the current tendering and bidding document, it is specifically used for:
[0157] Preprocess the current tendering and bidding document to obtain the preprocessed current tendering and bidding document;
[0158] Based on the extraction times corresponding to each extraction model, determine the extraction sequence corresponding to the current tendering and bidding document. The extraction sequence includes at least two extraction model groups arranged in sequence. Each extraction model group includes at least one extraction model, and the extraction models in each extraction model group are arranged in the extraction order;
[0159] Determine the current extraction model group, and perform the extraction step on the current tendering and bidding document to obtain the initial data corresponding to the current tendering and bidding document;
[0160] The extraction step includes:
[0161] Determine whether the extraction times of the Go language model in the current extraction model group is 0. If the extraction times of the Go language model is not 0, input the preprocessed current tendering and bidding document into the Go language model, and obtain the text data output by the Go language model;
[0162] Determine whether the extraction times of the CNN language model in the current extraction model group is 0. If the extraction times of the CNN language model is not 0, input the text data into the CNN language model, and obtain the first key data and local features output by the CNN language model;
[0163] Determine whether the extraction times of the RNN language model in the current extraction model group is 0. If the extraction times of the RNN language model is not 0, input the text data into the RNN language model, and obtain the second key data and sequence features output by the RNN language model;
[0164] When the extraction times of the RNN language model and the CNN language model in the current extraction model group are both not 0, splice the local features and sequence features to obtain combined features, and extract the data corresponding to the combined features to obtain the third key data.
[0165] In a possible implementation manner of the embodiment of the present application, when the file type corresponding to the current bidding document is the second type, when the extraction module 202 performs data extraction on the current bidding document based on the file type to obtain the bidding data corresponding to the current bidding document, it is specifically used for:
[0166] Input the current bidding document into the Embedding model and the sparse coding model respectively, and obtain the text vector output by the Embedding model and the feature representation output by the sparse coding model;
[0167] Obtain the feature vector constructed based on artificial features corresponding to the decision tree model, splice the text vector and the feature vector to form a first comprehensive feature vector, and splice the feature representation and the feature vector to form a second comprehensive feature vector;
[0168] Input the first comprehensive feature vector and the second comprehensive feature vector into the decision tree model respectively, and obtain the first extraction data and the second extraction data output by the decision tree model;
[0169] Based on the first extraction data and the second extraction data, determine the bidding data corresponding to the current bidding document.
[0170] In a possible implementation manner of the embodiment of the present application, when the file type corresponding to the current bidding document is the third type, when the extraction module 202 performs data extraction on the current bidding document based on the file type to obtain the bidding data corresponding to the current bidding document, it includes:
[0171] Sequentially divide the current bidding document into multiple sequence segments, and sequentially input the sequence segments into the deep learning model;
[0172] Obtain the intermediate features corresponding to each sequence segment output by the deep learning model;
[0173] Calculate the correlation score corresponding to each intermediate feature as the attention weight corresponding to each intermediate feature;
[0174] Perform weighted summation based on the attention weight corresponding to each intermediate feature to obtain the weighted feature, and identify the weighted feature to obtain the bidding data corresponding to the current bidding document.
[0175] In a possible implementation manner of the embodiment of the present application, when the analysis module 203 analyzes the bidding data to obtain the analysis result corresponding to the current bidding document, it is specifically used for:
[0176] Obtain historical bidding data, where the historical bidding data includes the first historical bidding data of the current enterprise corresponding to the current bidding document and the second historical bidding data corresponding to the bidding data;
[0177] Based on the historical bidding moments corresponding to the first historical bidding data and the second historical bidding data respectively, perform time series analysis on the historical bidding data to obtain historical change patterns;
[0178] Based on the historical change patterns, determine the characteristic sub-data corresponding to the bidding data, and determine the correlation degree between every two pieces of characteristic sub-data;
[0179] Based on the characteristic sub-data and the correlation degree between every two pieces of characteristic sub-data, obtain the analysis result corresponding to the current bidding document.
[0180] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0181] See Figure 3 , this embodiment of the application also introduces an electronic device from the perspective of an entity device, such as Figure 3 shown, Figure 3 The electronic device 300 shown in the figure includes: a processor 301 and a memory 303. Among them, the processor 301 and the memory 303 are connected, such as connected through a bus 302. Optionally, the electronic device 300 may further include a transceiver 304. It should be noted that in actual applications, the transceiver 304 is not limited to one, and the structure of the electronic device 300 does not constitute a limitation to this embodiment of the application.
[0182] The processor 301 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, data signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of this application. The processor 301 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0183] The bus 302 may include a path for transmitting information among the above components. The bus 302 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 302 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 only a thick line is used in Figure 3 , but it does not mean that there is only one bus or one type of bus.
[0184] The memory 303 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0185] The memory 303 is used to store the application program code for executing the solution of this application, and is controlled and executed by the processor 301. The processor 301 is used to execute the application program code stored in the memory 303 to implement the content shown in the foregoing method embodiments.
[0186] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc., and can also be a server, etc. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.
[0187] The embodiments of this application provide a computer-readable storage medium on which a computer program is stored. When it runs on a computer, it enables the computer to execute the corresponding content in the foregoing method embodiments.
[0188] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0189] The above are only some embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A bidding data processing method, characterized in that: include: Receive a current bidding document and determine the file type corresponding to the current bidding document; Based on the file type, extract data from the current bidding file to obtain bidding data corresponding to the current bidding file; Analyze the bidding data to obtain analysis results corresponding to the current bidding document; Based on the analysis results, generate an analysis report corresponding to the current bidding document; Wherein, determining the file type corresponding to the current bidding document includes: Identify basic information corresponding to the current bidding document, wherein the basic information includes the issuing unit, the length of the document, and the document keywords, wherein the document keywords include time nodes, project budgets, and project requirements; Determine the development status of the issuing unit corresponding to the current bidding document, wherein the development status is mature development or immature development; If the development status of the issuing unit corresponding to the current bidding document is mature, determining whether the length of the document exceeds the length threshold; if the length of the document exceeds the length threshold, determining that the file type corresponding to the bidding document is the first type; Wherein, when the file type corresponding to the current bidding document is the first type, extracting data from the current bidding document based on the file type to obtain bidding data corresponding to the current bidding document includes: Determine an extraction weight corresponding to the file type, and determine an extraction mode corresponding to the file type based on the extraction weight, wherein the extraction mode includes an extraction model and an extraction number corresponding to each extraction model; Extracting the current bidding document based on each extraction model and according to the number of extractions corresponding to each extraction model to obtain initial data corresponding to the current bidding document; Performing a data check on the initial data to obtain a check result, wherein the check result includes the accuracy of the data; Iteratively extract and check the initial data based on the check result until a preset condition is met, and use the initial data that meets the preset condition as the bidding data corresponding to the current bidding document, wherein the preset condition is that the number of iterations reaches an iteration number threshold or the data accuracy is greater than an accuracy threshold; The extraction model includes a Go language model, a CNN language model, and an RNN language model. The current bidding document is extracted based on each extraction model and according to the number of extractions corresponding to each extraction model to obtain the initial data corresponding to the current bidding document, including: Preprocessing the current bidding document to obtain the preprocessed current bidding document; Based on the number of extractions corresponding to each extraction model, determining an extraction sequence corresponding to the current bidding document, wherein the extraction sequence includes at least two extraction model groups arranged in sequence, each extraction model group includes at least one extraction model, and the extraction models in each extraction model group are arranged in an extraction sequence; Determine a current extraction model group, perform an extraction step on the current bidding document, and obtain initial data corresponding to the current bidding document; The extraction step comprises: Determine whether the number of extractions of the Go language model in the current extraction model group is 0; if the number of extractions of the Go language model is not 0, input the preprocessed current bidding document into the Go language model, and obtain text data output by the Go language model; Determine whether the number of extractions of the CNN language model in the current extraction model group is 0, and if the number of extractions of the CNN language model is not 0, input the text data into the CNN language model, and obtain the first key data and local features output by the CNN language model; Determine whether the number of extractions of the RNN language model in the current extraction model group is 0, and if the number of extractions of the RNN language model is not 0, input the text data into the RNN language model, and obtain the second key data and sequence features output by the RNN language model; When the extraction times of the RNN language model in the current extraction model group is not 0 and the extraction times of the CNN language model is not 0, obtain the local features output by the CNN language model and the sequence features output by the RNN language model, concatenate the local features and the sequence features to obtain combined features, and extract the data corresponding to the combined features to obtain the third key data.
2. The bidding data processing method according to claim 1, characterized in that: The determining of the file type corresponding to the current bidding document further includes: If the length of the document does not exceed the length threshold, determining that the file type corresponding to the bidding document is the second type; If the development status of the issuing unit corresponding to the current bidding document is immature, then based on the identified file keywords, it is determined whether there is any inconsistency in keyword content in the current bidding document; if there is no inconsistency in keyword content in the current bidding document, it is determined that the file type corresponding to the current bidding document is the third type; if there is any inconsistency in keyword content in the current bidding document, it is determined that the file type corresponding to the current bidding document is the fourth type.
3. The bidding data processing method according to claim 2, characterized in that: When the file type corresponding to the current bidding document is the second type, extracting data from the current bidding document based on the file type to obtain bidding data corresponding to the current bidding document includes: Input the current bidding document into the Embedding model and the sparse coding model respectively, and obtain the text vector output by the Embedding model and the feature representation output by the sparse coding model; Obtaining a feature vector constructed based on artificial features corresponding to the decision tree model, and concatenating the text vector with the feature vector to form a first comprehensive feature vector, and concatenating the feature representation with the feature vector to form a second comprehensive feature vector; Inputting the first comprehensive feature vector and the second comprehensive feature vector into the decision tree model respectively, and obtaining the first extracted data and the second extracted data output by the decision tree model; Based on the first extracted data and the second extracted data, the bidding data corresponding to the current bidding document is determined.
4. The bidding data processing method according to claim 2, characterized in that: When the file type corresponding to the current bidding document is the third type, extracting data from the current bidding document based on the file type to obtain bidding data corresponding to the current bidding document includes: Sequentially dividing the current bidding document into multiple sequence segments, and sequentially inputting the sequence segments into the deep learning model; Obtaining intermediate features corresponding to each sequence segment output by the deep learning model; Calculate the relevance score corresponding to each intermediate feature as the attention weight corresponding to each intermediate feature; A weighted sum is performed based on the attention weight corresponding to each intermediate feature to obtain a weighted feature, and the weighted feature is identified to obtain the bidding data corresponding to the current bidding document.
5. The bidding data processing method according to any one of claims 1 to 4, characterized in that: The step of analyzing the bidding data to obtain the analysis result corresponding to the current bidding document includes: Acquire historical bidding data, the historical bidding data including first historical bidding data of the current enterprise corresponding to the current bidding document and second historical bidding data corresponding to the bidding data; Based on the historical bidding moments corresponding to the first historical bidding data and the second historical bidding data, respectively, a time series analysis is performed on the historical bidding data to obtain a historical change rule; Based on the historical change rules, determining the characteristic sub-data corresponding to the bidding data, and determining the degree of association between every two characteristic sub-data; Based on the characteristic sub-data and the degree of association between every two characteristic sub-data, an analysis result corresponding to the current bidding document is obtained.
6. A bidding data processing device, characterized in that: include: A receiving module, used to receive the current bidding document and determine the file type corresponding to the current bidding document; An extraction module, used to extract data from the current bidding document based on the document type, so as to obtain bidding data corresponding to the current bidding document; An analysis module, used to analyze the bidding data to obtain analysis results corresponding to the current bidding document; A generating module, used to generate an analysis report corresponding to the current bidding document based on the analysis result; Wherein, when determining the file type corresponding to the current bidding document, the receiving module is specifically used to: Identify basic information corresponding to the current bidding document, wherein the basic information includes the issuing unit, the length of the document, and the document keywords, wherein the document keywords include time nodes, project budgets, and project requirements; Determine the development status of the issuing unit corresponding to the current bidding document, wherein the development status is mature development or immature development; If the development status of the issuing unit corresponding to the current bidding document is mature, determining whether the length of the document exceeds the length threshold; if the length of the document exceeds the length threshold, determining that the file type corresponding to the bidding document is the first type; Wherein, when the file type corresponding to the current bidding document is the first type, the extraction module extracts data from the current bidding document based on the file type to obtain the bidding data corresponding to the current bidding document, specifically for: Determine an extraction weight corresponding to the file type, and determine an extraction mode corresponding to the file type based on the extraction weight, wherein the extraction mode includes an extraction model and an extraction number corresponding to each extraction model; Extracting the current bidding document based on each extraction model and according to the number of extractions corresponding to each extraction model to obtain initial data corresponding to the current bidding document; Performing a data check on the initial data to obtain a check result, wherein the check result includes the accuracy of the data; Iteratively extract and check the initial data based on the check result until a preset condition is met, and use the initial data that meets the preset condition as the bidding data corresponding to the current bidding document, wherein the preset condition is that the number of iterations reaches an iteration number threshold or the data accuracy is greater than an accuracy threshold; The extraction model includes a Go language model, a CNN language model and an RNN language model. When the extraction module extracts the current bidding document based on each extraction model and according to the number of extractions corresponding to each extraction model to obtain the initial data corresponding to the current bidding document, it is specifically used to: Preprocessing the current bidding document to obtain the preprocessed current bidding document; Based on the number of extractions corresponding to each extraction model, determining an extraction sequence corresponding to the current bidding document, wherein the extraction sequence includes at least two extraction model groups arranged in sequence, each extraction model group includes at least one extraction model, and the extraction models in each extraction model group are arranged in an extraction sequence; Determine a current extraction model group, perform an extraction step on the current bidding document, and obtain initial data corresponding to the current bidding document; The extraction step comprises: Determine whether the number of extractions of the Go language model in the current extraction model group is 0; if the number of extractions of the Go language model is not 0, input the preprocessed current bidding document into the Go language model, and obtain text data output by the Go language model; Determine whether the number of extractions of the CNN language model in the current extraction model group is 0, and if the number of extractions of the CNN language model is not 0, input the text data into the CNN language model, and obtain the first key data and local features output by the CNN language model; Determine whether the number of extractions of the RNN language model in the current extraction model group is 0, and if the number of extractions of the RNN language model is not 0, input the text data into the RNN language model, and obtain the second key data and sequence features output by the RNN language model; When the extraction times of the RNN language model in the current extraction model group is not 0 and the extraction times of the CNN language model is not 0, obtain the local features output by the CNN language model and the sequence features output by the RNN language model, concatenate the local features and the sequence features to obtain combined features, and extract the data corresponding to the combined features to obtain the third key data.
7. An electronic device, characterized in that: The electronic device includes: at least one processor; Memory; At least one application, wherein the at least one application is stored in a memory and configured to be executed by at least one processor, and the at least one application is configured to: execute the bidding data processing method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed in a computer, the computer is caused to execute the bidding data processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Video processing method and device, equipment and storage medium
CN115223083A
Enterprise data management platform and method
CN117667841A